1. Key conclusions
- The most useful replication targets this week are execution improvements whose effects can be explained, rather than a single highest success rate. PACE turns semantic intent at transitions such as stairs and doorways into traversable poses; SeekVLN actively observes when evidence is insufficient; EdgeVLN evaluates both a complete continuous-navigation benchmark and device resources. Our assessment: these approaches address different bottlenecks: knowing where to go but being unable to pass, moving before seeing enough, and running in simulation but failing to fit on the device. Diagnose them separately.
- R2R-CE results keep increasing, but information access and evaluation scope differ substantially. Among the 242 independent works, 8 match the R2R-CE metadata field; EdgeVLN additionally specifies “R2R VLN-CE,” bringing the total to 9 works involving this continuous-environment benchmark. Of these, 8 give absolute SR/SPL results or improvement claims, while 1 only names the benchmark. EdgeVLN reports SR 58.02% over all 1,839 val-unseen episodes; InsightMap reports 56.9%/54.9% on R2R-CE/RxR-CE val-unseen; AVERT-VLN’s 76.2%/66.3% includes human correction. PACE evaluates only a cross-floor subset. PanoVLN and SeekVLN express improvements with percentage signs in their abstracts, without clarifying relative gains versus percentage points. Our assessment: these figures do not support a unified leaderboard.
- Harnesses are starting to be tested on how experience transfers across tasks. Lifelong-navigation NavHarness reports GOAT-Bench s-SR 83.7 and e-SR 36.9 with SLAM-estimated poses, and compares structured recovery handoffs against equal-length summaries. MemTransfer finds that trajectory memory with high success at the original demonstration start loses 48–49 percentage points after changing the start. Our assessment: long-term memory is valuable when it remains usable after changes to the start, route, and goal; storing more is insufficient evidence.
- Spatial memory is incorporating evidence about when an object was seen and whether it could have been seen. ECROM interprets detections and non-detections through observation opportunities. On a long-term retrieval benchmark spanning ten HM3D homes, it improves support-level AP by 4.5 points and search SPL by 4.2 points over the strongest memory baseline for each metric. EvolvingNav predicts the target location at the arrival and inspection time, retaining probability that the target has moved outside known candidate locations. Our assessment: persistent navigation needs beliefs conditioned on time and visibility; a static last-seen location cannot adequately handle movement and occlusion.
- Safety, correct stopping, and continuous operation are becoming separate evaluation targets. EdgeVLN’s stopping head directly reuses backbone hidden states; the social-navigation world model infers risk from the gap between nominal and executable actions; STARS exposes remaining weaknesses in social-scene understanding relative to rule-based baselines. Our assessment: report task completion, timely response, collision avoidance, and correct stopping separately. Reasoning evaluated with physical simulation paused does not establish deployment capability.
This collection added 242 records corresponding to 242 independent arXiv works, with 0 new WeChat articles. Automatic primary-pool assignments comprise 48 navigation works and 33 embodied-agent works, totaling 81; the secondary pool contains 133 VLA/manipulation works, 8 autonomous-driving works, and 20 others, totaling 161. The automatic primary pool still includes boundary cases such as code retrieval, loss landscapes, and computer operation. We retain these as low-relevance records rather than evidence for ground navigation. The coverage interval reflects the actual publication dates of newly collected works; the newest entry is dated 10-02.
2. Priority reading list
- EdgeVLN · Continuous R2R-CE and device deployment
- Contribution: jointly optimizes quantization, runtime, memory pruning, and stopping decisions.
- Evidence: all 1,839 val-unseen episodes; 4-bit model SR 58.02%; resident memory 11.35 GB on Orin NX 16 GB; the stopping head adds 0.013 seconds per step.
- Why read: evaluation scope and resource conditions are especially explicit, making it a useful deployment reference.
- PACE · Continuous R2R-CE/RxR-CE cross-floor subsets
- Contribution: connects frozen semantic planners to short-horizon actions through traversable poses, learning corrections from failure preferences.
- Evidence: mean SR across six zero-shot navigators rises from 16.35% → 27.65% on the R2R-CE subset and 4.76% → 12.06% on the RxR-CE subset.
- Why read: directly addresses doorways, stairs, and narrow transitions where ground robots often fail.
- SeekVLN · Active evidence acquisition in continuous R2R-CE/RxR-CE
- Contribution: judges whether progress evidence is sufficient before observing or moving; counterfactual branches from the same state assign rewards to evidence acquisition.
- Evidence: the abstract claims SR gains of 12.7%/7.5% over the base model, without absolute SR or clarification of percentage units.
- Why read: directly connects the value of perception actions to subsequent navigation gains.
- PanoVLN · Panoramic RGB continuous navigation
- Contribution: jointly designs long action sequences, confidence-based execution, branching-route supervision, and semantic-geometric representations.
- Evidence: 4B, RGB-only; the abstract claims R2R-CE/RxR-CE val-unseen SR exceeds the previous best by 11.9%/8.7%; absolute results and gain units require verification.
- Why read: panoramic perception benefits arise from a complete policy design that merits component-wise ablation.
- InsightMap · Explicit maps and continuous navigation
- Contribution: a shared backbone learns navigation actions alongside post-action map generation.
- Evidence: R2R-CE/RxR-CE val-unseen SR 56.9%/54.9%; map-prediction supervision improves R2R-CE SR/SPL by 4.3/3.2 percentage points.
- Why read: direct navigation ablations help test the causal contribution of auxiliary spatial supervision.
- FINE · Data efficiency in continuous navigation
- Contribution: uses future landmarks in existing demonstrations to learn semantic, geometric, and counterfactual future representations.
- Evidence: with the full training set, InternVLA-N1 gains 2.6/4.5 percentage points in R2R-CE/RxR-CE val-unseen SR.
- Why read: added supervision comes from existing trajectories, making the approach relevant to limited demonstration budgets.
- AVERT-VLN · Human-assisted recovery in continuous navigation
- Contribution: an independent monitor asynchronously detects deviations, requests human help, and converts correction interactions into local preference learning.
- Evidence: human-assisted R2R-CE/RxR-CE val-unseen SR 76.2%/66.3%.
- Why read: useful for studying when to ask for help; quoted results must retain the human-assistance condition.
- NavHarness: Towards Lifelong Embodied Navigation · Cross-task navigation and persistent memory
- Contribution: maintains maps, task records, and corrections across sessions, with structured handoffs for recovery.
- Evidence: GOAT-Bench with SLAM-estimated poses gives s-SR 83.7 and e-SR 36.9; IR2R-CE s-SR 85.9 belongs to a different continuous-navigation task.
- Why read: direct navigation evidence from the embodied-agent track; inspect the gap between whole-episode and subtask completion.
3. Analysis of key works
1. EdgeVLN: evaluate quantization together with execution paths and stopping
Problem. A smaller model can still be undeployable because of memory transfers, growing history, and stopping errors.
Method. Quantizes StreamVLN, reconstructs streaming context, and prunes memory tokens. The LATTE stopping head reuses hidden states without an additional visual encoder or second backbone forward pass.
Evidence. The abstract reports 4-bit SR 58.02% on the complete R2R-CE val-unseen set and resident device memory of 11.35 GB. Even within four-bit formats, different execution paths produce up to a 36.8-fold difference in per-step energy. The 20.8-fold speed and 13.3-fold energy improvements use storage-streamed BF16 as the reference. Two-bit quantization fails.
Value. Our assessment: deployment gains depend on model and runtime choices together; evaluate the stopping head within the same resource budget.
Limitations. Navigation SR comes from simulation, while resource measurements come from Orin NX; these do not constitute equally extensive physical-route validation. The abstract omits absolute total per-step latency and complete SPL.
Recommendation. Replicate the four-bit execution paths and stopping head, recording SR, SPL, stopping errors, memory, per-step latency, and energy consistently.
2. SeekVLN: active observation must improve subsequent actions
Problem. A robot may confidently move before seeing a turning landmark, carrying an incorrect progress estimate into later decisions.
Method. Generates supplementary views and evidence labels from offline expert trajectories, then compares downstream returns from evidence-acquisition and direct-movement branches starting at the same state to learn when to observe.
Evidence. The abstract claims R2R-CE/RxR-CE SR gains of 12.7%/7.5%, without specifying gain units, absolute results, or observation costs.
Value. Our assessment: counterfactual branches provide relatively direct credit assignment for evidence acquisition and can test whether observation prevents later deviations.
Limitations. The abstract does not quantify observation steps, inference costs, or physical-robot gains. Training uses future expert actions, so inference inputs and data budgets need verification.
Recommendation. Compare fixed observation, confidence-triggered observation, and counterfactual training on the same policy, including head turns, pauses, and total elapsed time.
3. PACE: correct plans still need spatially grounded local execution
Problem. Reaching stairs or a doorway does not guarantee passage; high-level semantic goals lack executable pose and path conditions.
Method. Generates robot-centered traversable poses followed by short-horizon actions. Execution failures supply preference pairs involving normal behavior, recovery, and amplified deviations.
Evidence. The paper reports mean cross-floor-subset SR across six frozen zero-shot navigators: R2R-CE 16.35% → 27.65%, RxR-CE 4.76% → 12.06%.
Value. Our assessment: traversable poses offer an inspectable interface between high-level intent and low-level control, helping distinguish recognition errors from execution errors.
Limitations. Results cover cross-floor subsets only. The local module receives supervised and preference training, so the complete system is not entirely training-free. The abstract does not provide physical-trial counts or SR.
Recommendation. Replicate transition regions separately, distinguish goal localization, passage failures, and recovery after deviations, and retain a purely geometric execution baseline.
4. AVERT-VLN: human-assisted SR alone cannot establish a monitor’s value
Problem. A controller may continue incorrect execution after departing the route, while constant human supervision is difficult to scale.
Method. An independent monitor detects semantic deviation using the instruction, history, and current observation, asynchronously issuing a LOST decision. Once accepted, the controller pauses and requests guidance, then learns preferences from the corresponding decisions.
Evidence. The abstract gives human-assisted R2R-CE/RxR-CE val-unseen SR 76.2%/66.3%. Risk data contain 20K counterfactual trajectories, while progress pretraining uses 40K normal trajectories.
Value. Our assessment: monitoring, help requests, and offline improvement are independently reusable modules for testing whether infrequent human intervention is cost-effective.
Limitations. The abstract omits help-request counts, human information volume, false-positive rates, and autonomous results under the same budget. Recovery-system performance cannot be attributed solely to the navigation policy.
Recommendation. Fix human minutes or correction counts and compare random help requests, policy confidence, and independent monitoring. Report false positives, missed detections, and recovery success.
5. Lifelong-navigation NavHarness: separate subtask and whole-episode success
Problem. Cross-task reuse encounters incomplete maps, conflicts between old records and new observations, and state loss when restarting after failure.
Method. Checks and corrects maps, search records, and house knowledge within sessions, preserving experience across new sessions. Recovery receives structured state handoffs, followed by consolidation after task completion.
Evidence. This work reports GOAT-Bench s-SR 83.7 and e-SR 36.9 with SLAM-estimated poses. Relative to independent sessions that only retain context, Astra/Opus 5 gain 18.6/22.6 points in s-SR. It is a different arXiv work from this week’s adaptive-goal NavHarness; do not merge them by method name.
Value. Our assessment: comparing structured recovery handoffs with equal-length summaries explains the reuse mechanism better than simply comparing memory on versus off.
Limitations. High subtask performance alongside low whole-episode performance leaves long-term stability unresolved. IR2R-CE cannot substitute for R2R-CE route-following results. The abstract does not fully specify closed-model costs or experience-access budgets.
Recommendation. Combine this evaluation with MemTransfer-style changes of start, blocked routes, and irrelevant-history perturbations to distinguish transferable knowledge from accumulated records of repeated goals.
4. Transferable methods
- Progress estimation: ProgressCompass.
- Source: manipulation benchmark ContextProgress-Bench, with 24 tasks and 120 episodes. The authors report that adding correct context reduces errors by 77–82% across the same five progress models.
- Integration point: navigation sub-instruction completion; explicitly provide the monitor with completed actions, revisited rooms, and historical landmark states.
- Our assessment: first use manually verified context as a diagnostic upper bound, then test automatically generated context. Manipulation progress error is not navigation SR.
- Executable skill transfer: RoboBridge.
- Source: LIBERO-PRO and corresponding physical manipulation tasks. Task intent, observations, tool calls, and outcome verification become revisable programs; the abstract has no verifiable numerical results.
- Integration point: express navigation recovery as skills with preconditions and outcome checks, retaining general structure while revising environment-dependent steps.
- Our assessment: transfer requires stable action, pose, and stopping interfaces in navigation tools; manipulation experiments do not establish route generalization.
- Asynchronous timing: DiffWAM / FastDreamer.
- Source: drone navigation. Predicted features from a frozen video model feed directly into trajectory generation, with asynchronous timestamped handoffs. The Flash pipeline has 1.08-second latency on Jetson AGX Thor.
- Integration point: when high-level visual reasoning overlaps low-level execution on a ground robot, verify the observation time and current execution progress associated with each candidate trajectory.
- Our assessment: reuse timestamps and trajectory handoffs first. Flight-trajectory RMSE or endpoint success cannot replace VLN route metrics.
- Local interaction evidence: JRDB-AVR.
- Source: active visual question answering from real robot video, evaluating both answers and supporting observation evidence; the abstract gives no verifiable accuracy.
- Integration point: evaluate visible evidence for navigation judgments such as reaching a turn or identifying the target doorway.
- Our assessment: useful for diagnosing correct guesses based on common sense; question-answering performance does not directly establish closed-loop navigation success.
5. Classified overview
A denotes direct ground language/semantic navigation, B a clear transferable mechanism from a different task, and C low-relevance observations. This overview covers all 242 independent works. Automatic pool assignment is distinct from research relevance: primary-pool boundary cases appear under their actual low-relevance tasks, while secondary-pool works are indexed by theme.
5.1 Ground VLN / ObjectNav / language and semantic navigation
- PACE (A): Traversable-pose execution across floors; see the detailed analysis.
- FINE (A): Uses semantic, geometric, and counterfactual representations of future landmarks to improve demonstration efficiency.
- RAO-Nav (A): Uses an omni-modal language model for zero-shot semantic audio-visual navigation, with latent reasoning guiding relevant observation acquisition.
- RECAST (A): Grounds traversable surfaces, risk, orientation, and corridor judgments in cost maps.
- Query, Align, and Distill (A): Explicit navigable queries form a teacher-student distillation interface; the abstract omits specific benchmarks and navigation figures.
- EdgeVLN (A): Jointly evaluates the complete continuous benchmark and Orin device resources; see the detailed analysis.
- NavHarness: Lifelong Navigation (A): Maintains maps across sessions with structured recovery handoffs; see the detailed analysis.
- NavJev (A): Compresses per-step multimodal generation into candidate-action evidence and structured selection; R2R-CE SR 27.0%, SPL 22.4%, and 0.65 seconds per step; the abstract does not specify the split.
- PanoVLN (A): Jointly designs panoramic visibility, long actions, confidence-based execution, and branching supervision.
- Reliability-Aware Route Memory (A): Queries outbound geometric anchors in reverse order, combining action arbitration with endpoint verification; only 50 reverse-paired episodes.
- SOR-Nav (A): Explicitly chooses between searching the current region and relocating across regions; full MP3D validation SR/SPL 61.8%/38.5%.
- BCNav (A): Decouples sound-direction estimation from the depth-navigation policy and outputs continuous ground-robot velocities.
- CGPI (A): Extracts credit from action-induced observation changes, retains verified adaptation updates, and rolls back unsupported updates.
- InsightMap (A): Uses explicit maps both as historical references and as auxiliary post-action prediction targets.
- Risk-Aware Semantic Grounding (A): Distinguishes ambiguity, hallucination, and semantic conflict before planning to choose execution, clarification, or refusal.
- SeekVLN (A): Actively acquires evidence when progress information is insufficient; see the detailed analysis.
- Astra Cross-Domain Embodied Policy Evaluation (A): The abstract reports RxR SR 92% and HM3D object-search SR 82%, without specifying complete splits or CE. Physical simulation waits for inference in control examples, so these are not evidence of real-time navigation.
- ASENA (A): A coding agent calls an optional 4B navigation policy and saves skills; the abstract does not explicitly specify CE for R2R/RxR. Ten passes over a 100-task subset cannot substitute for a single independent test.
- AVERT-VLN (A): Asynchronous deviation monitoring and human-guided recovery; see the detailed analysis.
- NavHarness: Adaptive Goals (A): Separates goals, verification, memory, and execution; the abstract has no R2R-CE/RxR-CE figures. Eight physical routes are each evaluated three times, with SR 83.3% and navigation error 1.51 meters.
- UniTrackPLA (A): Unifies language navigation and dynamic person tracking through panoramic spatiotemporal encoding and future-consistency checks.
- GeoScaffold (A): Reconstructs depth, connectivity, and traversability during training, removing geometric-supervision components at deployment; the abstract gives no specific benchmark figures.
5.2 Memory, maps, planning, social navigation, and evaluation
- TRACKGRAPH (B): Tracks short-term mask identities in image streams, then fuses them into open-vocabulary 3D scene graphs.
- VideoSocNav (A): Reconstructs traversable maps and pedestrian motion in policy state space from online walking videos, reducing dependence on photorealistic simulation.
- MemTransfer (A): Tests actual memory transfer through changed starts, blocked routes, and history-relevance controls.
- 3D Point Tracking with State Space Models (B): A fixed-size recurrent state corrects monocular metric depth, offering a possible dynamic-geometry frontend.
- EM-EQA Viewpoint Selection (B): Projects panoramas into perspective views, selecting historical evidence by question relevance and diversity.
- HEIR (B): Jointly evaluates complete human-entity events and local relations; closed-loop navigation benefits remain unverified.
- JRDB-AVR (B): Evaluates active visual question answering together with supporting evidence; see transferable methods.
- General Asynchronous Agents (B): Uses concurrent reasoning coroutines for streaming video, games, and monitoring; the abstract has no ground-navigation experiments.
- Multi-Scale Semantic Mapping (B): Calibrates urban observations by category and distance, reducing redundant coupling between mapping policies.
- SAIL (B): Preserves correspondences between events, directions, and distances by sound source, offering observations for audio-visual navigation frontends.
- MAVLN / TRISS (A): Introduces dependencies and resource constraints into multi-robot navigation, combining shared topological memory with conflict handling.
- VCN-Bench (A): Distinguishes goal recognition in prior videos from closed-loop arrival on MP3D, with 1,250 evaluation episodes.
- Human Motion Prediction During Daily Tasks (B): Analyzes inertia, occupancy, semantics, gaze, and explicit intent in indoor human-motion prediction.
- DeCOD LiDAR SLAM (B): Uses corridor cross-section landmarks to constrain degenerate axial drift; does not involve a language-navigation policy.
- RGB-Only CBF Distillation (B): Distills dynamic obstacle avoidance from a privileged teacher into a safety filter using only RGB history and velocity.
- BRAID / Generative Interactions (B): Models group interaction states separately from individual variation, offering a potential context interface for social navigation.
- Learning to Plan from Random Exploration (B): Learns multiscale reachability from temporal relations in random exploration, without action or reward labels for training the relation model.
- ECROM (A): Calibrates detections and non-detections through observation opportunities, enabling long-term object search with concepts specified only at query time.
- EvolvingNav (A): Time-indexed location beliefs, arrival-time prediction, and visibility-conditioned updates from negative observations.
- DODGER (B): Uses CBF references and constraint violations to guide policy training, with no runtime safety filter at deployment.
- Terrain Traversability Continual Learning (B): Learns traversability measures such as slip from contact experience, using a validation gate to limit degradation on historical performance.
- Hallway Legibility (A): Two studies with 45 participants each compare passing-side intent, goal intent, and pedestrian distraction conditions.
- Social-WM (A): The gap between predicted latent futures and executable actions provides a safety signal for social navigation.
- STARS / SocialNav-SUB (B): Question answering in real social-navigation scenes tests spatiotemporal relations and intent understanding; the best VLM still trails some rule-based and human baselines.
- Uruqi (B): Continuous visual experience jointly supervises ego-motion tracking, persistent object mapping, and spatial reasoning.
- Spatial Memory Intelligence (B): Spatial clustering, sparsification, action-relevant retrieval, and reliability filtering for long-term world-model memory.
- Token Communication for CEAI (B): Studies task-driven semantic-token communication and cooperative object transport; navigation transfer remains unverified.
- OmniAct3D (B): Handles panoramic versus perspective geometry, supporting three-dimensional detection with local evidence.
- Representational Alignment (B): Theoretically adds sensorimotor anchors only for representational differences affecting current interaction outcomes.
Navigation and spatial-perception boundary works from the secondary pool remain indexed; they do not underpin the primary-pool analysis:
- Navigation and spatial perception (B/C), 7 works: InfraVLA, TUDF Scene Completion, S4VY, PERSEPHONE Spatial Perception, WayFinder, GroundingPI, PAGER.
5.3 Embodied agents, VLA, and mobile manipulation
Agent and manipulation works from the primary pool:
- SciHorizon-eLab (C): Compiles experimental protocols into verifiable long-horizon laboratory manipulation tasks.
- RoboFoundry (B): Evolves context and skill systems as a verifiable whole policy, with main evidence from general embodied and manipulation tasks.
- Beyond Tasks (C): A position paper on persistent behavior coordination and long-term interaction, without navigation-benchmark evidence.
- Robot-GST (B): An RGB-D-reconstructed Gaussian-SAM environment supports pre-execution simulation and outcome verification.
- SkillWeaver (B): An agent explores closed-loop interactive skills, generating manipulation demonstrations through verifier-guided tree search.
- RoboSkill (B): A loop of exploration, execution, reuse, and evolution stores manipulation experience in text and code.
- ProgressCompass (B): Explicitly supplies context needed for progress estimation; see transferable methods.
- RobotEQ 3.0 (C): Predicts user expectations for proactive assistance from individual characteristics.
- Video2Skill (B): Diagnoses skill classification, reuse, and expansion from streaming observations; integrating familiar skills does not mean learning new skills.
- ChronoGraph (B): Writes past and anticipated actions, affordance-bearing parts, and state changes into a shared four-dimensional graph interface.
- Game-Guided Skill Discovery (B): Self-play produces composable skills accessible to human control; this is not yet a language-navigation policy.
- Embodied Agent Arena (B): Uses 1,000 cases to separately measure geometric precision, functional grounding, and complete task success.
- PyRUA-Lean (B): Conditional programs compose primitives with feedback on demand; across 700 manipulation cases with the same call budget, success rises from 63.1% → 71.7%.
- RoboBridge (B): Continues sim-to-real learning through verified skill revisions; see transferable methods.
The following secondary-pool themes are not compared as ground-VLN results; each work appears in exactly one theme:
- Failure recovery and system evolution (C), 18 works: Causeway, FIND, Kintsugi-VLA, SEES, ActionGround, Recursive Harness Distillation across Agents for Robot Manipulation, F4R, FailPatch, Self-Evolving Coding Agents, MotorMind, ProAct-VLM, Skill-Space Shooting for Autonomous Robot Policy Improvement, FailBank, InterEvolve, Recova, SocialVLA, MobiAgent, Execution Error Compensation.
- Action interfaces, chunking, and policy structure (C), 11 works: Fast Plans, Faithful Actions, DS-VLA, ActionUNet, Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning, Quantile Head for Vision-Language-Action Models, CATok, Discrete Forcing, DSDyn-VLA, ChunkVLA-AM, TOAST, Linear Representation Hypothesis.
- Quantization, distillation, and low-latency inference (C), 12 works: FRAM, Action Upcycling, EdgeDAE, RAVEL, Token Caching, DriftOPD, Asynchronous Distribution Alignment, Urgency-Aware Denoising, Spike-driven VLA, Two-Step Flow Denoising, CHASE-VLA, FastOPD.
- Language grounding and compositional generalization (C), 10 works: GT-VLA, Spatial Grafting, Layer Selection, Referential Guidance, RawVLA, Cue the Flow, Same Scene, Different Task, Instruction-Action Binding / ECT, WorldAuditBench, MixVLA.
- Future prediction and world models (C), 16 works: Towards VLA-Dreamer, Devol-ONE, SLIP-VLA, RoboFL, WorldGuide, DILL, V-JEPA Policy, Predictive Supervision Placement, EWAM, MotionWeave, Planning Limits of Latent World Models, Token-World, ATI-VLA, UniWAM, World-Calibrated Proposal-to-Action, IG-VLA.
- Reinforcement learning and policy optimization (C), 14 works: VLaRL, PF-RL, SAMBAR, Principal Steering Subspaces for Online Adaptation of Frozen Generative Robot Policies, TimelyDAgger, PolicyWeave, Adjoint Guidance Flow, ChronoSRL, StructRL, Low-Rank RL, UTMPO, Online-ES, PRICE the Action Chunks, eRLT.
- Safety, faults, and behavioral diagnostics (C), 17 works: One-Step Observation Perturbations, Adversarial Training / View Collapse, Do Not Cut When Uncertain, MAIL-Bench, State Readout Diagnostics, RoboIRGBench, Acceleration Benchmark Diagnostics, Memorize, Adapt, Ignore, Blackout vs. Freeze, Exploiting Vulnerabilities, Multi-Link Safety Filtering for VLA Policies Around Moving Hazards, Behavioural Robustness Evaluation, WBAG, Detect and Suppress, ManiPhysicsBench, Multi-Agent Action Collapse, Reactive Obstacle Avoidance.
- Touch, human input, and specialized manipulation (C), 9 works: CLAP, BrainVLA, Gaze Prompts, mmHRI, Tactile Curiosity, EMG Task Conditioning, RoboChemGym, SARI, SimpleTouch.
- Memory and persistent state (C), 11 works: RecastVLA, SSVR, D²-VLA, Ledger, LexiconVLA, Action-History Memory, T²Mem, ECoMEM, Optimus-R, MIKASA-Robo-VLA, Divide-and-Remember.
- Humanoids, dual arms, mobile manipulation, and cooperation (C), 11 works: Fiatlux, TAO-DA, Humanoid Loco-Manipulation With Discrete VLA Model, Uni-VLaT, Cooperative Multi-Agent VLA, EgoAlign, EgoHumanoid-V2, FineART, IronMind, Whole-Body Human Pretraining, DuoMind.
5.4 Drones, autonomous driving, and other low-relevance directions
- SatNav (C): Builds city-scale drone VLN from satellite imagery; evaluation is not from a ground viewpoint.
- SemNav: Code Repository (C): “Navigation” for locating issues in code repositories, outside embodied navigation.
- Wireless Evidence Acquisition (C): Deadline-constrained multi-sensor wireless scheduling, without closed-loop VLN validation.
- AquaBEV-Nav (C): Predicts BEV occupancy directly from underwater monocular images for exploration.
- AquaWAM (C): Jointly predicts underwater actions and passive dynamics; results do not map directly to ground navigation.
- DroneWAM (B): A drone latent world model adapts prediction depth to the scene; ground applicability needs separate testing.
- ForeFly (B): Distinguishes near-term futures from route-critical futures in drone VLN; retained as a mechanism to watch.
- Just-In-Time Agent Memory (C): General history retrieval and query-time context construction; the abstract has no embodied-navigation evidence.
- Normative Loss Landscape Navigation (C): “Navigation” in parameter loss landscapes, outside robot spatial navigation.
- VehicleArena (C): Evaluates multi-vehicle driving with independent goals and traffic externalities.
- Pixels to Keys (C): Infers actions from game videos, without ground-robot experiments.
- DiffWAM (B): Connects drone prediction features to trajectories through asynchronous timestamped handoffs; see transferable methods.
- OSWorld-Science (C): Computer operation of scientific software, outside physical embodied navigation.
- Reward as Observation (C): Transfers using reward and action histories, which cannot be directly integrated into deployed navigation that typically lacks dense rewards.
- MEAN Movement and Compression (C): A communication model jointly optimizing movement, compression, and power, without navigation-benchmark experiments.
- LiDARFlow (C): Airborne LiDAR potential-flow geometric obstacle avoidance, without language or semantic-navigation mechanisms.
Remaining secondary-pool directions:
- Flight and pure control (C), 4 works: LQR-ArUco Fusion, AeroManip-VLA, Scene-Scale Aerial Manipulation, RL-Guided PAC-NMPC.
- Driving and traffic (C), 9 works: CausalDriveBench, RCVLA, CAR-VLA, RefineDrive, Embodied Reasoning Interfaces, Class 8 Truck VLA, Speed in the Blind Spot, Vision-Language-Action Autonomous Driving Agent with Language-based Memory, VLALight.
- General agents, games, and surveys (C), 9 works: CueKFS, GameBoyWorlds, ExpVoyager, Reflect Reverse, AutoDataBench, EngramBench, JevSpawn, Coco, Evolutionary Computation Survey.
- Morphology, collective behavior, and medical robotics (C), 3 works: Bridging Body and Brain, Fish Schools, Endovascular BCI Navigation.
5.5 News and non-paper resources
- No new WeChat articles or independent non-paper news were added this time; Docker was not running and the manual WeChat URL list was empty.
6. Trends and recommended actions
Trends
- Training supervision and inference budgets are increasingly optimized separately. FINE extracts supervision from future landmarks, GeoScaffold internalizes geometry during training, EdgeVLN measures deployment execution paths, and NavJev replaces per-step generation with structured action selection. Our assessment: report both additional training resources and actual closed-loop costs.
- Memory is moving from trajectory records toward conditional decision evidence. MemTransfer exposes the effects of start and route changes; ECROM calibrates observation opportunities; EvolvingNav advances time-dependent beliefs; lifelong-navigation NavHarness checks old records against new observations. Our assessment: future evaluations should cover memory failures and correction, rather than repeated tasks alone.
- Monitoring, evidence acquisition, and recovery can be studied independently. SeekVLN’s active evidence acquisition, AVERT-VLN’s help-request monitor, and CGPI’s update rollback target insufficient information, execution deviation, and incorrect adaptation respectively. Our assessment: compare module effects under the same base policy and additional budget.
- Larger embodied-model results remain constrained by execution and evaluation conditions. This week’s cross-domain Astra evaluation does not fully specify the RxR/HM3D splits or CE setting in the abstract; motion-control examples pause physical simulation while waiting for inference. Embodied Agent Arena and VCN-Bench also distinguish local judgments from final task completion. Our assessment: evaluate precise estimation, correct goal localization, and complete action success separately.
Research gaps
- Under a shared budget, how much are observation, panoramas, and human help worth? PanoVLN expands the field of view, SeekVLN adds evidence acquisition, and AVERT-VLN introduces human correction. Comparisons on the same routes with equal perception, inference, and human costs are missing.
- Can systems detect and repair incorrect maps or memories themselves? Average abstract results do not explain the cost of actively correcting old experience, the risk of mistaken corrections, or worst-case task losses.
- How does device latency affect outcomes in dynamic environments? Monitoring, stopping decisions, evidence acquisition, and low-level safety control must all run on the real clock; reasoning evaluated only while simulation is paused is insufficient.
Recommended actions
High priority
- Replicate EdgeVLN’s stopping and deployment comparisons: first fix the complete R2R-CE val-unseen set, execution formats, and common hardware, then record SR/SPL, stopping errors, resources, and total elapsed time.
- Read PACE and SeekVLN closely: verify PACE’s subset definition and local-module training data; verify SeekVLN’s gain units, absolute results, and active-observation costs before choosing an execution module or evidence-acquisition policy.
- Verify PanoVLN and FINE’s resource conditions: inspect panorama acquisition, branching-route data, and whether future supervision introduces extra expert information. FINE’s low-data-budget gains are not explicitly tied to benchmarks in the abstract, so defer quantitative cross-method comparisons.
- Replicate lifelong-navigation NavHarness’s experience handoffs: standardize task sequences, history access, and restart budgets, adding MemTransfer-style changed starts, blocked routes, and irrelevant memories.
Medium priority
- Track InsightMap, ECROM, and EvolvingNav: these target post-action spatial supervision, observation-opportunity calibration, and arrival-time beliefs respectively. Validate each module before combining them.
- Diagnose the AVERT-VLN monitor: fix the human-intervention budget and separate detection, recovery, and offline-training contributions.
- Verify code and data ownership for both NavHarness works: lifelong navigation is 2609.34276 and adaptive goals is 2609.39915. Keep complete titles and IDs in replication records.
- Inspect ASENA’s protocol: the entry uses R2R/RxR names without explicitly specifying CE. Record the 100-task subset results from ten repeated passes separately from a single independent test.
Low priority
- Watch VLA memory, quantization, and recovery: maintain a thematic secondary-pool index; promote a work for close reading when it offers a clear ground-navigation integration point and corresponding closed-loop comparisons.
- Defer choosing systems by highest SR: first verify conditions for human assistance, subsets, panoramas, RGB-D, repeated experience learning, model choices, and inference settings.
- Defer pure-control and non-embodied “navigation” entries: they do not provide evidence for this week’s language or semantic navigation trends.