A curated, continuously-updated reading list of World Action Models (WAM), Vision-Language-Action (VLA) models, and Embodied AI — organized by a survey-grounded taxonomy.
🌐 Website: hyperboliccurve.github.io/Awesome-World-Action-Model
The push toward general-purpose robots has produced two converging families of foundation models:
- Vision-Language-Action (VLA) models inherit the language grounding and visual understanding of pretrained Vision-Language Models (VLMs) and adapt them to emit actions — a scalable route to language-conditioned policies.
- World Action Models (WAM) start from a world model / video backbone that predicts how a scene evolves, and adapt that predictive prior to emit actions — trading the "language→motion" grounding gap for a "dynamics→action" one.
These two families overlap: a WAM built on a pretrained VLM is simultaneously a VLA and a WAM. This list maps that landscape with a taxonomy grounded in the recent survey literature (see Surveys), so each category has a clear, defensible scope rather than an ad-hoc label.
Note
Legend — 📄 arXiv · 🌐 project page · 💻 code · 📊 dataset/benchmark. Tables are sorted newest-first within each category. The 🆕 Latest Papers section is refreshed daily from arXiv by a GitHub Action; everything else is hand-curated.
flowchart TD
A[Robot Foundation Models] --> B[Vision-Language-Action<br/>VLA]
A --> C[World & World-Action Models<br/>WM / WAM]
A --> R[Action Representations]
A --> P[Foundational Policies]
B --> B1[By Action Representation:<br/>Autoregressive · Diffusion · Flow-Matching]
B --> B2[By Capability:<br/>Reasoning/Dual-System · 3D-4D · Efficient · RL Fine-Tuning]
C --> C1[Foundation / General World Models]
C --> C2[WAM from Video Generation]
C --> C3[WAM from VLMs]
C --> C4[WAM from Scratch · Latent / JEPA]
C --> C5[Domain: Driving · Navigation]
R --> R1[Discrete / Autoregressive Tokenizers]
R --> R2[Diffusion & Flow-Matching Policies]
- 🔑 Key Definitions
- 🆕 Latest Papers (Auto-updated)
- 📚 Surveys
- 🤖 Vision-Language-Action (VLA) Models
- 🌎 World & World-Action Models
- 🧩 Action Representations & Tokenization
- 🦾 Foundational Robot Policies
- 📦 Resources
- 🗂️ Extended Paper Index (Auto-Curated)
- 📋 Full Paper Index & Baselines
- 🤝 Contributing
| Term | Definition | Canonical reference |
|---|---|---|
| Vision-Language-Action (VLA) | A robot policy that adapts a pretrained VLM to map images + language instructions to actions. | RT-2 (Brohan et al., 2023) |
| World Model (WM) | A learned model that predicts future states of an environment (in pixels, latents, or 3D/4D), used for planning, simulation, or representation. | World Models (Ha & Schmidhuber, 2018) |
| World Action Model (WAM) | A policy that leverages world-modeling capability (predicting future states) for action prediction — typically by adapting a video / world-model backbone to emit actions. | GR-1 (Wu et al., 2023) |
Important
VLA ∩ WAM. The families intersect: a WAM built on a pretrained VLM is both. The split in this list is by what prior the model starts from — VLM-style vision-language priors (VLA) vs. video/dynamics priors (WAM) — and, within VLA, by how actions are represented, the axis most surveys agree is the field's clearest discriminator.
Papers are automatically fetched daily from arXiv. Last updated: 2026-07-21
| Paper | Date | Code |
|---|---|---|
| BadWAM: When World-Action Models Dream Right but Act Wrong Qi Li, Xingyi Yang et al. |
2026-07-16 | |
| AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight Xinhong Zhang, Qiyuan Zhu et al. |
2026-07-16 | |
| WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Yusen Feng, Bingchen Han et al. |
2026-07-08 | |
| HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models Angen Ye, Weijie Ke et al. |
2026-07-05 | |
| Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors Zixing Wang, Kausik Sivakumar et al. |
2026-06-30 |
Recent surveys that define the field and motivate the taxonomy used here.
| Title | Authors | Year | Links |
|---|---|---|---|
| Understanding World or Predicting Future? A Comprehensive Survey of World Models | Ding et al. | 2024 | 📄 |
| A Comprehensive Survey on World Models for Embodied AI | Li et al. | 2025 | 📄 |
| 3D and 4D World Modeling: A Survey | Kong et al. | 2025 | 📄 |
| Learning Embodied Intelligence from Physical Simulators and World Models | Long et al. | 2025 | 📄 |
| Embodied AI: From LLMs to World Models | Feng et al. | 2025 | 📄 |
| World Model for Robot Learning: A Comprehensive Survey | Hou et al. | 2026 | 📄 |
| Modeling the Mental World for Embodied AI: A Comprehensive Review | Liu et al. | 2026 | 📄 |
| The Role of World Models in Shaping Autonomous Driving: A Survey | Tu et al. | 2025 | 📄 |
| Title | Authors | Year | Links |
|---|---|---|---|
| A Survey on Vision-Language-Action Models for Embodied AI | Ma et al. | 2024 | 📄 |
| A Survey on VLA Models: An Action Tokenization Perspective | Zhong et al. | 2025 | 📄 |
| VLA Models: Concepts, Progress, Applications and Challenges | Sapkota et al. | 2025 | 📄 |
| Large VLM-based VLA Models for Robotic Manipulation: A Survey | Shao et al. | 2025 | 📄 |
| Efficient VLA Models for Embodied Manipulation: A Systematic Survey | Guan et al. | 2025 | 📄 |
| VLA Models for Robotics: A Review Towards Real-World Applications | Kawaharazuka et al. | 2025 | 📄 · 🌐 |
| An Anatomy of VLA Models: From Modules to Milestones and Challenges | — | 2025 | 📄 |
| Pure Vision-Language-Action Models: A Comprehensive Survey | — | 2025 | 📄 |
| A Survey on Efficient Vision-Language-Action Models | Yu et al. | 2025 | 📄 |
| A Survey on VLA Models for Autonomous Driving | Jiang et al. | 2025 | 📄 |
| VLA in Robotics: A Survey of Datasets, Benchmarks, and Data Engines | Wang et al. | 2026 | 📄 |
| Title | Authors | Year | Links |
|---|---|---|---|
| Foundation Models in Robotics: Applications, Challenges, and the Future | Firoozi et al. | 2023 | 📄 |
| Toward General-Purpose Robots via Foundation Models: A Survey | Hu et al. | 2023 | 📄 |
| Aligning Cyber Space with Physical World: A Survey on Embodied AI | Liu et al. | 2024 | 📄 |
| What Foundation Models Can Bring for Robot Learning in Manipulation: A Survey | Li et al. | 2024 | 📄 |
| Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes | Tang et al. | 2024 | 📄 |
| Generative AI in Robotic Manipulation: A Survey | Zhang et al. | 2025 | 📄 |
| A Survey of Sim-to-Real Methods in RL with Foundation Models | Da et al. | 2025 | 📄 |
| Behavior Foundation Model: Towards Next-Generation Whole-Body Control of Humanoids | Yuan et al. | 2025 | 📄 |
| Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey | Bai et al. | 2025 | 📄 |
| Robotic Foundation Models for Industrial Control: A Survey & Readiness Assessment | Kube et al. | 2026 | 📄 |
Following the action-tokenization view (Zhong et al., 2025), the primary split is by how actions are represented; capability-oriented subsections (reasoning, 3D/4D, efficiency, RL) cut across it. A few pre-/non-VLM generalist policies (e.g., RT-1, Octo) are listed alongside their successors to show lineage — see Foundational Robot Policies for the strictly non-VLA baselines.
Actions are binned into discrete tokens and decoded like text. Simple and VLM-native; high-frequency dexterity needs better tokenizers (see FAST).
| Model | Title | Year | Links |
|---|---|---|---|
| VLA-0 | Building SOTA VLAs with Zero Modification | 2025 | 📄 · 🌐 |
| UniVLA | Unified Vision-Language-Action Model (native multimodal tokens) | 2025 | 📄 · 🌐 |
| π0-FAST | Autoregressive π0 variant using the FAST action tokenizer | 2025 | 📄 · 🌐 |
| OpenVLA | An Open-Source Vision-Language-Action Model | 2024 | 📄 · 🌐 · 💻 |
| RT-2 | VLA Models Transfer Web Knowledge to Robotic Control | 2023 | 📄 · 🌐 |
| RT-1 | Robotics Transformer for Real-World Control at Scale | 2022 | 📄 · 🌐 · 💻 |
A diffusion action head denoises continuous action chunks conditioned on vision-language features.
| Model | Title | Year | Links |
|---|---|---|---|
| RoboVLMs | Towards Generalist Robot Policies: What Matters in Building VLAs | 2024 | 📄 · 🌐 |
| CogACT | A Foundational VLA Model for Synergizing Cognition and Action | 2024 | 📄 |
| TinyVLA | Fast, Data-Efficient VLA Models for Manipulation | 2024 | 📄 · 🌐 |
| Octo | An Open-Source Generalist Robot Policy | 2024 | 📄 · 🌐 · 💻 |
A conditional flow/vector field transports noise to action chunks — the dominant head for current SOTA generalist VLAs.
| Model | Title | Year | Links |
|---|---|---|---|
| π*0.6 | A VLA That Learns From Experience | 2025 | 📄 · 🌐 |
| X-VLA | Soft-Prompted Transformer as a Scalable Cross-Embodiment VLA | 2025 | 📄 · 🌐 · 💻 |
| SmolVLA | A VLA for Affordable and Efficient Robotics | 2025 | 📄 · 💻 |
| π0.5 | A VLA with Open-World Generalization | 2025 | 📄 · 🌐 |
| Gemini Robotics | Bringing AI into the Physical World | 2025 | 📄 · 🌐 |
| GR00T N1 | An Open Foundation Model for Generalist Humanoid Robots | 2025 | 📄 · 💻 |
| EO-1 | An Open Unified Embodied Foundation Model (interleaved reasoning + acting) | 2025 | 📄 · 🌐 |
| GR-3 | Large-Scale Vision-Language-Action Model (Technical Report) | 2025 | 📄 |
| FLOWER | Democratizing Generalist Robot Policies with Efficient VLA Flow Policies | 2025 | 📄 |
| π0 | A Vision-Language-Action Flow Model for General Robot Control | 2024 | 📄 · 🌐 |
Explicit chain-of-thought / embodied reasoning, or a slow System-2 planner paired with a fast System-1 controller.
| Model | Title | Year | Links |
|---|---|---|---|
| ACoT-VLA | Action Chain-of-Thought for VLA Models | 2026 | 📄 · 💻 |
| Gemini Robotics 1.5 | Embodied Reasoning & Motion Transfer | 2025 | 📄 |
| ThinkAct | VLA Reasoning via Reinforced Visual Latent Planning | 2025 | 📄 |
| OpenHelix | A Short Survey & Open-Source Dual-System VLA | 2025 | 📄 |
| FiS-VLA | Fast-in-Slow: A Dual-System Foundation Model for Unified Fast–Slow Reasoning | 2025 | 📄 |
| WALL-OSS | Igniting VLMs toward the Embodied Space | 2025 | 📄 · 💻 |
| CoT-VLA | Visual Chain-of-Thought Reasoning for VLA | 2025 | 📄 |
Policies that reason over explicit 3D/4D structure (point clouds, occupancy, predicted future frames) rather than 2D images alone. (VoxPoser, a zero-shot 3D value-map planner, lives under Foundational Robot Policies.)
| Model | Title | Year | Links |
|---|---|---|---|
| 3D-VLA | A 3D Vision-Language-Action Generative World Model | 2024 | 📄 |
Compression, caching, parallel decoding, and distillation to make VLAs small and fast enough for real-time / edge control (Guan et al., 2025).
| Model | Title | Year | Links |
|---|---|---|---|
| FASTER | Rethinking Real-Time Flow VLAs | 2026 | 📄 |
| RTC | Real-Time Chunking: Running VLAs at Real-Time Speed | 2025 | 📄 |
| NanoVLA | Routing-Decoupled VLA for Nano-Sized Generalist Policies | 2025 | 📄 |
| VLA-Adapter | A Tiny-Scale VLA Paradigm | 2025 | 📄 |
| OpenVLA-OFT | Fine-Tuning VLAs: Optimizing Speed and Success | 2025 | 📄 · 🌐 |
| TinyVLA | Fast, Data-Efficient VLA Models | 2024 | 📄 · 🌐 |
Reinforcement learning (often on top of flow-/diffusion-based VLAs) to improve over imitation-only training.
| Model | Title | Year | Links |
|---|---|---|---|
| π_RL | Online RL Fine-Tuning for Flow-based VLAs | 2025 | 📄 |
| VLA-RFT | RL Fine-Tuning with Verified Rewards in World Simulators | 2025 | 📄 |
| SimpleVLA-RL | Scaling VLA Training via Reinforcement Learning | 2025 | 📄 |
| ConRFT | A Reinforced Fine-Tuning Method for VLA via Consistency Policy | 2025 | 📄 |
Organized by what the model predicts and how it is built, following the embodied-world-model taxonomy of Li et al., 2025 and the WAM split popularized by awesome-vla-wam.
General-purpose models of environment dynamics — spanning classical latent world models for model-based RL (World Models, DreamerV3) and modern large-scale video / foundation world models — used for planning, neural simulation, or as backbones for WAMs.
| Model | Title | Year | Links |
|---|---|---|---|
| Cosmos-Predict2.5 | World Simulation with Video Foundation Models for Physical AI | 2025 | 📄 · 💻 |
| Cosmos-Reason1 | From Physical Common Sense to Embodied Reasoning | 2025 | 📄 |
| Cosmos | World Foundation Model Platform for Physical AI | 2025 | 📄 · 🌐 |
| V-JEPA 2 | Self-Supervised Video Models Enable Understanding, Prediction & Planning | 2025 | 📄 |
| iVideoGPT | Interactive VideoGPTs are Scalable World Models | 2024 | 📄 |
| Genie | Generative Interactive Environments | 2024 | 📄 |
| DreamerV3 | Mastering Diverse Domains through World Models | 2023 | 📄 · 💻 |
| UniSim | Learning Interactive Real-World Simulators | 2023 | 📄 |
| World Models | Recurrent latent world model + controller (origin of the term) | 2018 | 📄 |
A (text-/image-conditioned) video generator imagines future frames; actions are recovered via an inverse-dynamics / action head.
| Model | Title | Year | Links |
|---|---|---|---|
| DreamZero | World Action Models are Zero-shot Policies | 2026 | 📄 · 🌐 |
| DiT4DiT | Jointly Modeling Video Dynamics and Actions | 2026 | 📄 |
| Cosmos Policy | Fine-Tuning Video Models for Visuomotor Control & Planning | 2026 | 📄 · 🌐 |
| Video2Act | A Dual-System Video Diffusion Policy | 2025 | 📄 |
| GR-2 | A Generative Video-Language-Action Model with Web-Scale Knowledge | 2024 | 📄 |
| GR-1 | Large-Scale Video Generative Pre-training for Visual Robot Manipulation | 2023 | 📄 |
A pretrained VLM is turned into a world model (e.g., predicting goal images / object-centric futures) that then drives action.
| Model | Title | Year | Links |
|---|---|---|---|
| DreamVLA | A VLA Model Dreamed with Comprehensive World Knowledge | 2025 | 📄 |
| Goal-VLA | Image-Generative VLMs as Object-Centric World Models for VLA | 2025 | 📄 |
Single architectures that jointly learn to act and to predict world dynamics, blurring the VLA/WAM boundary.
| Model | Title | Year | Links |
|---|---|---|---|
| RynnVLA-002 | A Unified Vision-Language-Action and World Model | 2025 | 📄 · 💻 |
| WholeBodyVLA | Unified Latent VLA for Whole-Body Loco-Manipulation | 2025 | 📄 · 💻 |
| WorldVLA | Towards an Autoregressive Action World Model | 2025 | 📄 · 💻 |
Self-supervised latent predictive models (non-reconstructive joint-embedding / JEPA). The JEPA foundations (I-JEPA) learn to predict in representation space; the action-conditioned variant (V-JEPA 2-AC) turns that prior into a world model for planning.
| Model | Title | Year | Links |
|---|---|---|---|
| V-JEPA 2-AC | Action-Conditioned Latent World Model for Zero-Shot Planning | 2025 | 📄 |
| I-JEPA | Image-based Joint-Embedding Predictive Architecture (representation foundation) | 2023 | 📄 |
| Model | Title | Year | Links |
|---|---|---|---|
| GAIA-2 | A Controllable Multi-View Generative World Model for Autonomous Driving | 2025 | 📄 |
| Navigation World Models | Conditional Diffusion Transformer for Navigation | 2024 | 📄 |
| GAIA-1 | A Generative World Model for Autonomous Driving | 2023 | 📄 |
Building blocks shared across VLA and WAM policies — how continuous actions become learnable targets.
| Method | Title | Year | Links |
|---|---|---|---|
| FAST | Efficient (DCT-based) Action Tokenization for VLAs | 2025 | 📄 · 🌐 |
| BeT | Behavior Transformers: Cloning k Modes with One Stone | 2022 | 📄 |
Heads that emit continuous action chunks — by denoising diffusion (Diffusion Policy) or by chunked sequence prediction with a CVAE (ACT). Flow-matching heads (π0, SmolVLA, …) are listed with their models under Flow-Matching VLA.
| Method | Title | Year | Links |
|---|---|---|---|
| Diffusion Policy | Visuomotor Policy Learning via Action Diffusion | 2023 | 📄 · 🌐 |
| ACT / ALOHA | Action Chunking with Transformers | 2023 | 📄 · 🌐 |
Non-VLA policies and planners that remain standard baselines in the experimental tables of the papers above. (Diffusion Policy, ACT, and BeT are described under Action Representations.)
| Method | Title | Year | Links |
|---|---|---|---|
| CrossFormer | Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion & Flight | 2024 | 📄 · 💻 |
| RoboFlamingo | Vision-Language Foundation Models as Effective Robot Imitators | 2023 | 📄 · 💻 |
| VoxPoser | Composable 3D Value Maps for Robotic Manipulation (zero-shot LLM + 3D planner) | 2023 | 📄 · 🌐 |
| RT-1 | Robotics Transformer for Real-World Control at Scale | 2022 | 📄 |
| Name | Description | Scale | Links |
|---|---|---|---|
| Open X-Embodiment | Cross-embodiment aggregation behind the RT-X models | 1M+ traj · 22 embodiments | 📄 · 🌐 |
| AgiBot World | Large-scale real-world manipulation (Colosseo) | 1M+ traj · 217 tasks | 📄 · 🌐 |
| EgoScale | Scaling dexterous manipulation with diverse egocentric human data | egocentric · 2026 | 📄 |
| DexCanvas | Human demos ↔ robot learning for dexterous manipulation | dexterous | 📄 |
| Galaxea Open-World | Mobile-bimanual dataset paired with the G0 dual-system VLA | 500 hrs · 150 tasks | 📄 · 💻 |
| DROID | In-the-wild Franka manipulation across 3 continents | 76K traj · 564 scenes | 📄 · 🌐 |
| RoboMIND | Multi-embodiment teleop incl. labeled failures | 107K traj · 479 tasks | 📄 · 🌐 |
| BridgeData V2 | WidowX manipulation w/ language + goal images | 60K traj · 24 envs | 📄 · 🌐 |
| RH20T | Contact-rich skills w/ paired human demos | 110K+ seq · 147 tasks | 📄 · 🌐 |
| Ego-Exo4D | Simultaneous ego + exo video of skilled activity | 1,286 hrs | 📄 · 🌐 |
| Ego4D | Massive egocentric daily-life video | 3,670 hrs | 📄 · 🌐 |
| Name | Description | Links |
|---|---|---|
| LIBERO | Lifelong robot-learning, 130 manipulation tasks (de-facto VLA eval) | 📄 · 💻 |
| CALVIN | Long-horizon language-conditioned manipulation | 📄 · 💻 |
| SimplerEnv | Real-to-sim evaluation for manipulation policies | 📄 · 🌐 |
| RoboCasa | Large-scale kitchen simulation (100 tasks) | 📄 · 🌐 |
| VLABench | World-knowledge & long-horizon language tasks | 📄 · 🌐 |
| ManiSkill3 | GPU-parallel manipulation (30K+ FPS) | 📄 · 🌐 |
| THE COLOSSEUM | Robustness under 14 environmental perturbations | 📄 · 🌐 |
| RoboArena | Distributed crowd-sourced real-world policy eval | 📄 · 🌐 |
| RoboChallenge | Large-scale real-robot evaluation of embodied policies | 📄 |
| RobotArena ∞ | Scalable robot benchmarking via real-to-sim translation | 📄 · 🌐 |
| WorldArena | Perception & functional-utility benchmark for embodied world models | 📄 |
| Meta-World | 50 tabletop tasks for multi-task / meta-RL | 📄 · 💻 |
| RLBench | 100 hand-designed manipulation tasks | 📄 · 💻 |
| Name | Description | Links |
|---|---|---|
| Isaac Sim / Isaac Lab | GPU-native robotics sim + RL/IL framework (Omniverse/USD) | 🌐 |
| MuJoCo / MJX | Standard rigid-body engine + JAX/XLA parallel variant | 🌐 |
| Genesis | Generative, multi-solver physics platform (up to ~43M FPS) | 🌐 |
| ManiSkill | GPU-parallel manipulation simulator on SAPIEN | 🌐 |
| SAPIEN | Part-level articulated-object simulator (PartNet-Mobility) | 📄 · 🌐 |
| Habitat | Photorealistic indoor navigation & rearrangement | 🌐 |
| ThreeDWorld | Multimodal Unity3D sim (vision + audio + physics) | 📄 · 🌐 |
| Newton | Open, differentiable GPU physics engine (NVIDIA + DeepMind + Disney) | 🌐 |
| Name | Description | Links |
|---|---|---|
| LeRobot | End-to-end PyTorch robot-learning library + datasets + low-cost HW | 📄 · 💻 |
| openpi | Open models & training/inference for π0, π0-FAST, π0.5 | 💻 |
| Isaac GR00T | Open humanoid foundation-model framework + checkpoints | 💻 |
| OpenVLA | Training / LoRA fine-tuning for the 7B OpenVLA model | 💻 |
| Octo | JAX/Flax generalist transformer policy on OXE | 💻 |
| robomimic / robosuite | Learning-from-demonstration framework + MuJoCo manipulation sim | 💻 |
| HIL-SERL | Human-in-the-loop, sample-efficient real-world RL | 💻 |
A broader, continuously-mined index of recent arXiv work that complements the curated highlights above — 183 additional papers, newest first. Last updated: 2026-07-16. Auto-generated from
data/*.jsonbyscripts/expand_papers.py; papers already highlighted above are omitted here to avoid duplication.
VLA — General & Manipulation · 45 papers
VLA — Reasoning, Planning & Dual-System · 6 papers
| Paper | Authors | Date | Links |
|---|---|---|---|
| ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception | Weichen Zhang, Shiquan Yu et al. | 2026-07-11 | |
| Do World Action Models Generalize Better than VLAs? A Robustness Study | Zhanguang Zhang, Zhiyuan Li et al. | 2026-03-23 | |
| Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models | Riccardo Andrea Izzo, Gianluca Bardaro et al. | 2026-03-05 | |
| Chain of World: World Model Thinking in Latent Motion | Fuxiang Yang, Donglin Di et al. | 2026-03-03 | |
| FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment | Han Zhao, Jingbo Wang et al. | 2026-02-19 | |
| VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation | Changhua Xu, Jie Lu et al. | 2026-02-07 |
VLA — Autonomous Driving · 12 papers
VLA — Dexterous & Humanoid · 1 papers
| Paper | Authors | Date | Links |
|---|---|---|---|
| Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models | Ruixing Jin, Zicheng Zhu et al. | 2026-03-24 |
VLA — 3D / 4D & Spatial · 6 papers
| Paper | Authors | Date | Links |
|---|---|---|---|
| CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking | Ruilong Ren, Songsheng Cheng et al. | 2026-07-16 | |
| VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation | Mohan Liu, Zhihao Gu et al. | 2026-07-14 | |
| See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models | Byungkun Lee, Dongyoon Hwang et al. | 2026-07-13 | |
| TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging | Shengzhuo Yang, Ronghao Yu et al. | 2026-07-10 | |
| LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior | Xinkai Wang, Chenyi Wang et al. | 2026-03-26 | |
| 3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models | Bin Yu, Shijie Lian et al. | 2026-03-25 |
VLA — RL & Post-Training · 6 papers
| Paper | Authors | Date | Links |
|---|---|---|---|
| ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning | Yilun Kong, Yunpeng Qing et al. | 2026-07-14 | |
| VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation | Zhide Zhong, Haodong Yan et al. | 2026-03-27 | |
| On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning | Changyu Liu, Yiyang Liu et al. | 2026-01-11 | |
| VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning | Si-Cheng Wang, Tian-Yu Xiang et al. | 2025-09-30 | |
| The arc-shaped radio source at the center of NGC 6334A: Is it a colliding wind region of two young massive stars or the bow shock of a runaway star? | Vanessa Yanza, Sergio A. Dzib et al. | 2025-02-24 | |
| VLA 22 GHz Imaging of Massive Star Formation in Local Wolf-Rayet Galaxies | Nicholas G. Ferraro, Jean L. Turner et al. | 2024-11-09 |
VLA — Efficient & Real-Time · 19 papers
VLA — Safety, Robustness & Evaluation · 13 papers
World Models — General & Foundation · 12 papers
World Models — Video Generation & WAM · 8 papers
World Models — Driving & Navigation · 4 papers
| Paper | Authors | Date | Links |
|---|---|---|---|
| Enhancing Policy Learning with World-Action Model | Yuci Han, Alper Yilmaz | 2026-03-30 | |
| Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving | Linbo Wang, Yupeng Zheng et al. | 2026-03-25 | |
| NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation | Tianshuai Hu, Zeying Gong et al. | 2026-03-16 | |
| AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation | Ge Yuan, Qiyuan Qiao et al. | 2026-02-23 |
Policies — Diffusion & Flow · 27 papers
Policies — Imitation & Behavior Learning · 10 papers
Policies — Robot Learning & Manipulation · 14 papers
📊 Click to expand the complete paper list and baseline methods
The curated tables above highlight landmark and representative work. For the exhaustive, auto-maintained index and the baseline methods extracted from experimental tables, see:
- 📋 Complete Paper List — full index, sorted newest-first
- 📊 Baseline Methods — comparison methods from major VLA/WAM papers
| Family | Key baselines |
|---|---|
| VLA | RT-1, RT-2, OpenVLA, Octo, π0, π0.5, X-VLA, UniVLA, SmolVLA |
| Policy | Diffusion Policy, ACT, BeT, RoboFlamingo, CrossFormer |
| World Model | DreamerV3, I-JEPA, V-JEPA 2, Genie, Cosmos, GR-1/GR-2 |
Contributions are very welcome! To add or fix a paper:
- Add a paper — open a PR placing it in the appropriate category (keep tables sorted newest-first), or open an issue with the arXiv link.
- Fix an error — submit a PR with the correction.
- New papers appear automatically — the 🆕 Latest Papers section and the 🗂️ Extended Paper Index are regenerated daily by the scraper; do not hand-edit content between the auto markers.
To run the discovery pipeline locally:
pip install -r requirements.txt
python scripts/arxiv_scraper.py --max-results 50 --days-back 30 # writes data/papers.json
python scripts/update_readme.py # refreshes the 🆕 auto section
python scripts/expand_papers.py # refreshes the 🗂️ extended index
python scripts/build_site.py # rebuilds the GitHub Pages site (index.html)Released under the MIT License.
Inspired by awesome-vla-wam, awesome-physical-ai, and awesome-vla-study. Taxonomy grounded in the surveys listed above.