Skip to content

HyperbolicCurve/Awesome-World-Action-Model

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

131 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🤖 Awesome World Action Models Awesome

A curated, continuously-updated reading list of World Action Models (WAM), Vision-Language-Action (VLA) models, and Embodied AI — organized by a survey-grounded taxonomy.

Last Update PRs Welcome License: MIT Stars

🌐 Website: hyperboliccurve.github.io/Awesome-World-Action-Model


Overview

The push toward general-purpose robots has produced two converging families of foundation models:

  • Vision-Language-Action (VLA) models inherit the language grounding and visual understanding of pretrained Vision-Language Models (VLMs) and adapt them to emit actions — a scalable route to language-conditioned policies.
  • World Action Models (WAM) start from a world model / video backbone that predicts how a scene evolves, and adapt that predictive prior to emit actions — trading the "language→motion" grounding gap for a "dynamics→action" one.

These two families overlap: a WAM built on a pretrained VLM is simultaneously a VLA and a WAM. This list maps that landscape with a taxonomy grounded in the recent survey literature (see Surveys), so each category has a clear, defensible scope rather than an ad-hoc label.

Note

Legend — 📄 arXiv · 🌐 project page · 💻 code · 📊 dataset/benchmark. Tables are sorted newest-first within each category. The 🆕 Latest Papers section is refreshed daily from arXiv by a GitHub Action; everything else is hand-curated.


Taxonomy at a Glance

flowchart TD
    A[Robot Foundation Models] --> B[Vision-Language-Action<br/>VLA]
    A --> C[World &amp; World-Action Models<br/>WM / WAM]
    A --> R[Action Representations]
    A --> P[Foundational Policies]

    B --> B1[By Action Representation:<br/>Autoregressive · Diffusion · Flow-Matching]
    B --> B2[By Capability:<br/>Reasoning/Dual-System · 3D-4D · Efficient · RL Fine-Tuning]

    C --> C1[Foundation / General World Models]
    C --> C2[WAM from Video Generation]
    C --> C3[WAM from VLMs]
    C --> C4[WAM from Scratch · Latent / JEPA]
    C --> C5[Domain: Driving · Navigation]

    R --> R1[Discrete / Autoregressive Tokenizers]
    R --> R2[Diffusion &amp; Flow-Matching Policies]
Loading

Table of Contents


🔑 Key Definitions

Term Definition Canonical reference
Vision-Language-Action (VLA) A robot policy that adapts a pretrained VLM to map images + language instructions to actions. RT-2 (Brohan et al., 2023)
World Model (WM) A learned model that predicts future states of an environment (in pixels, latents, or 3D/4D), used for planning, simulation, or representation. World Models (Ha & Schmidhuber, 2018)
World Action Model (WAM) A policy that leverages world-modeling capability (predicting future states) for action prediction — typically by adapting a video / world-model backbone to emit actions. GR-1 (Wu et al., 2023)

Important

VLA ∩ WAM. The families intersect: a WAM built on a pretrained VLM is both. The split in this list is by what prior the model starts from — VLM-style vision-language priors (VLA) vs. video/dynamics priors (WAM) — and, within VLA, by how actions are represented, the axis most surveys agree is the field's clearest discriminator.


🆕 Latest Papers (Auto-updated)

Papers are automatically fetched daily from arXiv. Last updated: 2026-07-21

VLA

Paper Date Code
FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
Ruicheng Li, Qixiu Li et al.
2026-07-20
Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation
Peng Ren, Haoyang Ge et al.
2026-07-20
Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models
Tuan Duong Trinh, Naveed Akhtar et al.
2026-07-20
VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking with Highly Similar Objects
Yanrong Qin, Xiaoyan Cao et al.
2026-07-19
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
Kalpana Panda, Wesley Maia et al.
2026-07-18
Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models
Yuhan Liu, Xinyu Zhang et al.
2026-07-17
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
Haoran Sun, Wentao Zhang et al.
2026-07-17
AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning
Xiaojiang Peng, Kai Peng et al.
2026-07-17
Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving
Yun Li, Jiachen Gong et al.
2026-07-17
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo et al.
2026-07-16

World Model

Paper Date Code
WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory
Haisheng Su, Zongdai Liu et al.
2026-07-21
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Ziqin Wang, Hao Li et al.
2026-07-21
GeoWorldAD: Geometry World Action Model for Autonomous Driving
Songyan Zhang, Jinyuan Tian et al.
2026-07-20
Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation
Zesen Zhao, Minkyoung Cho et al.
2026-07-20
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
Jihoon Hong, Julian Skifstad et al.
2026-07-16
FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
Yixiang Chen, Peiyan Li et al.
2026-07-14
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Yuanzhi Liang, Xufeng Zhan et al.
2026-07-13
EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
Baoyu Li, Xinchen Yin et al.
2026-07-08
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
Haoyu Zhao, Xingyue Zhao et al.
2026-07-07
Learning 4D Geometric Priors for Inference-Efficient World Action Models
Jianjun Zhang, Jian Zhu et al.
2026-07-06

Policy

Paper Date Code
BadWAM: When World-Action Models Dream Right but Act Wrong
Qi Li, Xingyi Yang et al.
2026-07-16
AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight
Xinhong Zhang, Qiyuan Zhu et al.
2026-07-16
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Yusen Feng, Bingchen Han et al.
2026-07-08
HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
Angen Ye, Weijie Ke et al.
2026-07-05
Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors
Zixing Wang, Kausik Sivakumar et al.
2026-06-30

📚 Surveys

Recent surveys that define the field and motivate the taxonomy used here.

World & Embodied World Models

Title Authors Year Links
Understanding World or Predicting Future? A Comprehensive Survey of World Models Ding et al. 2024 📄
A Comprehensive Survey on World Models for Embodied AI Li et al. 2025 📄
3D and 4D World Modeling: A Survey Kong et al. 2025 📄
Learning Embodied Intelligence from Physical Simulators and World Models Long et al. 2025 📄
Embodied AI: From LLMs to World Models Feng et al. 2025 📄
World Model for Robot Learning: A Comprehensive Survey Hou et al. 2026 📄
Modeling the Mental World for Embodied AI: A Comprehensive Review Liu et al. 2026 📄
The Role of World Models in Shaping Autonomous Driving: A Survey Tu et al. 2025 📄

Vision-Language-Action

Title Authors Year Links
A Survey on Vision-Language-Action Models for Embodied AI Ma et al. 2024 📄
A Survey on VLA Models: An Action Tokenization Perspective Zhong et al. 2025 📄
VLA Models: Concepts, Progress, Applications and Challenges Sapkota et al. 2025 📄
Large VLM-based VLA Models for Robotic Manipulation: A Survey Shao et al. 2025 📄
Efficient VLA Models for Embodied Manipulation: A Systematic Survey Guan et al. 2025 📄
VLA Models for Robotics: A Review Towards Real-World Applications Kawaharazuka et al. 2025 📄 · 🌐
An Anatomy of VLA Models: From Modules to Milestones and Challenges 2025 📄
Pure Vision-Language-Action Models: A Comprehensive Survey 2025 📄
A Survey on Efficient Vision-Language-Action Models Yu et al. 2025 📄
A Survey on VLA Models for Autonomous Driving Jiang et al. 2025 📄
VLA in Robotics: A Survey of Datasets, Benchmarks, and Data Engines Wang et al. 2026 📄

Foundation Models & Embodied AI

Title Authors Year Links
Foundation Models in Robotics: Applications, Challenges, and the Future Firoozi et al. 2023 📄
Toward General-Purpose Robots via Foundation Models: A Survey Hu et al. 2023 📄
Aligning Cyber Space with Physical World: A Survey on Embodied AI Liu et al. 2024 📄
What Foundation Models Can Bring for Robot Learning in Manipulation: A Survey Li et al. 2024 📄
Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes Tang et al. 2024 📄
Generative AI in Robotic Manipulation: A Survey Zhang et al. 2025 📄
A Survey of Sim-to-Real Methods in RL with Foundation Models Da et al. 2025 📄
Behavior Foundation Model: Towards Next-Generation Whole-Body Control of Humanoids Yuan et al. 2025 📄
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey Bai et al. 2025 📄
Robotic Foundation Models for Industrial Control: A Survey & Readiness Assessment Kube et al. 2026 📄

🤖 Vision-Language-Action (VLA) Models

Following the action-tokenization view (Zhong et al., 2025), the primary split is by how actions are represented; capability-oriented subsections (reasoning, 3D/4D, efficiency, RL) cut across it. A few pre-/non-VLM generalist policies (e.g., RT-1, Octo) are listed alongside their successors to show lineage — see Foundational Robot Policies for the strictly non-VLA baselines.

By Action Representation

Autoregressive / Discrete-Token VLA

Actions are binned into discrete tokens and decoded like text. Simple and VLM-native; high-frequency dexterity needs better tokenizers (see FAST).

Model Title Year Links
VLA-0 Building SOTA VLAs with Zero Modification 2025 📄 · 🌐
UniVLA Unified Vision-Language-Action Model (native multimodal tokens) 2025 📄 · 🌐
π0-FAST Autoregressive π0 variant using the FAST action tokenizer 2025 📄 · 🌐
OpenVLA An Open-Source Vision-Language-Action Model 2024 📄 · 🌐 · 💻
RT-2 VLA Models Transfer Web Knowledge to Robotic Control 2023 📄 · 🌐
RT-1 Robotics Transformer for Real-World Control at Scale 2022 📄 · 🌐 · 💻

Diffusion-based VLA

A diffusion action head denoises continuous action chunks conditioned on vision-language features.

Model Title Year Links
RoboVLMs Towards Generalist Robot Policies: What Matters in Building VLAs 2024 📄 · 🌐
CogACT A Foundational VLA Model for Synergizing Cognition and Action 2024 📄
TinyVLA Fast, Data-Efficient VLA Models for Manipulation 2024 📄 · 🌐
Octo An Open-Source Generalist Robot Policy 2024 📄 · 🌐 · 💻

Flow-Matching VLA

A conditional flow/vector field transports noise to action chunks — the dominant head for current SOTA generalist VLAs.

Model Title Year Links
π*0.6 A VLA That Learns From Experience 2025 📄 · 🌐
X-VLA Soft-Prompted Transformer as a Scalable Cross-Embodiment VLA 2025 📄 · 🌐 · 💻
SmolVLA A VLA for Affordable and Efficient Robotics 2025 📄 · 💻
π0.5 A VLA with Open-World Generalization 2025 📄 · 🌐
Gemini Robotics Bringing AI into the Physical World 2025 📄 · 🌐
GR00T N1 An Open Foundation Model for Generalist Humanoid Robots 2025 📄 · 💻
EO-1 An Open Unified Embodied Foundation Model (interleaved reasoning + acting) 2025 📄 · 🌐
GR-3 Large-Scale Vision-Language-Action Model (Technical Report) 2025 📄
FLOWER Democratizing Generalist Robot Policies with Efficient VLA Flow Policies 2025 📄
π0 A Vision-Language-Action Flow Model for General Robot Control 2024 📄 · 🌐

By Capability

Reasoning & Dual-System (Fast–Slow) VLA

Explicit chain-of-thought / embodied reasoning, or a slow System-2 planner paired with a fast System-1 controller.

Model Title Year Links
ACoT-VLA Action Chain-of-Thought for VLA Models 2026 📄 · 💻
Gemini Robotics 1.5 Embodied Reasoning & Motion Transfer 2025 📄
ThinkAct VLA Reasoning via Reinforced Visual Latent Planning 2025 📄
OpenHelix A Short Survey & Open-Source Dual-System VLA 2025 📄
FiS-VLA Fast-in-Slow: A Dual-System Foundation Model for Unified Fast–Slow Reasoning 2025 📄
WALL-OSS Igniting VLMs toward the Embodied Space 2025 📄 · 💻
CoT-VLA Visual Chain-of-Thought Reasoning for VLA 2025 📄

3D / 4D-Aware VLA

Policies that reason over explicit 3D/4D structure (point clouds, occupancy, predicted future frames) rather than 2D images alone. (VoxPoser, a zero-shot 3D value-map planner, lives under Foundational Robot Policies.)

Model Title Year Links
3D-VLA A 3D Vision-Language-Action Generative World Model 2024 📄

Efficient & Real-Time VLA

Compression, caching, parallel decoding, and distillation to make VLAs small and fast enough for real-time / edge control (Guan et al., 2025).

Model Title Year Links
FASTER Rethinking Real-Time Flow VLAs 2026 📄
RTC Real-Time Chunking: Running VLAs at Real-Time Speed 2025 📄
NanoVLA Routing-Decoupled VLA for Nano-Sized Generalist Policies 2025 📄
VLA-Adapter A Tiny-Scale VLA Paradigm 2025 📄
OpenVLA-OFT Fine-Tuning VLAs: Optimizing Speed and Success 2025 📄 · 🌐
TinyVLA Fast, Data-Efficient VLA Models 2024 📄 · 🌐

RL Fine-Tuning for VLA

Reinforcement learning (often on top of flow-/diffusion-based VLAs) to improve over imitation-only training.

Model Title Year Links
π_RL Online RL Fine-Tuning for Flow-based VLAs 2025 📄
VLA-RFT RL Fine-Tuning with Verified Rewards in World Simulators 2025 📄
SimpleVLA-RL Scaling VLA Training via Reinforcement Learning 2025 📄
ConRFT A Reinforced Fine-Tuning Method for VLA via Consistency Policy 2025 📄

🌎 World & World-Action Models

Organized by what the model predicts and how it is built, following the embodied-world-model taxonomy of Li et al., 2025 and the WAM split popularized by awesome-vla-wam.

General World Models

General-purpose models of environment dynamics — spanning classical latent world models for model-based RL (World Models, DreamerV3) and modern large-scale video / foundation world models — used for planning, neural simulation, or as backbones for WAMs.

Model Title Year Links
Cosmos-Predict2.5 World Simulation with Video Foundation Models for Physical AI 2025 📄 · 💻
Cosmos-Reason1 From Physical Common Sense to Embodied Reasoning 2025 📄
Cosmos World Foundation Model Platform for Physical AI 2025 📄 · 🌐
V-JEPA 2 Self-Supervised Video Models Enable Understanding, Prediction & Planning 2025 📄
iVideoGPT Interactive VideoGPTs are Scalable World Models 2024 📄
Genie Generative Interactive Environments 2024 📄
DreamerV3 Mastering Diverse Domains through World Models 2023 📄 · 💻
UniSim Learning Interactive Real-World Simulators 2023 📄
World Models Recurrent latent world model + controller (origin of the term) 2018 📄

WAM from Video Generation

A (text-/image-conditioned) video generator imagines future frames; actions are recovered via an inverse-dynamics / action head.

Model Title Year Links
DreamZero World Action Models are Zero-shot Policies 2026 📄 · 🌐
DiT4DiT Jointly Modeling Video Dynamics and Actions 2026 📄
Cosmos Policy Fine-Tuning Video Models for Visuomotor Control & Planning 2026 📄 · 🌐
Video2Act A Dual-System Video Diffusion Policy 2025 📄
GR-2 A Generative Video-Language-Action Model with Web-Scale Knowledge 2024 📄
GR-1 Large-Scale Video Generative Pre-training for Visual Robot Manipulation 2023 📄

WAM from VLMs

A pretrained VLM is turned into a world model (e.g., predicting goal images / object-centric futures) that then drives action.

Model Title Year Links
DreamVLA A VLA Model Dreamed with Comprehensive World Knowledge 2025 📄
Goal-VLA Image-Generative VLMs as Object-Centric World Models for VLA 2025 📄

Unified VLA–World Models

Single architectures that jointly learn to act and to predict world dynamics, blurring the VLA/WAM boundary.

Model Title Year Links
RynnVLA-002 A Unified Vision-Language-Action and World Model 2025 📄 · 💻
WholeBodyVLA Unified Latent VLA for Whole-Body Loco-Manipulation 2025 📄 · 💻
WorldVLA Towards an Autoregressive Action World Model 2025 📄 · 💻

Latent & JEPA World Models

Self-supervised latent predictive models (non-reconstructive joint-embedding / JEPA). The JEPA foundations (I-JEPA) learn to predict in representation space; the action-conditioned variant (V-JEPA 2-AC) turns that prior into a world model for planning.

Model Title Year Links
V-JEPA 2-AC Action-Conditioned Latent World Model for Zero-Shot Planning 2025 📄
I-JEPA Image-based Joint-Embedding Predictive Architecture (representation foundation) 2023 📄

Domain World Models (Driving & Navigation)

Model Title Year Links
GAIA-2 A Controllable Multi-View Generative World Model for Autonomous Driving 2025 📄
Navigation World Models Conditional Diffusion Transformer for Navigation 2024 📄
GAIA-1 A Generative World Model for Autonomous Driving 2023 📄

🧩 Action Representations & Tokenization

Building blocks shared across VLA and WAM policies — how continuous actions become learnable targets.

Discrete / Autoregressive Tokenizers

Method Title Year Links
FAST Efficient (DCT-based) Action Tokenization for VLAs 2025 📄 · 🌐
BeT Behavior Transformers: Cloning k Modes with One Stone 2022 📄

Continuous & Chunked Action Policies

Heads that emit continuous action chunks — by denoising diffusion (Diffusion Policy) or by chunked sequence prediction with a CVAE (ACT). Flow-matching heads (π0, SmolVLA, …) are listed with their models under Flow-Matching VLA.

Method Title Year Links
Diffusion Policy Visuomotor Policy Learning via Action Diffusion 2023 📄 · 🌐
ACT / ALOHA Action Chunking with Transformers 2023 📄 · 🌐

🦾 Foundational Robot Policies

Non-VLA policies and planners that remain standard baselines in the experimental tables of the papers above. (Diffusion Policy, ACT, and BeT are described under Action Representations.)

Method Title Year Links
CrossFormer Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion & Flight 2024 📄 · 💻
RoboFlamingo Vision-Language Foundation Models as Effective Robot Imitators 2023 📄 · 💻
VoxPoser Composable 3D Value Maps for Robotic Manipulation (zero-shot LLM + 3D planner) 2023 📄 · 🌐
RT-1 Robotics Transformer for Real-World Control at Scale 2022 📄

📦 Resources

Datasets

Name Description Scale Links
Open X-Embodiment Cross-embodiment aggregation behind the RT-X models 1M+ traj · 22 embodiments 📄 · 🌐
AgiBot World Large-scale real-world manipulation (Colosseo) 1M+ traj · 217 tasks 📄 · 🌐
EgoScale Scaling dexterous manipulation with diverse egocentric human data egocentric · 2026 📄
DexCanvas Human demos ↔ robot learning for dexterous manipulation dexterous 📄
Galaxea Open-World Mobile-bimanual dataset paired with the G0 dual-system VLA 500 hrs · 150 tasks 📄 · 💻
DROID In-the-wild Franka manipulation across 3 continents 76K traj · 564 scenes 📄 · 🌐
RoboMIND Multi-embodiment teleop incl. labeled failures 107K traj · 479 tasks 📄 · 🌐
BridgeData V2 WidowX manipulation w/ language + goal images 60K traj · 24 envs 📄 · 🌐
RH20T Contact-rich skills w/ paired human demos 110K+ seq · 147 tasks 📄 · 🌐
Ego-Exo4D Simultaneous ego + exo video of skilled activity 1,286 hrs 📄 · 🌐
Ego4D Massive egocentric daily-life video 3,670 hrs 📄 · 🌐

Benchmarks

Name Description Links
LIBERO Lifelong robot-learning, 130 manipulation tasks (de-facto VLA eval) 📄 · 💻
CALVIN Long-horizon language-conditioned manipulation 📄 · 💻
SimplerEnv Real-to-sim evaluation for manipulation policies 📄 · 🌐
RoboCasa Large-scale kitchen simulation (100 tasks) 📄 · 🌐
VLABench World-knowledge & long-horizon language tasks 📄 · 🌐
ManiSkill3 GPU-parallel manipulation (30K+ FPS) 📄 · 🌐
THE COLOSSEUM Robustness under 14 environmental perturbations 📄 · 🌐
RoboArena Distributed crowd-sourced real-world policy eval 📄 · 🌐
RoboChallenge Large-scale real-robot evaluation of embodied policies 📄
RobotArena ∞ Scalable robot benchmarking via real-to-sim translation 📄 · 🌐
WorldArena Perception & functional-utility benchmark for embodied world models 📄
Meta-World 50 tabletop tasks for multi-task / meta-RL 📄 · 💻
RLBench 100 hand-designed manipulation tasks 📄 · 💻

Simulation Platforms

Name Description Links
Isaac Sim / Isaac Lab GPU-native robotics sim + RL/IL framework (Omniverse/USD) 🌐
MuJoCo / MJX Standard rigid-body engine + JAX/XLA parallel variant 🌐
Genesis Generative, multi-solver physics platform (up to ~43M FPS) 🌐
ManiSkill GPU-parallel manipulation simulator on SAPIEN 🌐
SAPIEN Part-level articulated-object simulator (PartNet-Mobility) 📄 · 🌐
Habitat Photorealistic indoor navigation & rearrangement 🌐
ThreeDWorld Multimodal Unity3D sim (vision + audio + physics) 📄 · 🌐
Newton Open, differentiable GPU physics engine (NVIDIA + DeepMind + Disney) 🌐

Tools & Frameworks

Name Description Links
LeRobot End-to-end PyTorch robot-learning library + datasets + low-cost HW 📄 · 💻
openpi Open models & training/inference for π0, π0-FAST, π0.5 💻
Isaac GR00T Open humanoid foundation-model framework + checkpoints 💻
OpenVLA Training / LoRA fine-tuning for the 7B OpenVLA model 💻
Octo JAX/Flax generalist transformer policy on OXE 💻
robomimic / robosuite Learning-from-demonstration framework + MuJoCo manipulation sim 💻
HIL-SERL Human-in-the-loop, sample-efficient real-world RL 💻

🗂️ Extended Paper Index (Auto-Curated, Newest First)

A broader, continuously-mined index of recent arXiv work that complements the curated highlights above — 183 additional papers, newest first. Last updated: 2026-07-16. Auto-generated from data/*.json by scripts/expand_papers.py; papers already highlighted above are omitted here to avoid duplication.

VLA — General & Manipulation · 45 papers
Paper Authors Date Links
Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Yao He, Gan Sun et al. 2026-07-16
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models Wei Li, Peijin Jia et al. 2026-07-16
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Yufeng Ji, Wenhao Tang et al. 2026-07-16
DiMaS: Distribution Matching for Steering Vision-Language-Action Models Pegah Khayatan, Sara Meziane et al. 2026-07-15
An Empirical Study on Stage-Information Interfaces for VLA Fine-Tuning Yingwei Ji 2026-07-15
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Dwip Dalal, Shivansh Patel et al. 2026-07-15
On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation Finn Ferchau, Daniel Pommer et al. 2026-07-11
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model Ronghan Chen, Yandan Yang et al. 2026-07-01
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models Yichi Zhang, Weihao Yuan et al. 2026-03-30
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation Hongyu Yan, Qiwei Li et al. 2026-03-29
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation Yang Liu, Pengxiang Ding et al. 2026-03-26
ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making Young-Chae Son, Dae-Kwan Ko et al. 2026-03-26
$π$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial Manipulation Johnathan Tucker, Denis Liu et al. 2026-03-26
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models Jiaying Zhou, Zhihao Zhan et al. 2026-03-25
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs Haoran Yuan, Weigang Yi et al. 2026-03-24
Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation Anupam Pani, Yanchao Yang 2026-03-24
CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models Youzhi Liu, Li Gao et al. 2026-03-24
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models Zhou Fang, Jiaqi Wang et al. 2026-03-18
The Steep-spectrum Radio-loud AGN Luminosity Function and Its Implications for Black Hole Growth and Star Formation Wenjie Wang, Zunli Yuan et al. 2026-03-16
Building Explicit World Model for Zero-Shot Open-World Object Manipulation Xiaotong Li, Gang Chen et al. 2026-03-14
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation Minghao Jin, Mozheng Liao et al. 2026-03-13
Adaptive Capacity Allocation for Vision Language Action Fine-tuning Donghoon Kim, Minji Bae et al. 2026-03-08
HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild Ziyang Zhao, Shuheng Wang et al. 2026-03-06
CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuning Yike Zhang, Yaonan Wang et al. 2026-02-13
HoloBrain-0 Technical Report Horizon Robotics 2026-02-07 🌐
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion Ralf Römer, Yi Zhang et al. 2026-01-14
Diverse stages of star formation in the IRAS 18162-2048 region. Emergence of UV Feedback R. Fedriani, G. Anglada et al. 2025-12-08
Mixture of Horizons in Action Chunking Dong Jing, Gang Wang et al. 2025-11-24
A scaling relationship for non-thermal radio emission from ordered magnetospheres - II. Investigating the efficiency of relativistic electron production in magnetospheres of BA-type stars P. Leto, S. Owocki et al. 2025-11-07
First X-ray and radio polarimetry of the neutron star low-mass X-ray binary GX 17+2 Unnati Kashyap, Thomas J. Maccarone et al. 2025-10-06
Deciphering the radio-star formation correlation on kpc scales. IV. Radio halos of highly-inclined Virgo cluster spiral galaxies B. Vollmer, M. Soida et al. 2025-10-03
Masses, Star-Formation Efficiencies, and Dynamical Evolution of 18,000 HII Regions Debosmita Pathak, Adam K. Leroy et al. 2025-09-26
X-ray and radio polarimetry of the neutron star low mass X-ray binary GX 13+1 Unnati Kashyap, Thomas J. Maccarone et al. 2025-08-07
Protostellar Outflows at the EarliesT Stages (POETS). VIII. The jets in the intermediate-mass star-forming region G105.42+9.88 (alias LkHα 234) Luca Moscadelli, Fabrizio Massi et al. 2025-08-05
Star formation histories and gas content limits of three ultra-faint dwarfs on the periphery of M31 Michael G. Jones, David J. Sand et al. 2025-08-01
Quenching Through Tidal Gas Removal: Molecular Gas and Star Formation in Tidal Tails of z ~ 0.7 Post-Starburst Galaxies Vincenzo R. D'Onofrio, Justin S. Spilker et al. 2025-07-28
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation Sixiang Chen, Jiaming Liu et al. 2025-07-02
The Radio Spectral Energy Distribution and Star Formation Calibration in MIGHTEE-COSMOS Highly Star-Forming Galaxies at 1.5 < z < 3.5 Fatemeh Tabatabaei, Maryam Khademi et al. 2025-06-19
Semi-empirical constraints on the HI mass function of star-forming galaxies and $Ω_{\rm HI}$ at $z\sim 0.37$ from interferometric surveys Francesco Sinigaglia, Alessandro Bianchetti et al. 2025-06-12
A persistent disk wind and variable jet outflow in the neutron-star low-mass X-ray binary GX 13+1 Daniele Rogantini, Jeroen Homan et al. 2025-04-07
X-ray and radio data obtained by XMM-Newton and VLA constrain the stellar wind of the magnetic quasi-Wolf-Rayet star in HD45166 P. Leto, L. M. Oskinova et al. 2025-03-10
The Arp 240 Galaxy Merger: A Detailed Look at the Molecular Kennicutt-Schmidt Star Formation Law on Sub-kpc Scales Alejandro Saravia, Eduardo Rodas-Quito et al. 2024-12-10
Runaway O and Be stars found using Gaia DR3, new stellar bow shocks and search for binaries M. Carretero-Castrillo, M. Ribó et al. 2024-12-10
A-VL: Adaptive Attention for Large Vision-Language Models Junyang Zhang, Mu Yuan et al. 2024-09-23
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers Jianke Zhang, Yanjiang Guo et al. 2024-09-12
VLA — Reasoning, Planning & Dual-System · 6 papers
Paper Authors Date Links
ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception Weichen Zhang, Shiquan Yu et al. 2026-07-11
Do World Action Models Generalize Better than VLAs? A Robustness Study Zhanguang Zhang, Zhiyuan Li et al. 2026-03-23
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models Riccardo Andrea Izzo, Gianluca Bardaro et al. 2026-03-05
Chain of World: World Model Thinking in Latent Motion Fuxiang Yang, Donglin Di et al. 2026-03-03
FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment Han Zhao, Jingbo Wang et al. 2026-02-19
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Changhua Xu, Jie Lu et al. 2026-02-07
VLA — Autonomous Driving · 12 papers
Paper Authors Date Links
Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection Yi Wang, Wendi Chen et al. 2026-07-15
S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving Jianguo Yu, Rukang Wang et al. 2026-07-15
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation Yiran Shi, Dongqi Guo et al. 2026-03-30
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving Qiqi Liu, Huan Xu et al. 2026-03-28
Vega: Learning to Drive with Natural Language Instructions Sicheng Zuo, Yuxuan Li et al. 2026-03-26
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving Zehao Wang, Huaide Jiang et al. 2026-03-26
ETA-VLA: Efficient Token Adaptation via Temporal Fusion and Intra-LLM Sparsification for Vision-Language-Action Models Yiru Wang, Anqing Jiang et al. 2026-03-26
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models Jintao Cheng, Haozhe Wang et al. 2026-03-24
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving Zihan You, Hongwei Liu et al. 2026-03-09
VLANeXt: Recipes for Building Strong VLA Models Liu, Xiangyu et al. 2026-02-18 🌐
A global view on star formation: The GLOSTAR Galactic plane survey XII. Effelsberg's continuum view and data release Y. Gong, W. Reich et al. 2025-12-17
Low-frequency spectra of neutron star + OB supergiant binaries: Does wind density drive persistent and flaring modes of accretion? J. van den Eijnden, L. Sidoli et al. 2025-08-06
VLA — Dexterous & Humanoid · 1 papers
Paper Authors Date Links
Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models Ruixing Jin, Zicheng Zhu et al. 2026-03-24
VLA — 3D / 4D & Spatial · 6 papers
Paper Authors Date Links
CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking Ruilong Ren, Songsheng Cheng et al. 2026-07-16
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation Mohan Liu, Zhihao Gu et al. 2026-07-14
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models Byungkun Lee, Dongyoon Hwang et al. 2026-07-13
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Shengzhuo Yang, Ronghao Yu et al. 2026-07-10
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior Xinkai Wang, Chenyi Wang et al. 2026-03-26
3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models Bin Yu, Shijie Lian et al. 2026-03-25
VLA — RL & Post-Training · 6 papers
Paper Authors Date Links
ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning Yilun Kong, Yunpeng Qing et al. 2026-07-14
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation Zhide Zhong, Haodong Yan et al. 2026-03-27
On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning Changyu Liu, Yiyang Liu et al. 2026-01-11
VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning Si-Cheng Wang, Tian-Yu Xiang et al. 2025-09-30
The arc-shaped radio source at the center of NGC 6334A: Is it a colliding wind region of two young massive stars or the bow shock of a runaway star? Vanessa Yanza, Sergio A. Dzib et al. 2025-02-24
VLA 22 GHz Imaging of Massive Star Formation in Local Wolf-Rayet Galaxies Nicholas G. Ferraro, Jean L. Turner et al. 2024-11-09
VLA — Efficient & Real-Time · 19 papers
Paper Authors Date Links
Reflex: Real-Time VLA Control through Streaming Inference Yuanchun Guo, Bingyan Liu 2026-07-16
Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference Yuzhou Wu, Yuxin Zheng et al. 2026-07-14
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Yi Chen, Yuying Ge et al. 2026-03-31
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate Chen Yang, Yucheng Hu et al. 2026-03-27
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching Jiayi Chen, Wenxuan Song et al. 2026-03-27
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance Wenxuan Song, Jiayi Chen et al. 2026-03-26
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models Peiju Liu, Jinming Liu et al. 2026-03-26
Agile-VLA: Few-Shot Industrial Pose Rectification via Implicit Affordance Anchoring Teng Yan, Zhengyang Pei et al. 2026-03-24
Fast-WAM: Do World Action Models Need Test-time Future Imagination? Tianyuan Yuan, Zibin Dong et al. 2026-03-17
FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic Manipulation Yao Li, Peiyuan Tang et al. 2026-02-27
Learning Native Continuation for Action Chunking Flow Policies Yufeng Liu, Hang Yu et al. 2026-02-13
Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models Yuting Huang, Leilei Ding et al. 2026-01-31
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching Yujie Wei, Jiahan Fan et al. 2026-01-31
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation Wenda Yu, Tianshi Wang et al. 2026-01-27
U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation Linzhi Wu, Aoran Mei et al. 2025-09-29
Leave No Observation Behind: Real-time Correction for VLA Action Chunks Kohei Sendai, Maxime Alvarez et al. 2025-09-27
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding Wenxuan Song, Jiayi Chen et al. 2025-03-04
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning Zhiwei Hao, Jianyuan Guo et al. 2024-10-23
VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks Yi-Lin Sung, Jaemin Cho et al. 2021-12-13
VLA — Safety, Robustness & Evaluation · 13 papers
Paper Authors Date Links
Lights, Camera, Malfunction: When Illumination Robustness Leaves VLA Models Blind to Color Marino Watanabe, Takami Sato et al. 2026-07-16
TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors Pinhan Fu, Xianda Guo et al. 2026-07-14
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models Xiyang Wu, Guangyao Shi et al. 2026-03-26
SOMA: Strategic Orchestration and Memory-Augmented System for Vision-Language-Action Model Robustness via In-Context Adaptation Zhuoran Li, Zhiyang Li et al. 2026-03-25
ROBOGATE: Adaptive Failure Discovery for Safe Robot Policy Deployment via Two-Stage Boundary-Focused Sampling Azuki Kim 2026-03-23
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models Zhilong Zhang, Haoxiang Ren et al. 2026-03-21
Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control Zunzhe Zhang, Runhan Huang et al. 2026-03-18
World2Act: Latent Action Post-Training via Skill-Compositional World Models An Dinh Vuong, Tuan Van Vo et al. 2026-03-11
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model Yuanjie Lu, Beichen Wang et al. 2026-03-09
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models Xiaoquan Sun, Zetian Xu et al. 2026-03-09
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models Hyeongjun Heo, Seungyeon Woo et al. 2026-03-06
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models Hyeonbeom Choi, Daechul Ahn et al. 2026-02-04
SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models Bingxin Xu, Yuzhang Shang et al. 2026-01-20
World Models — General & Foundation · 12 papers
Paper Authors Date Links
KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation Xinyu Shao, Keru Zhou et al. 2026-07-06
VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation Shuai Tian, Yupeng Zheng et al. 2026-07-02
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation GigaWorld Team, Angyuan Ma et al. 2026-07-02
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model Quankai Gao, Jiawei Yang et al. 2026-03-28
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation Yuhang Zheng, Songen Gu et al. 2026-03-19
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation Jacob Levy, Tyler Westenbroek et al. 2026-03-16
WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems Yuchen Wang, Jiangtao Kong et al. 2026-03-15
ResWM: Residual-Action World Model for Visual RL Jseen Zhang, Gabriel Adineera et al. 2026-03-11
MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation Yutong Shen, Hangxu Liu et al. 2026-03-09
Foundational World Models Accurately Detect Bimanual Manipulator Failures Isaac R. Ward, Michelle Ho et al. 2026-03-07
What if? Emulative Simulation with World Models for Situated Reasoning Ruiping Liu, Yufan Chen et al. 2026-03-06
Self-adapting Robotic Agents through Online Continual Reinforcement Learning with World Model Feedback Fabian Domberg, Georg Schildbach 2026-03-04
World Models — Video Generation & WAM · 8 papers
Paper Authors Date Links
From World Models to World Action Models: A Concise Tutorial for Robotics Xiaoxiong Zhang, Xiong Zeng et al. 2026-07-01
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation Ziyu Shan, Zhenyu Wu et al. 2026-06-30
HCLSM: Hierarchical Causal Latent State Machines for Object-Centric World Modeling Jaber Jaber, Osama Jaber 2026-03-31
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning Jai Bardhan, Patrik Drozdik et al. 2026-03-26
EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards Ruixiang Wang, Qingming Liu et al. 2026-03-18
DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models Emily Yue-Ting Jia, Weiduo Yuan et al. 2026-03-17
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation Mutian Xu, Tianbao Zhang et al. 2026-03-17
PlayWorld: Learning Robot World Models from Autonomous Play Tenny Yin, Zhiting Mei et al. 2026-03-09
World Models — Driving & Navigation · 4 papers
Paper Authors Date Links
Enhancing Policy Learning with World-Action Model Yuci Han, Alper Yilmaz 2026-03-30
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving Linbo Wang, Yupeng Zheng et al. 2026-03-25
NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation Tianshuai Hu, Zeying Gong et al. 2026-03-16
AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation Ge Yuan, Qiyuan Qiao et al. 2026-02-23
Policies — Diffusion & Flow · 27 papers
Paper Authors Date Links
Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy Adrien Jacquet Crétides, Mouad Abrini et al. 2026-03-17
ReMAP-DP: Reprojected Multi-view Aligned PointMaps for Diffusion Policy Xinzhang Yang, Renjun Wu et al. 2026-03-16
REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning Zhaoyuan Gu, Yipu Chen et al. 2026-03-14
PPGuide: Steering Diffusion Policies with Performance Predictive Guidance Zixing Wang, Devesh K. Jha et al. 2026-03-11
SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation Youqiang Gui, Yuxuan Zhou et al. 2026-03-05
Diffusion Policy through Conditional Proximal Policy Optimization Ben Liu, Shunpeng Yang et al. 2026-03-05
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy Pengyuan Wu, Pingrui Zhang et al. 2026-03-02
ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent Manipulation Enyi Wang, Wen Fan et al. 2026-02-25
Preference Aligned Visuomotor Diffusion Policies for Deformable Object Manipulation Marco Moletta, Michael C. Welle et al. 2026-02-10
SERFN: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Chenyu Yang, Denis Tarasov et al. 2026-02-10
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation Yuxuan Hu, Xiangyu Chen et al. 2026-02-07
Moving On, Even When You're Broken: Fail-Active Trajectory Generation via Diffusion Policies Conditioned on Embodiment and Task Gilberto G. Briscoe-Martinez, Yaashia Gautam et al. 2026-02-02
RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human Feedback Amitesh Vatsa, Zhixian Xie et al. 2026-01-31
Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation Runhua Zhang, Junyi Hou et al. 2026-01-30
Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion Policies Ce Hao, Xuanran Zhai et al. 2026-01-29
ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot Manipulation Weize Xie, Yi Ding et al. 2026-01-19
Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning Kangye Ji, Yuan Meng et al. 2026-01-19
CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action Space Bingyi Liu, Jinbo He et al. 2026-01-09
Learning Diffusion Policy from Primitive Skills for Robot Manipulation Zhihao Gu, Ming Yang et al. 2026-01-05
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control Wonhyeok Choi, Shutong Ding et al. 2026-01-05
Flexible Multitask Learning with Factorized Diffusion Policy Chaoqi Liu, Haonan Chen et al. 2025-12-26
Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation Kangchen Lv, Mingrui Yu et al. 2025-12-19
ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision Wenlong Xia, Jinhao Zhang et al. 2025-12-17
Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks Aileen Liao, Dong-Ki Kim et al. 2025-12-08
CAPE: Context-Aware Diffusion Policy Via Proximal Mode Expansion for Collision Avoidance Rui Heng Yang, Xuan Zhao et al. 2025-11-27
Learning Diffusion Policies for Robotic Manipulation of Timber Joinery under Fabrication Uncertainty Salma Mozaffari, Daniel Ruan et al. 2025-11-21
UltraDP: Generalizable Carotid Ultrasound Scanning with Force-Aware Diffusion Policy Ruoqu Chen, Xiangjie Yan et al. 2025-11-19
Policies — Imitation & Behavior Learning · 10 papers
Paper Authors Date Links
Bi-AQUA: Bilateral Control-Based Imitation Learning for Underwater Robot Arms via Lighting-Aware Action Chunking with Transformers Takeru Tsunoori, Masato Kobayashi et al. 2025-11-20
Temporal Action Selection for Action Chunking Yueyang Weng, Xiaopeng Zhang et al. 2025-11-06
FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task Ryo Watanabe, Maxime Alvarez et al. 2025-09-27
Action Chunking with Transformers for Image-Based Spacecraft Guidance and Control Alejandro Posadas-Nava, Andrea Scorsoglio et al. 2025-09-04
LiPo: A Lightweight Post-optimization Framework for Smoothing Action Chunks Generated by Learned Policies Dongwoo Son, Suhan Park 2025-06-05
Bi-LAT: Bilateral Control-Based Imitation Learning via Natural Language and Action Chunking with Transformers Takumi Kobayashi, Masato Kobayashi et al. 2025-04-02
Cross-Embodiment Robotic Manipulation Synthesis via Guided Demonstrations through CycleVAE and Human Behavior Transformer Apan Dastider, Hao Fang et al. 2025-03-11
Memorized action chunking with Transformers: Imitation learning for vision-based tissue surface scanning Bochen Yang, Kaizhong Deng et al. 2024-11-06
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling Yuejiang Liu, Jubayer Ibn Hamid et al. 2024-08-30
Surgical Robot Transformer (SRT): Imitation Learning for Surgical Tasks Ji Woong Kim, Tony Z. Zhao et al. 2024-07-17
Policies — Robot Learning & Manipulation · 14 papers
Paper Authors Date Links
Learning Multi-View Spatial Reasoning from Cross-View Relations Suchae Jeong, Jaehwi Song et al. 2026-03-30
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation Motonari Kambara, Koki Seno et al. 2026-03-26
Chunk-Boundary Artifact in Action-Chunked Generative Policies: A Noise-Sensitive Failure Mechanism Rui Wang 2026-03-12
Real-Time Robot Execution with Masked Action Chunking Haoxuan Wang, Gengyu Zhang et al. 2026-01-27
Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward Jiarui Yang, Bin Zhu et al. 2025-08-15
Reinforcement Learning with Action Chunking Qiyang Li, Zhiyuan Zhou et al. 2025-07-10
Real-Time Execution of Action Chunking Flow Policies Kevin Black, Manuel Y. Galliker et al. 2025-06-09
Learning Bimanual Manipulation via Action Chunking and Inter-Arm Coordination with Transformers Tomohiro Motoda, Ryo Hanai et al. 2025-03-18
MissionGPT: Mission Planner for Mobile Robot based on Robotics Transformer Model Vladimir Berman, Artem Bazhenov et al. 2024-11-07
VQ-ACE: Efficient Policy Search for Dexterous Robotic Manipulation via Action Chunking Embedding Chenyu Yang, Davide Liconti et al. 2024-11-05
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation Andrew Lee, Ian Chuang et al. 2024-09-12
Bringing the RT-1-X Foundation Model to a SCARA robot Jonathan Salzer, Arnoud Visser 2024-09-05
Logically Constrained Robotics Transformers for Enhanced Perception-Action Planning Parv Kapoor, Sai Vemprala et al. 2024-08-09
SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust Attention Isabel Leal, Krzysztof Choromanski et al. 2023-12-04

📋 Full Paper Index & Baselines

📊 Click to expand the complete paper list and baseline methods

The curated tables above highlight landmark and representative work. For the exhaustive, auto-maintained index and the baseline methods extracted from experimental tables, see:

Quick Reference (common baselines)

Family Key baselines
VLA RT-1, RT-2, OpenVLA, Octo, π0, π0.5, X-VLA, UniVLA, SmolVLA
Policy Diffusion Policy, ACT, BeT, RoboFlamingo, CrossFormer
World Model DreamerV3, I-JEPA, V-JEPA 2, Genie, Cosmos, GR-1/GR-2

🤝 Contributing

Contributions are very welcome! To add or fix a paper:

  1. Add a paper — open a PR placing it in the appropriate category (keep tables sorted newest-first), or open an issue with the arXiv link.
  2. Fix an error — submit a PR with the correction.
  3. New papers appear automatically — the 🆕 Latest Papers section and the 🗂️ Extended Paper Index are regenerated daily by the scraper; do not hand-edit content between the auto markers.

To run the discovery pipeline locally:

pip install -r requirements.txt
python scripts/arxiv_scraper.py --max-results 50 --days-back 30   # writes data/papers.json
python scripts/update_readme.py                                  # refreshes the 🆕 auto section
python scripts/expand_papers.py                                  # refreshes the 🗂️ extended index
python scripts/build_site.py                                     # rebuilds the GitHub Pages site (index.html)

License

Released under the MIT License.

Acknowledgments

Inspired by awesome-vla-wam, awesome-physical-ai, and awesome-vla-study. Taxonomy grounded in the surveys listed above.


If you find this repository useful, please consider giving it a ⭐

About

A curated list of academic papers and resources on Vision-Language-Action (VLA) and World Action Models (WAM)

Topics

Resources

License

Stars

30 stars

Watchers

1 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors

Languages