| DIAL |
Decoupling Intent and Action via Latent World Modeling |
2026 |
Intent-action decoupling with latent world modeling |
| VLANeXt |
Recipes for Building Strong VLA Models |
2026 |
VLA architecture recipes |
| FocusVLA |
Focused Visual Utilization for VLA Models |
2026 |
Focused visual attention in VLA |
| StreamingVLA |
Streaming VLA with Action Flow Matching |
2026 |
Streaming with flow matching |
| ABot-M0 |
VLA with Action Manifold Learning |
2026 |
Action manifold learning |
| SimVLA |
Simple VLA Baseline for Robotic Manipulation |
2026 |
Simple VLA baseline |
| Lingbot-VLA |
Pragmatic VLA Foundation Model |
2026 |
Pragmatic VLA approach |
| Gemini Robotics |
Bringing AI into the Physical World |
2025 |
Gemini-based VLA |
| π*0.6 |
A VLA That Learns From Experience |
2025 |
Experience-learning VLA |
| X-VLA |
Soft-Prompted Cross-Embodiment VLA |
2025 |
Soft-prompted cross-embodiment |
| UniVLA |
Unified Vision-Language-Action Model |
2025 |
Unified VLA with world model |
| SmolVLA |
VLA for Affordable and Efficient Robotics |
2025 |
Compact VLA for edge devices |
| NORA |
Small Open-Sourced Generalist VLA |
2025 |
Small parameter VLA |
| VLA-0 |
Building State-of-the-Art VLAs with Zero Modification |
2025 |
Zero-modification VLA |
| CronusVLA |
Efficient Multi-Frame VLA |
2025 |
Multi-frame VLA modeling |
| OpenVLA-OFT |
Fine-Tuning VLAs: Optimizing Speed and Success |
2025 |
Online fine-tuning |
| AsyncVLA |
Asynchronous Flow Matching for VLA |
2025 |
Asynchronous flow matching |
| AVA-VLA |
VLA with Active Visual Attention |
2025 |
Active visual attention |
| OpenVLA |
An Open-Source VLA |
2024 |
First open-source VLA |
| Octo |
Open-Source Generalist Robot Policy |
2024 |
Modular generalist policy |
| π0 |
A Vision-Language-Action Flow Model for Robot Control |
2024 |
Flow matching + action expert |
| AC2-VLA |
Action-Context-Aware Adaptive Computation in VLA |
Yu et al. |
2026 |
| APPLV |
Adaptive Planner Parameter Learning from VLA |
Lu et al. |
2026 |
| Act, Think or Abstain |
Complexity-Aware Adaptive Inference for VLA |
Izzo et al. |
2026 |
| AnyCamVLA |
Zero-Shot Camera Adaptation for Viewpoint Robust VLA |
Heo et al. |
2026 |
| CLARE |
Continuous Learning for VLA via Adapter Routing |
Römer et al. |
2026 |
| DIAL |
Decoupling Intent and Action via Latent World Modeling for VLA |
Chen et al. |
2026 |
| DiT4DiT |
Jointly Modeling Video Dynamics and Actions |
Ma et al. |
2026 |
| EAPruning |
Adaptive Pruning with Interleaved Inference for VLA |
Huang et al. |
2026 |
| ETA-VLA |
Efficient Token Adaptation |
Wang et al. |
2026 |
| FAVLA |
Force-Adaptive Fast-Slow VLA |
Li et al. |
2026 |
| HarvestFlex |
Harvesting via VLA Policy Adaptation |
Zhao et al. |
2026 |
| On-the-Fly VLA |
VLA Adaptation via Test-Time RL |
Liu et al. |
2026 |
| ProbeFlow |
Training-Free Adaptive Flow Matching for VLA |
Fang et al. |
2026 |
| RAFT |
Adapting VLA Models via Force-aware Curriculum |
Zhang et al. |
2026 |
| ROBOGATE |
Adaptive Failure Discovery for Safe Robot Policy |
Kim et al. |
2026 |
| SAMoE-VLA |
Scene Adaptive Mixture-of-Experts VLA |
You et al. |
2026 |
| SCALE |
Self-Uncertainty Adaptive Looking for VLA |
Choi et al. |
2026 |
| SOMA |
Memory-Augmented System for VLA Robustness |
Li et al. |
2026 |
| VGAS |
Adaptive Capacity Allocation for VLA |
Kim et al. |
2026 |
| VLA-Acceleration |
Accelerate VLA through Visual Token Caching |
Wei et al. |
2026 |
| VGAS |
Value-Guided Action-Chunk Selection for VLA |
Xu et al. |
2026 |
| AC-DiT |
AC-DiT: Adaptive Coordination Diffusion Transformer |
Chen et al. |
2025 |
| U-DiT |
U-DiT: U-shaped Diffusion Transformers |
Wu et al. |
2025 |
| VLA-Adapter |
VLA-Adapter: Tiny-Scale VLA Paradigm |
Wang et al. |
2025 |
| A-VL |
Adaptive Attention for Large VLA |
Zhang et al. |
2024 |
| ADEM-VL |
Adaptive and Embedded Fusion for VLA |
Hao et al. |
2024 |
| VL-Adapter |
VL-Adapter: Parameter-Efficient Transfer Learning |
Sung et al. |
2021 |
| RT-2 |
Vision-Language-Action Models |
2023 |
VLA paradigm establishment |
| RT-2-55B |
Vision-Language-Action Models |
2023 |
55B parameter version |
| RT-2-1B |
Vision-Language-Action Models |
2023 |
1B parameter version |
| RT-1 |
Robotics Transformer for Real-World Control |
2022 |
First large-scale robotics transformer |
| RT-1-35M |
Robotics Transformer for Real-World Control |
2022 |
35M parameter version |