Skip to content

Latest commit

 

History

History
128 lines (115 loc) · 8.88 KB

File metadata and controls

128 lines (115 loc) · 8.88 KB

Comparison Methods & Baselines

This document contains baseline methods extracted from experimental tables of major VLA/WAM papers.

VLA Baselines

Method Paper Year Description
DIAL Decoupling Intent and Action via Latent World Modeling 2026 Intent-action decoupling with latent world modeling
VLANeXt Recipes for Building Strong VLA Models 2026 VLA architecture recipes
FocusVLA Focused Visual Utilization for VLA Models 2026 Focused visual attention in VLA
StreamingVLA Streaming VLA with Action Flow Matching 2026 Streaming with flow matching
ABot-M0 VLA with Action Manifold Learning 2026 Action manifold learning
SimVLA Simple VLA Baseline for Robotic Manipulation 2026 Simple VLA baseline
Lingbot-VLA Pragmatic VLA Foundation Model 2026 Pragmatic VLA approach
Gemini Robotics Bringing AI into the Physical World 2025 Gemini-based VLA
π*0.6 A VLA That Learns From Experience 2025 Experience-learning VLA
X-VLA Soft-Prompted Cross-Embodiment VLA 2025 Soft-prompted cross-embodiment
UniVLA Unified Vision-Language-Action Model 2025 Unified VLA with world model
SmolVLA VLA for Affordable and Efficient Robotics 2025 Compact VLA for edge devices
NORA Small Open-Sourced Generalist VLA 2025 Small parameter VLA
VLA-0 Building State-of-the-Art VLAs with Zero Modification 2025 Zero-modification VLA
CronusVLA Efficient Multi-Frame VLA 2025 Multi-frame VLA modeling
OpenVLA-OFT Fine-Tuning VLAs: Optimizing Speed and Success 2025 Online fine-tuning
AsyncVLA Asynchronous Flow Matching for VLA 2025 Asynchronous flow matching
AVA-VLA VLA with Active Visual Attention 2025 Active visual attention
OpenVLA An Open-Source VLA 2024 First open-source VLA
Octo Open-Source Generalist Robot Policy 2024 Modular generalist policy
π0 A Vision-Language-Action Flow Model for Robot Control 2024 Flow matching + action expert
AC2-VLA Action-Context-Aware Adaptive Computation in VLA Yu et al. 2026
APPLV Adaptive Planner Parameter Learning from VLA Lu et al. 2026
Act, Think or Abstain Complexity-Aware Adaptive Inference for VLA Izzo et al. 2026
AnyCamVLA Zero-Shot Camera Adaptation for Viewpoint Robust VLA Heo et al. 2026
CLARE Continuous Learning for VLA via Adapter Routing Römer et al. 2026
DIAL Decoupling Intent and Action via Latent World Modeling for VLA Chen et al. 2026
DiT4DiT Jointly Modeling Video Dynamics and Actions Ma et al. 2026
EAPruning Adaptive Pruning with Interleaved Inference for VLA Huang et al. 2026
ETA-VLA Efficient Token Adaptation Wang et al. 2026
FAVLA Force-Adaptive Fast-Slow VLA Li et al. 2026
HarvestFlex Harvesting via VLA Policy Adaptation Zhao et al. 2026
On-the-Fly VLA VLA Adaptation via Test-Time RL Liu et al. 2026
ProbeFlow Training-Free Adaptive Flow Matching for VLA Fang et al. 2026
RAFT Adapting VLA Models via Force-aware Curriculum Zhang et al. 2026
ROBOGATE Adaptive Failure Discovery for Safe Robot Policy Kim et al. 2026
SAMoE-VLA Scene Adaptive Mixture-of-Experts VLA You et al. 2026
SCALE Self-Uncertainty Adaptive Looking for VLA Choi et al. 2026
SOMA Memory-Augmented System for VLA Robustness Li et al. 2026
VGAS Adaptive Capacity Allocation for VLA Kim et al. 2026
VLA-Acceleration Accelerate VLA through Visual Token Caching Wei et al. 2026
VGAS Value-Guided Action-Chunk Selection for VLA Xu et al. 2026
AC-DiT AC-DiT: Adaptive Coordination Diffusion Transformer Chen et al. 2025
U-DiT U-DiT: U-shaped Diffusion Transformers Wu et al. 2025
VLA-Adapter VLA-Adapter: Tiny-Scale VLA Paradigm Wang et al. 2025
A-VL Adaptive Attention for Large VLA Zhang et al. 2024
ADEM-VL Adaptive and Embedded Fusion for VLA Hao et al. 2024
VL-Adapter VL-Adapter: Parameter-Efficient Transfer Learning Sung et al. 2021
RT-2 Vision-Language-Action Models 2023 VLA paradigm establishment
RT-2-55B Vision-Language-Action Models 2023 55B parameter version
RT-2-1B Vision-Language-Action Models 2023 1B parameter version
RT-1 Robotics Transformer for Real-World Control 2022 First large-scale robotics transformer
RT-1-35M Robotics Transformer for Real-World Control 2022 35M parameter version

World Model Baselines

Method Paper Year Description
DreamZero World Action Models are Zero-shot Policies 2026 World model as zero-shot policy
HCLSM Hierarchical Causal Latent State Machines 2026 Hierarchical causal world model
WAM World-Action Model 2026 Action-regularized world model
DreamerV3 Mastering Atari from Pixels 2023 World model RL algorithm
DreamerV2 Mastering Atari from Pixels 2022 Improved world model
DreamerV1 Learning Behaviors from Pixels 2020 Original world model
I-JEPA Image-based Joint-Embedding Predictive Architecture 2023 Joint embedding predictive
V-JEPA Video Joint-Embedding Predictive Architecture 2024 Video joint embedding
Cosmos Policy Fine-Tuning Video Models for Visuomotor Control & Planning 2026 NVIDIA world-action policy

Policy Baselines

Method Paper Year Description
Diffusion Policy Diffusion for Robot Control 2023 Diffusion model for action generation
ACT Action Chunking Transformer 2023 CVAE-based action chunking
ACT-ALOHA Action Chunking with ALOHA 2023 Bimanual manipulation
BeT Behavior Transformers 2022 Multimodal action discretization
PerAct Behavior Primitive Discovery 2023 Per-act primitive learning
MVP Masked Visual Pre-training 2022 Masked visual encoder pretraining
R3M Universal Visual Representation 2022 Self-supervised visual representation
CQL Conservative Q-Learning 2020 Conservative offline RL
IQL Implicit Q-Learning 2021 Implicit offline RL

Dataset Baselines

Dataset Source Tasks Size
OXE Open X-Embodiment 547 500k+ episodes
RT-1 Dataset Google Research 9 130k episodes
BridgeData Berkeley AI Research 8 70k episodes
ALOHA Stanford 8 10k+ episodes
AgiBot World Agibot 100+ 1M+ episodes
UMI CMU 6 15k episodes
DROID CMU 56 80k episodes
Fractal Berkeley 20 10M+ demonstrations

Benchmark Baselines

Benchmark Tasks Evaluation Metric
Libero Object manipulation Success rate
Libero-Goal Goal-conditioned manipulation Success rate
Libero-Object Object manipulation Success rate
RLBench 100+ tasks Success rate
CALVIN Sequential manipulation Success rate
ManiSkill2 Manipulation tasks Success rate
Metaworld 50 tasks Success rate

Common Metrics

Metric Description
Success Rate Percentage of successfully completed tasks
Average Return Cumulative reward over episodes
Completion Rate Tasks completed within time limit
Efficiency Time taken to complete tasks
Robustness Performance across variations
Generalization Performance on unseen tasks/objects
Sim2Real Gap Performance drop from sim to real
Inference Time Time per action prediction
Throughput Actions per second