Demystifying When and Why VLAs Fail in Contact-Rich Tasks and How to Fix Them

Stanford University
0:00 / 0:00
Abstract

We address the problem of understanding when and why Vision-Language-Action models struggle with contact-rich manipulation tasks that require precise physical interaction. Prior work has primarily focused on addressing contact failures through force-augmented architectures and training-time regularizers, yet the root causes of these failures remain underexplored. We identify two distinct failure modes underlying this gap. Precision failures are rooted in a flow-matching policy training mismatch, and force failures arise from the distinctive structure of force signals. We address each failure mode with a targeted mechanism and combine them into FACT, which achieves 66% average success rate across five contact-rich tasks against 41% for the best prior baseline, in an evaluation spanning almost 2,500 real-world rollouts.

Diagnosis

Compared to general manipulation tasks, contact-rich tasks require both fine-grained correction and reasoning about contact forces. Below, we illustrate these requirements through plug insertion, contrasting a successful execution with two distinct failure modes: precision failure and force failure.

Success

The plug approaches, aligns with the socket, enters and fully seats.

Precision failure

The policy fails to align the plug with the socket and stalls at the rim.

Force failure

The policy aligns the plug with the socket, but then fails to detect when the plug is fully seated.

Tracing each failure mode back to its root cause reveals two distinct sets of properties. One is rooted in how flow-matching policies are trained, the other in the structure of force signals themselves.

Precision failures
  • Delta collapse. In free space, action deltas are large and variable. During contact they collapse to near zero, and this low-delta regime must be reproduced precisely at the moment visual feedback is least informative.
  • Training starvation. Fine-grained corrections are generated in the low-τ denoising regime. Commonly used Beta noise schedules in VLAs such as π₀ [1], π₀.₅ [2], GR00T N1 [3], and SmolVLA [4] allocate little gradient signal there, starving the contact-correction regime.
Force failures
  • Contact sparsity. Force signals are approximately zero over free-space timesteps and become informative only during short contact intervals. Near-zero samples can dominate the objective and bias gradients toward ignoring force.
  • Temporal structure. The instantaneous measurement encodes the current interaction state. Recent history encodes the dynamics that led to it, so omitting either timescale discards task-relevant contact information.
  • Sensitivity modulation. The influence of force should depend on contact state. In free space, force readings should be effectively ignored, whereas at contact even small deviations should trigger corrective actions.
Approach

Each failure mode is addressed with a targeted mechanism, and together they form FACT, Force-Aware Contact-rich manipulation via Timestep modulation. Both preserve the base VLA architecture, require no additional data, and add negligible inference-time overhead.

Precision · Logit-Normal noise schedule

We propose to improve the learning of sub-millimeter actions by reinforcing the low-noise, fine-correction regime. Specifically, we replace the Beta schedule with the Logit-Normal schedule, biasing the noise distribution towards the contact-rich regime. With mode m=1.5 instead of m=0, LN allocates 6× more gradient signal to τ < 0.2 than the Beta schedule. Critically, LN requires no changes to the model architecture, no additional training data, and no extra parameters, making it directly applicable to any flow-matching VLA.

Beta
Logit-Normal
LN allocates 6.0× the gradient signal of Beta below τ = 0.20
Gradient-signal density over noise level τ. Drag to explore the schedule.
Force · Time-aware force injection

Three design choices, one per force-failure property.

Contact state injected into AdaRMSNorm layers

For the full ablation study and hyperparameter details, please refer to the paper.

Results

FACT outperforms every baseline, with consistent gains on both precision- and force-critical tasks.

39%
π₀.₅ [2]
38%
TA-VLA [6]
41%
ForceVLA [5]
56%
π₀.₅ + LN (ours)
66%
FACT (ours)

Average success rate (%) across all five tasks.

Method PlugUSBButtonBoardKeyAll
π₀.₅ [2]30.037.512.510015.039.0
TA-VLA [6]30.025.020.097.515.037.5
ForceVLA [5]32.537.512.577.542.540.5
π₀.₅ + LN (ours)50.047.557.587.537.556.0
FACT (ours)57.547.575.090.060.066.0

Success rate (%) per task. All methods fine-tuned from π₀.₅ with LoRA. Tasks span precision-critical (plug, USB) and force-critical (button, board, key) regimes.

References
  1. K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. π0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024.
  2. P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al. π0.5: A vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025.
  3. J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al. GR00T N1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025.
  4. M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti, et al. SmolVLA: A vision-language-action model for affordable and efficient robotics. arXiv preprint arXiv:2506.01844, 2025.
  5. J. Yu, H. Liu, Q. Yu, J. Ren, C. Hao, H. Ding, G. Huang, G. Huang, Y. Song, P. Cai, et al. ForceVLA: Enhancing VLA models with a force-aware MoE for contact-rich manipulation. arXiv preprint arXiv:2505.22159, 2025.
  6. Z. Zhang, H. Xu, Z. Yang, C. Yue, Z. Lin, H.-a. Gao, Z. Wang, and H. Zhao. TA-VLA: Elucidating the design space of torque-aware vision-language-action models. In 9th Conference on Robot Learning (CoRL), 2025.
BibTeX
Cite this work
@article{paresmorlans2026fact,
  title={Demystifying When and Why VLAs Fail in Contact-Rich Tasks and How to Fix Them},
  author={Parés-Morlans, Carlota and Kuhn, Nils and Liu, Isabel and Longhini, Alberta and Bohg, Jeannette},
  journal={arXiv preprint arXiv:2608.01402},
  year={2026}
}
Acknowledgements

This work was supported in part by Agile Robotics. Carlota Parés-Morlans is supported by a graduate fellowship from Knight-Hennessy Scholars at Stanford University. Nils Kuhn is supported by scholarships from the Friedrich Ebert Foundation and the German Academic Exchange Service (DAAD). Alberta Longhini is supported by a Wallenberg–Bienenstock Postdoctoral Fellowship.

We thank Hila Chefer for helpful discussions, and Michelle Yi for discussions on environment setup, data collection, and camera footage for figures.