ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting

Image credit: ReWeight authors

Abstract

ReWeight uses egocentric human demonstrations to improve vision-language-action model post-training while reducing interference caused by differences between human and robot embodiments. A learned visuomotor representation combines observations and future actions to compare demonstrations. Optimal transport then supports retrieval of relevant human demonstrations and sample weighting that favors closer matches to robot data. Evaluations with pi-0.5 cover eight simulation tasks and four real-world tasks under clean and randomized conditions. The method reaches 57% average simulation success, compared with 39% for robot-only training and 44% for random human-robot mixing, and achieves 68.8% success in physical experiments.

Publication
arXiv preprint, 2026