Propose
The frozen policy proposes a motion chunk. Its end-effector waypoints make the planned movement inspectable.
π₀.₅ → proposed actionsThe policy places the 8 and 7. GPT-6 helps recover the grasp and straighten the 2, handing control back between corrections.
Same task. Same layout.Videos show control-step playback at 25 fps. Model inference waiting time is omitted. Play each rollout using its own controls.
A little help. A more capable robot.
Selective intervention improves a frozen policy while leaving most actions to it.
Astra on Call: When and How GPT-6 Should Intervene in a VLA Policy
Better assistance means recovering mistakes while preserving useful behavior. Selective takeover does both on the evaluated tasks.
Weak-set Score · 35 episodes
Policy only → selective takeover
Aggregate Score · 15 episodes
Policy only → selective takeover
Weak set · versus full GPT-6 control
Using the paper’s timing accounting
7 tasks · 35 episodes
Score = mean partial credit, from 0 to 100. It is not success rate.
Timing: 1,269 s vs 1,753 s per weak-set episode, accounting for 16 s per model call, 0.3 s per control step, and 3 s per consultation. These figures describe the study’s execution setup.
The biggest advantage over full GPT-6 control comes from keeping more of the policy’s original successes.
Fixed-15 across the matched weak and strong sets. Full GPT-6 control preserves 2/13 and recovers 7/37.
A scene, a proposed motion, and a choice. The model inspects what the policy intends to do before deciding whether to help. No policy retraining.
The frozen policy proposes a motion chunk. Its end-effector waypoints make the planned movement inspectable.
π₀.₅ → proposed actionsGPT-6 considers the scene, task, execution feedback, and proposed motion. Fixed-15 consults every 15 policy steps.
Scene + plan + feedbackKeep the policy in control, take over an action segment, or—in the editing variants—adjust the proposed waypoints.
Then observe and repeat ↺Bounded waypoint editing lowers weak-set Score at three of four evaluated Fixed-K intervals. At K = 15, segment takeover scores 42.9, compared with 25.1 for EditOnly.
The correction interface matters: a successful local adjustment in one rollout does not imply a better controller overall.
3 of 4 intervals favor segment takeover.RoboDojo manipulation tasks probe both recovery and preservation: seven tasks where the baseline struggles, and three where it already performs well.
These are results from simulated manipulation with a frozen π₀.₅ policy. Fixed-15 is an observed strong setting on this evaluation, and aggregate preservation does not mean that every individual success is retained.