SHARED ROBOT CONTROL

When should GPT-6
take the wheel?

01 / SEE THE DIFFERENCEActual RoboDojo rollouts · simulation
Arrange the largest number
Astra on Call Fixed-15
Six selective takeovers help finish the sequence.Success
Policy only Base-25
The policy leaves this layout unfinished.Incomplete

The policy places the 8 and 7. GPT-6 helps recover the grasp and straighten the 2, handing control back between corrections.

Same task. Same layout.

Videos show control-step playback at 25 fps. Model inference waiting time is omitted. Play each rollout using its own controls.

A little help. A more capable robot.
Selective intervention improves a frozen policy while leaving most actions to it.

Astra on Call: When and How GPT-6 Should Intervene in a VLA Policy

Anonymous authors

THE STUDY AT A GLANCEFigure 1 · from the paper
FIGURE 1The full paper teaser: how control is shared, when Fixed-15 intervenes or continues, and mean partial-credit Score across controller variants. Score is on a 0–100 scale.
Astra on Call overview: control allocation, selective intervention traces, and RoboDojo Scores
Astra on Call overview: control allocation, selective intervention traces, and RoboDojo Scores

The full paper teaser: how control is shared, when Fixed-15 intervenes or continues, and mean partial-credit Score across controller variants. Score is on a 0–100 scale.

02 / THE TAKEAWAYFixed-15 on RoboDojo

Help where it matters.
Keep what works.

Better assistance means recovering mistakes while preserving useful behavior. Selective takeover does both on the evaluated tasks.

18.1 42.9

Better on difficult tasks

Weak-set Score · 35 episodes
Policy only → selective takeover

76.7 76.7

Strong-set performance retained

Aggregate Score · 15 episodes
Policy only → selective takeover

−27.6%

Less time per episode

Weak set · versus full GPT-6 control
Using the paper’s timing accounting

How should control be shared?

7 tasks · 35 episodes

Mean Score050100
Policy onlyBase-25
18.1
GPT-6 onlyAlways-Segment
37.7
Astra on CallFixed-15
42.9

Score = mean partial credit, from 0 to 100. It is not success rate.

Successful episodesBase-252 / 35Always-Segment6 / 35Fixed-158 / 35

Timing: 1,269 s vs 1,753 s per weak-set episode, accounting for 16 s per model call, 0.3 s per control step, and 3 s per consultation. These figures describe the study’s execution setup.

FINDING / 01

Recover a failure.
Protect a success.

The biggest advantage over full GPT-6 control comes from keeping more of the policy’s original successes.

11 / 13prior successes preserved
8 / 37prior failures recovered

Fixed-15 across the matched weak and strong sets. Full GPT-6 control preserves 2/13 and recovers 7/37.

RESULT FIGURESelective takeover balances recovery of policy failures with preservation of policy successes.
Recovery and preservation across controller variants
Recovery and preservation across controller variants

Selective takeover balances recovery of policy failures with preservation of policy successes.

03 / HOW IT WORKSFrozen π₀.₅ policy · Inspect-EEF harness

The policy acts.
Astra stays on call.

A scene, a proposed motion, and a choice. The model inspects what the policy intends to do before deciding whether to help. No policy retraining.

01

Propose

The frozen policy proposes a motion chunk. Its end-effector waypoints make the planned movement inspectable.

π₀.₅ → proposed actions
02

Inspect

GPT-6 considers the scene, task, execution feedback, and proposed motion. Fixed-15 consults every 15 policy steps.

Scene + plan + feedback
03

Continue or correct

Keep the policy in control, take over an action segment, or—in the editing variants—adjust the proposed waypoints.

Then observe and repeat ↺
METHOD FIGUREController variants change when to consult, whether continuation is allowed, and how corrections are executed. The model can also terminate an episode with give_up.
Inspect-EEF consultation and execution loop
Inspect-EEF consultation and execution loop

Controller variants change when to consult, whether continuation is allowed, and how corrections are executed. The model can also terminate an episode with give_up.

RESULT FIGURETakeover and EditOnly variants share consultation schedules but differ in correction representation and execution.
Segment takeover compared with bounded waypoint editing
Segment takeover compared with bounded waypoint editing

Takeover and EditOnly variants share consultation schedules but differ in correction representation and execution.

FINDING / 02

Small edits are
not always enough.

Bounded waypoint editing lowers weak-set Score at three of four evaluated Fixed-K intervals. At K = 15, segment takeover scores 42.9, compared with 25.1 for EditOnly.

The correction interface matters: a successful local adjustment in one rollout does not imply a better controller overall.

3 of 4 intervals favor segment takeover.
04 / THE TEST BED10 tasks · 5 layouts per task

Different tasks.
The same question.

RoboDojo manipulation tasks probe both recovery and preservation: seven tasks where the baseline struggles, and three where it already performs well.

BENCHMARK FIGUREHead-camera views of the evaluated task families. Weak and strong sets are defined from Base-25 outcomes on the evaluated layouts.
Ten RoboDojo task families with weak-set and strong-set examples
Ten RoboDojo task families with weak-set and strong-set examples

Head-camera views of the evaluated task families. Weak and strong sets are defined from Base-25 outcomes on the evaluated layouts.

These are results from simulated manipulation with a frozen π₀.₅ policy. Fixed-15 is an observed strong setting on this evaluation, and aggregate preservation does not mean that every individual success is retained.