RoboIRS: Inference-Time Intervention on Robot Policies
with Internal Representation Steering Vectors

Anonymous Authors

Abstract

Vision-language-action (VLA) and world-action models (WAMs) often degrade under out-of-distribution task variations despite retaining partial task capability. To recover such capability, we propose RoboIRS, an inference-time internal representation steering method that uses successful and failed rollouts to train linear classifiers, select outcome-relevant intervention locations, and derive task-specific steering directions without updating policy parameters. On 15 simulation tasks with a frozen π0.5 policy, RoboIRS improves the average success rate from 44.4% to 66.2%, outperforming alternative inference-time intervention baselines while adding little inference time. We further validate RoboIRS on real-robot manipulation using the same π0.5 policy and demonstrate its applicability to a world-action model Cosmos Policy, where the average success rate improves from 35.4% to 55.4%. These results show that directly steering internal robot-policy representations can improve the performance of robot policies at inference time.

Motivation

Failure under shift is not incapability

VLA and world-action policies degrade on out-of-distribution inputs — unseen scenes, novel object layouts, new appearances. But degradation is rarely total: the same policy still succeeds on a fraction of those rollouts. That partial competence suggests the internal computation required for success is still present, just not reliably expressed.

The same story as LLMs

Language models produce both truthful and untruthful answers to near-identical prompts, and mechanistic interpretability has shown those behaviors live along linear directions in the residual stream that can be causally manipulated. Inference-time intervention, contrastive activation addition, and representation engineering all exploit this.

The question

Can we steer the internal representations of a frozen robot policy at inference time to recover task performance under distribution shift?

Prior robotics work localizes interpretable features or steers a few predefined behavioral attributes (motion speed, end-effector height), or requires carefully matched nominal/perturbed observation pairs. RoboIRS instead derives its intervention directly from outcome-labeled rollouts — what succeeded and what failed — and targets general task success.

Real-robot setup under visual distribution shift: reference scene vs. unseen scene, and the same frozen policy with and without RoboIRS steering.
Fig. 1. A π0.5 policy fine-tuned on a clean background is evaluated under visual distribution shift. Without steering it deviates from the target; with RoboIRS the same frozen policy grasps successfully.

Method

RoboIRS intervenes between transformer blocks of a frozen policy. The steering direction comes from a linear classifier trained to separate success-like from failure-like residual streams; the intervention site is chosen by that classifier's validation AUC.

RoboIRS framework: (a) a scaled steering vector is added to the selected residual stream between transformer blocks of the frozen policy; (b) residual streams from successful and failed rollouts train a linear classifier whose weight is the steering direction.
Fig. 2. Overview of RoboIRS. (a) A scaled steering vector is added to the selected residual-stream representation between transformer blocks, without updating any policy parameter. (b) Residual streams from successful and failed rollouts train a linear classifier; the classifier weight provides the steering direction injected at inference.
1

Collect outcome-labeled residuals

Roll out the frozen policy and record the residual stream Rℓ,t,n at every timestep t, layer ℓ, and denoising step n. For π0.5 a 30-step rollout yields a (30, 18, 10, 10, 1024) tensor. Up to 10 successful and 10 failed rollouts per task — that is the entire data cost.

2

Fit a linear classifier → steering direction

Per layer and action-token position, train P(y=1 | x) = σ(w⊤x + b) where y=1 marks failure. Successful rollouts label all timesteps 0; a failed rollout labels timesteps 1 from the step its end-effector leaves the space swept by successful trajectories. The weight w is the steering vector. Average validation AUC: 0.88; all classifiers train in under three minutes on CPU.

3

Select the site by AUC, then inject

Exhaustive sweeps over intervention sites are prohibitive in robotics — each candidate needs multi-step closed-loop rollouts. Instead, within each layer pick the action-token position with the highest held-out failure-detection AUC, and steer only at the final denoising step:

rsteered = r + sα ‖r‖‖v‖ v,    s ∈ {−1, +1}

α sets steering strength as a fraction of the residual norm, which keeps the intervention scale-consistent across layers. The sign s is chosen empirically: we read v as a task-relevant latent factor that can be under- or over-activated, not as a fixed failure→success arrow.

Qualitative Results — Simulation

Each row shows the same π0.5 policy three times. First on the original LIBERO task, where it is in-distribution and succeeds. Then on the LIBERO-PRO perturbation of that task, where it fails. Then on the same perturbation with RoboIRS steering enabled — same frozen weights, same initial state as the failure, only the residual stream is modified. Clips play at 2× speed.

L10 · object · position swap

“Pick up the alphabet soup and place it in the basket”

Under perturbation: 39.3% → 90.0%
✓No perturbation
Original LIBERO task, unsteered
✗Perturbed, unsteered
Grasps the butter instead
✓Perturbed + RoboIRS
Selects the correct object

After the object-position swap the unsteered policy consistently confuses alphabet soup with butter. This is a decision-making failure, not an execution failure — the category steering helps most (+24.9 points on average).

L2 · goal · scene change

“Put the bowl on the plate”

Under perturbation: 32.7% → 96.7%
✓No perturbation
Original LIBERO task, unsteered
✗Perturbed, unsteered
Inaccurate grasp, bowl slips
✓Perturbed + RoboIRS
Clean grasp and place

The largest single-task gain in the benchmark: an action-precision failure mode, where the policy repeatedly re-attempts an imprecise grasp until timeout.

L9 · object · scene change

“Pick up the milk and place it in the basket”

Under perturbation: 19.3% → 72.7%
✓No perturbation
Original LIBERO task, unsteered
✗Perturbed, unsteered
Never secures the milk carton
✓Perturbed + RoboIRS
Completes the placement

Re-ranking reaches only 32.0% here and CAA 16.0%, while RoboIRS reaches 72.7% — the classifier-derived direction carries information that candidate scoring alone does not recover.

L4 · goal · scene change

“Put the wine bottle on top of the cabinet”

Under perturbation: 17.3% → 68.0%
✓No perturbation
Original LIBERO task, unsteered
✗Perturbed, unsteered
Fails to lift the bottle onto the cabinet
✓Perturbed + RoboIRS
Places the bottle successfully

Steered with s = +1. The sign is task-specific: we read the steering vector as a latent factor that can be under- or over-activated, not as a fixed failure→success arrow.

Qualitative Results — Real Robot

A π0.5 policy fine-tuned on demonstrations collected against a clean background, then evaluated with printed paper sheets scattered across the workcell — a pure visual distribution shift. 50 trials per condition per task.

H1 · real robot

Disassemble the RAM from the motherboard

24.0% (12/50) → 44.0% (22/50)
✓Reference scene
In-distribution, no steering
✗Shifted scene, unsteered
Trajectory deviates
✓Shifted scene + RoboIRS
Successful grasp
Policy observation view (base + wrist camera)
✗ Unsteered rollout
✓ RoboIRS rollout
H2 · real robot

Pull out the USB connector

12.0% (6/50) → 32.0% (16/50)
✓Reference scene
In-distribution, no steering
✗Shifted scene, unsteered
Misses the connector
✓Shifted scene + RoboIRS
Successful extraction
Policy observation view (base + wrist camera)
✗ Unsteered rollout
✓ RoboIRS rollout

Q3 — Transfer to a World-Action Model

Cosmos Policy has a different architecture (DiT) and a different training recipe: no separate action expert, actions decoded by averaging a 196-token action frame generated by a video diffusion transformer. We steer all 196 tokens of the final transformer block, with the classifier trained on mean-pooled residuals. Macro-average success rises 35.4% → 55.4%; 6 of 8 tasks improve meaningfully.

C1 · goal · object change

“Put the bowl on top of the cabinet”

Under perturbation: 36.7% → 92.2%
✗Perturbed, unsteered
✓Perturbed + RoboIRS
C5 · spatial · position swap

“Pick up the black bowl next to the plate and place it on the plate”

Under perturbation: 62.2% → 97.8%
✗Perturbed, unsteered
✓Perturbed + RoboIRS

Both pairs use matched episode seeds. Unlike the π0.5 rows above, no unperturbed Cosmos Policy rollouts were available, so these comparisons show the perturbed conditions only.

Baseline Comparison

All numbers are π0.5 on the same 15 LIBERO-PRO tasks, 150 rollouts over 3 random seeds per cell. Every baseline keeps the policy frozen and intervenes only at inference time, and each one is given the same successful and failed rollout data RoboIRS uses, with its strength chosen by the same validation protocol.

Average success rate over 15 LIBERO-PRO tasks (%)
0 20 40 60 80 Gradient guidance Gradient guidance: 43.6% average success rate, 137.9 ms average inference time 43.6 Unsteered Unsteered: 44.4% average success rate, 134.0 ms average inference time 44.4 Re-ranking (K=16) Re-ranking (K=16): 58.6% average success rate, 1074.9 ms average inference time 58.6 CAA CAA: 62.6% average success rate, 136.3 ms average inference time 62.6 RoboIRS (ours) RoboIRS (ours): 66.2% average success rate, 136.4 ms average inference time 66.2 unsteered baseline

Re-ranking is the closest, at 8× the latency

Sampling K=16 action chunks and scoring them reaches 58.6%, but pushes average inference from 134 ms to 1075 ms. RoboIRS adds 2.4 ms. Re-ranking also wins on only 7 of 15 tasks and by an average of 9.7 points, while RoboIRS wins the other 8 by 22.7.

CAA gains most of its ground on a few tasks

Contrastive activation addition uses the mean success−failure difference at the same intervention sites, reaching 62.6%. It is the strongest baseline overall and beats RoboIRS outright on L4 and L5, but trails on average because its direction is not selected by a held-out criterion.

Gradient guidance barely moves the policy

Guiding the flow velocity with ∇a log P(success) lands at 43.6%, below the unsteered policy. The classifier leans almost entirely on the encoded state, so the action-space gradient is weak; an action-only classifier yields 43.5%.

IDSuiteTask UnsteeredGrad.Re-rank K=16CAARoboIRS
L1goal_envOpen top drawer; put bowl inside12.012.744.734.034.0
L2goal_envPut bowl on plate32.730.774.068.796.7
L3goal_envPut bowl on cabinet72.058.779.375.389.3
L4goal_envPut wine bottle on cabinet17.319.351.394.068.0
L5goal_swapPut bowl on plate70.063.370.785.366.7
L6object_envAlphabet soup to basket59.356.779.366.071.3
L7object_envBBQ sauce to basket80.783.398.795.395.3
L8object_envKetchup to basket84.780.092.082.084.7
L9object_envMilk to basket19.320.732.016.072.7
L10obj_swapAlphabet soup to basket39.343.331.375.390.0
L11obj_swapButter to basket32.034.721.359.343.3
L12spatial_envBowl between plate/ramekin to plate40.743.364.048.036.7
L13spatial_envBowl in top drawer to plate25.329.332.036.034.0
L14spatial_envBowl next to plate to plate58.758.068.075.376.7
L15spatial_envBowl next to ramekin to plate22.020.740.728.733.3
Average SR (%)44.443.658.662.666.2
Win margin, baseline over RoboIRS3.7 (2/15)6.6 (1/15)9.7 (7/15)14.8 (5/15)—
Win margin, RoboIRS over baseline27.8 (12/15)24.6 (14/15)22.7 (8/15)15.9 (8/15)—
Average inference time (ms)134.0137.91074.9136.3136.4

Highest value in each row is marked. Win margin rows report the average success-rate difference over the subset of tasks where that method wins, with the number of tasks won in parentheses — so a method can win more often yet by a smaller margin.

Video Overview

Limitations

BibTeX

Citation information will be posted once the paper is published.

@inproceedings{TBD,
  title     = {TBD},
  author    = {TBD},
  booktitle = {TBD},
  year      = {TBD}
}