Residual Modeling for Regression Policies

Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning

  • 1 University of Pennsylvania
  • 2 Tsinghua University
  • 3 UC San Diego

* Equal contribution

Overview of residual modeling: heavy-tailed action residuals, the mismatch with Gaussian assumptions, and HT-Policy success rates compared with MSE-Policies and Flow-Policies or Diffusion Policy on GR1, SIMPLER, and RoboMimic.

We revisit the gap between MSE-Policies and Flow-Policies from the perspective of statistical modeling: how action-prediction residuals shape policy optimization. Our human demonstration data analysis reveals state-dependent residual scales and heavier-than-Gaussian tails. We introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails. HT-Policies predict action chunks with a single feed-forward pass and achieve success rates competitive with generative policy baselines across four simulation benchmarks and real-robot evaluations, despite being faster in training and inference.

Why Residual Modeling? Insights from Data and Policies

Sweep into Pile · Bridge

Example episodeAlign and grasp

Residual RMS · task average

Task-average residual RMS for Sweep into Pile varies with episode progress, from about 0.190 at 5% to 0.337 at 95%. The curve is measured across the task, not from the example video.
Illustration: combining zero-mean Gaussians with different scales produces heavier tails than a Gaussian with the same variance.

Illustration · Scale mixture vs Gaussian with the same variance

Residual fit correlates with policy success

Calibrated mean log-likelihood versus success rate for eight training objectives on RoboCasa-GR1, including MSE-Loss and HT-Loss. Spearman correlation is 0.86 and Pearson correlation is 0.83. HT-Policies achieve the highest success rate in this comparison.

The scale of policy residuals varies across different stages of a task. Gaussian residuals with different scales are combined together during training. Combining Gaussians with different scales naturally produces a heavy-tailed distribution. Across different objectives, likelihood correlates with success rate.

HT-Loss: Modeling the Heavy-Tailed Residuals

HT-Loss code
ℓHT(o,a)=d2log⁡σθ2(o)+ν+d2log⁡ ⁣(1+∥r∥22ν σθ2(o))\ell_{\mathrm{HT}}(o,a)=\frac{d}{2}\log \textcolor{#2b65ad}{\sigma_\theta^2(o)}+\textcolor{#168452}{\frac{\nu+d}{2}\log\!\left(1+\frac{\lVert r\rVert_2^2}{\nu\,\textcolor{#2b65ad}{\sigma_\theta^2(o)}}\right)}
ℓHT(o,a)=d2log⁡σθ2(o)+ν+d2log⁡ ⁣(1+∥r∥22ν σθ2(o))\begin{gathered}\ell_{\mathrm{HT}}(o,a)=\frac{d}{2}\log \textcolor{#2b65ad}{\sigma_\theta^2(o)}\\[0.5em]+\textcolor{#168452}{\frac{\nu+d}{2}\log\!\left(1+\frac{\lVert r\rVert_2^2}{\nu\,\textcolor{#2b65ad}{\sigma_\theta^2(o)}}\right)}\end{gathered}
r=fθ(o)−ar=f_\theta(o)-a dd: action-chunk dimension ν\nu: degrees of freedom

Input-dependent scales

The model learns a residual scale for each observation.

Robust to heavy tails

The gradient weight decreases as the standardized residual grows.

HT-Policies: Performance and Efficiency Across Tasks

Benchmark performance

MSE-Policy Flow-Policy / DP HT-Policy (ours)
Success rates for MSE-Policies, Flow-Policies or Diffusion Policy, and HT-Policies. Simulation compares GR00T N1.7 on RoboCasa-GR1 and SimplerEnv, pi0.5 on LIBERO, Cosmos 3 on LIBERO-10, and policies trained from scratch on RoboMimic. Real-world tasks are Push-T, Peel-Note, and Insert-T. Every bar is labeled with its success rate.

Inference efficiency

NVIDIA RTX 5090 · batch size 1

Whole-model latency per action chunk on NVIDIA RTX 5090, batch size 1. Flow-Policy versus HT-Policy: GR00T N1.7, 46.20 versus 24.52 ms (1.88 times faster); pi0.5, 123.58 versus 67.94 ms (1.82 times faster); Cosmos3-Nano, 1271.76 versus 71.34 ms (17.83 times faster). For each model, Flow-Policy is normalized to 100%; lower is better. Timings include observation encoding and action prediction from preprocessed GPU inputs.

Real world and simulation demos

Peel Note

Real World

HT-Policy 52%

vs Flow-Policy −4 pp

Push T

Real World

HT-Policy 88%

vs Flow-Policy +2 pp

Store Bottle

Simulation

HT-Policy 58.3%

vs Flow-Policy +16.3 pp

Tool Hang

Simulation

HT-Policy 97%

vs DP −3 pp

Understanding HT-Policies Through Gradients

Similar residual tails, different gradient allocation

RoboCasa-GR1 and Tool-Hang comparisons: MSE-Policies and Flow-Policies exhibit heavy-tailed action residuals. Gradient contributions concentrate on high-residual observations under MSE-Loss and high-noise flow matching, while low-noise Flow-Policy gradients are more evenly distributed.

For MSE-Policies, large-residual observations dominate gradient allocation. Flow-Policies exhibit similarly heavy-tailed action residuals, but distribute gradients more evenly in the low-noise regime (t > 0.5).

More balanced gradients, higher policy success

Success rates above and gradient shares by action-residual decile below for RoboCasa-GR1, Bridge, Fractal, and Tool-Hang. HT-Policies improve success rates over MSE-Policies in all four settings and exhibit more balanced gradient allocation, resembling low-noise Flow-Policies.

Across four settings, HT-Policies improve success rates over MSE-Policies (top). With HT-Loss, large-residual observations no longer dominate gradient allocation (bottom); the resulting distribution resembles that of low-noise Flow-Policies (t > 0.5).

Citation

BibTeX
@misc{zhou2026residualmodelingclosesregression,
  title={Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning},
  author={Yuchen Zhou and Jiacheng You and Weikang Wan and Weijun Dong and Yang Gao and Jiayuan Mao},
  year={2026},
  eprint={2610.12231},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2610.12231},
}