Residual Modeling for Regression Policies
Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning
We revisit the gap between MSE-Policies and Flow-Policies from the perspective of statistical modeling: how action-prediction residuals shape policy optimization. Our human demonstration data analysis reveals state-dependent residual scales and heavier-than-Gaussian tails. We introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails. HT-Policies predict action chunks with a single feed-forward pass and achieve success rates competitive with generative policy baselines across four simulation benchmarks and real-robot evaluations, despite being faster in training and inference.
Why Residual Modeling? Insights from Data and Policies
Residual RMS · task average
Illustration · Scale mixture vs Gaussian with the same variance
Residual fit correlates with policy success
The scale of policy residuals varies across different stages of a task. Gaussian residuals with different scales are combined together during training. Combining Gaussians with different scales naturally produces a heavy-tailed distribution. Across different objectives, likelihood correlates with success rate.
HT-Loss: Modeling the Heavy-Tailed Residuals
Input-dependent scales
The model learns a residual scale for each observation.
Robust to heavy tails
The gradient weight decreases as the standardized residual grows.
HT-Policies: Performance and Efficiency Across Tasks
Real world and simulation demos
Peel Note
HT-Policy 52%
vs Flow-Policy −4 pp
Push T
HT-Policy 88%
vs Flow-Policy +2 pp
Store Bottle
HT-Policy 58.3%
vs Flow-Policy +16.3 pp
Tool Hang
HT-Policy 97%
vs DP −3 pp
Understanding HT-Policies Through Gradients
Citation
@misc{zhou2026residualmodelingclosesregression,
title={Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning},
author={Yuchen Zhou and Jiacheng You and Weikang Wan and Weijun Dong and Yang Gao and Jiayuan Mao},
year={2026},
eprint={2610.12231},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2610.12231},
}