Narrated summary of the method, the reward landscape, simulation results and real-robot deployment.
Real-robot clips: Unitree G1 and Booster T1; the operator is masked.
Final and intermediate models of single-motion policies. Each error point is a scaled copy of the typical tracking error of that motion's CADD policy, measured as mean body-position error. Press Play or drag the frame slider: the ADD and BeyondMimic-reward curves stay fixed along the motion, while the CADD curve changes from frame to frame. Press the training Play button or drag its slider to see the learned rewards form.
Error axis on a log scale, from 0.3 cm to where BeyondMimic-reward falls to 3% of its peak; dotted line: the policy's median error. Learned-reward curves (ADD, CADD) end where the straight-line perturbation first pushes a joint past its range, or earlier where, after falling below half of their peak, they rise again by more than 0.02 with larger error: the discriminators never saw such errors and their output there is extrapolation. Each reward is divided by its maximum over the drawn range.