Conditional Adversarial Differential Discriminator for Whole-Body Humanoid Imitation
Narrated summary of the method, the reward landscape, simulation results and real-robot deployment.
Real-robot clips: Unitree G1 and Booster T1; the operator is masked.
Final and intermediate models of single-motion policies. Each error point is a scaled copy of the typical tracking error of that motion's CADD policy, measured as mean body-position error. Drag the frame slider: ADD and BeyondMimic-reward give the same curve at every frame, CADD does not. Drag the training slider to see the learned rewards form.
Error axis on a log scale, from 0.3 cm to where BeyondMimic-reward falls to 3% of its peak; dotted line: the policy's median error. Shaded: the straight-line perturbation exceeds the range of at least one joint. A learned-reward curve stops where it starts rising again with larger error: the discriminators never saw such errors and their output there is extrapolation. Each reward is divided by its maximum over the drawn range.
One CADD policy trained on all of LAFAN1, simulated offline in MuJoCo and replayed here. Pick a motion and a push; the green ghost is the reference motion, shown beside the robot. Booster T1, EngineAI PM01 and AgiBot X2 are being retrained with the current recipe.
Robot
LAFAN1 motion
Push at 5 s
Reference
Loading…