Humanoid motion tracking · RA-L submission
Conditional Adversarial Differential Discriminator for Whole-Body Humanoid Imitation
CADD learns the tracking reward with a discriminator that sees both the tracking error and the target reference pose, so the reward adapts to each motion. The same recipe trains Unitree G1, Booster T1, AgiBot X2 and EngineAI PM01 without per-robot reward tuning.
Narrated summary of the method, the reward landscape, simulation results and real-robot deployment.
Real-robot clips: Unitree G1 and Booster T1; the operator is masked.
The discriminator scores the tracking error Δ against a zero-error target, conditioned on the reference pose c. The action rate is part of Δ, so smoothness is learned without a separate reward.
Final and intermediate models of single-motion policies. Each error point is a scaled copy of the typical tracking error of that motion's CADD policy, measured as mean body-position error. Drag the frame slider: ADD and BeyondMimic give the same curve at every frame, CADD does not. Drag the training slider to see the learned rewards form.
Error axis on a log scale, from 0.3 cm to where BeyondMimic falls to 3% of its peak; dotted line: the policy's median error. Shaded: the straight-line perturbation exceeds the range of at least one joint. A learned-reward curve stops where it starts rising again with larger error: the discriminators never saw such errors and their output there is extrapolation. Each reward is divided by its maximum over the drawn range.
Success rate [%], mean over 5 seeds. One policy per dataset or robot.
| Dataset (G1) | DeepMimic | BeyondMimic | ADD | CADD |
|---|---|---|---|---|
| LAFAN1 | 79.3 | 89.9 | 89.7 | 97.4 |
| AMASS | 76.3 | 82.1 | 85.2 | 91.6 |
| OmniXtreme | 59.3 | 86.9 | 81.5 | 97.2 |
| Robot (LAFAN1) | DeepMimic | BeyondMimic | ADD | CADD |
|---|---|---|---|---|
| Unitree G1 | 79.3 | 89.9 | 89.7 | 97.4 |
| AgiBot X2 | 84.8 | 84.9 | 90.1 | 98.3 |
| EngineAI PM01 | 75.3 | 79.0 | 78.3 | 90.2 |
| Booster T1 | 70.4 | 78.2 | 85.4 | 89.1 |
One CADD policy trained on all of LAFAN1, simulated offline in MuJoCo and replayed here. Pick a motion and a push; the green ghost is the reference motion, shown beside the robot. Booster T1, EngineAI PM01 and AgiBot X2 are being retrained with the current recipe.
Robot
LAFAN1 motion
Push at 5 s
Reference
Loading…