AbstractReinforcement learning-based motion-tracking has become one of the most commonly used paradigms for creating controllers that enable humanoid robots to perform a wide range of motor skills. One of the key components of these tracking-based methods is the use of manually-designed tracking objectives that evaluate the difference between the reference motion and the motions produced by the agent. Such pre-defined, fixed objectives are unable to dynamically adapt to the characteristics of different motions or the changing behaviors of the controller over the course of training. This rigidity often limits tracking accuracy, particularly when imitating large diverse motion datasets. Moreover, the reward parameters typically require re-tuning for each new robot. This work introduces the Conditional Adversarial Differential Discriminator (CADD), which learns an adaptive tracking reward function using a differential discriminator conditioned on the target reference motion. This allows the discriminator to automatically adapt the reward function based on the characteristics of different target motions, thus enabling scalable motion tracking across diverse behaviors. CADD trains general motion tracking controllers that closely replicate a diverse range of behaviors on real humanoid robots. The same training recipe applies across distinct humanoid platforms, maintaining high tracking accuracy without manual reward-weight tuning or motion- and robot-specific reward engineering. CADD provides a scalable drop-in alternative to handcrafted tracking reward functions in most motion imitation frameworks.
Adaptive to each frame. Compared with ADD, CADD tracks up to 60% more accurately and lifts the success rate from 82–90% to 92–97% on three large motion datasets.
One reward for all. One learned reward, with no handcrafted terms, drives agile skills on real humanoids (94.6% success on the G1, 90.0% on the T1) and, in simulation, trains one policy on large datasets of highly dynamic motions. To our knowledge, no prior work trains agile real-humanoid skills with a single reward term.
Scalable. The same training setup, with no per-robot reward tuning, tracks most accurately on four different humanoids, and fall recovery emerges on its own from large motion datasets.