World grounding vs. behavior grounding

Written by Gilwoo Lee, PhD · Based on Getting Out and Getting Back by Samuel Liu et al.

A robot picks up an object perfectly in simulation. We deploy it on the real robot, and it fails. Or it succeeds, but moves in a way we’d never want on a factory floor.

These are two different problems. One is about how well simulation matches reality. The other is about what behavior we want the robot to learn. We call them world grounding and behavior grounding.

World grounding

Align the simulator with reality.

This includes visual fidelity, physics, robot kinematics, and control. In theory, if simulation matches the real world in every respect, the Sim2Real gap should be zero: what works in simulation should work in reality. World grounding means aligning all aspects of simulation with reality as closely as possible.

Behavior grounding

Align sim behaviors with desired behaviors.

A motion planner can generate a successful grasp using a very different motion than a human operator would. Depending on the task, we may prefer specific behavioral traits over pure task performance. Behavior grounding refers to aligning simulated trajectories to match the desired behaviors.

What happens when we change them separately?

We varied world and behavior grounding independently on a conveyor-picking task. We generated simulated demonstrations with and without world grounding, using either human-like or motion-planned behavior. Then we mixed each dataset with real demonstrations and evaluated the resulting policies on the physical robot.

Fully grounded co-training improved success from 52% to 86%. But watching the robot raised a more interesting question: which experience was it actually drawing on?

Sometimes the robot moves like the human. Sometimes it moves like the planner.

The policies often followed the real demonstrations in familiar situations. Outside that coverage, their movements looked more like the behaviors in the simulated data.

To make this easier to see, we trained a policy with just 10 real demonstrations and 1,500 ungrounded simulated trajectories. Watch the grasp below: the robot uses the two-finger pinch seen in the motion-planned data.

Top: Policy trained with 100 real demonstrations. Bottom: 10. With fewer real demonstrations, the robot uses behaviors unseen in human demos, like the two-finger pinch.

One possible explanation is that simulated demonstrations, whether close to real-world demonstrations or not, give the policy something to fall back on when real demonstrations run out. World grounding helps that experience transfer. Behavior grounding may help the robot find its way back to familiar real-world states.

That’s the idea behind “getting out” (to the simulation distribution) and “getting back” (to the real-world distribution). We see evidence for it in the rollouts and representation analysis.

See the real-world rollouts and representation analysis →

Does simulated motion have to be grounded in human behavior?

With better world grounding, a planner-generated motion should perform well even if it looks different from a human demonstration.

To test this hypothesis, we used video diffusion to make simulated camera views look more like real ones. Compared with the native Isaac RTX rendering, the simulated observations looked much closer to real observations, as you can see below.

Simulated camera views refined with video diffusion (top) and corresponding real camera views (bottom) for a reconstructed trajectory.

With 25 real and 75 simulated trajectories, human-like and planner-generated behavior both achieved 17 successes in 24 trials. For comparison, a policy trained on 100 real trajectories achieved 16.

In this experiment, making simulated behavior human-like no longer improved the observed success rate. That suggests better world grounding can reduce how much we need behavior grounding for task success.

But success isn’t always the whole objective. We may still want motions that are predictable, legible, or appropriate around people. Those preferences remain a reason to ground behavior.

See the visual-alignment experiment →

What we take from this

When simulated data doesn’t help, “make the simulation better” is too vague.

We should ask two separate questions: Does this world transfer? And does it teach the behavior we want? The answers can point to very different fixes.