P-BFMPerceptive Behavioral Foundation Modelsfor Humanoid Control via Unsupervised Reinforcement Learning

A single humanoid policy that adapts its motion to the terrain it sees.

RoboPartyLab

The idea

Same intent.
Different terrain.

P-BFM separates what behavior to perform from how to execute it. Depth perception guides terrain adaptation; explicit heading and velocity instructions guide root motion.

One frozen policy supports motion tracking, goal reaching, reward inference, and teleoperation on a Unitree G1.

How P-BFM works

Terrain adaptation

One walking intent, varied terrain.

Changes in gait and footholds emerge without terrain-specific motion demonstrations.

Walking reference and execution upstairs, downstairs, uphill, downhill, on discrete blocks and rough terrain.
Same forward-walking intent across terrains. Figure 3.Swipe to exploreOpen full-size figure

01 / Demonstrations

Input A target body state.

Watch how the robot approaches that state on flat ground, slopes, and stairs.

Show remaining 3 videosShow fewer videos

02 / Demonstrations

Input A reference motion.

Watch walking, dancing, and cartwheel motions adapted to the terrain.

Show remaining 3 videosShow fewer videos

03 / Demonstrations

Reward inference

Input A reward function for forward speed and pelvis-to-foot height.

Watch the resulting walking behavior.

Velocity target
1 m/s
Height target
0.4 m

Reward targets, not measured outcomes.

04 / Demonstrations

Input Live human motion through online retargeting.

Watch the robot follow the operator’s movements on stairs and platforms.

Show remaining 3 videosShow fewer videos

Method

From behavioral intent to physical action.

01

Encode the intent

A forward–backward representation maps motion references, target states, and rewards into a shared behavioral space.

02

Perceive the terrain

The actor uses depth observations to adapt whole-body execution to its surroundings.

03

Guide root motion

Explicit Root Instructions specify heading and planar velocity alongside the behavioral latent.

P-BFM pretraining architecture and frozen inference with teleoperation, reference motions, goals, and rewards.
Pretraining and zero-shot inference. Figure 2.Swipe to exploreOpen full-size figure