When a humanoid cannot avoid a fall, the goal is to reduce mechanical damage on impact and still stand up afterward. Ukemi-SafeFall learns this behavior from martial-arts ukemi (breakfall) techniques, with a Score-Matching Motion Prior to keep motions natural.
Simulation surrogates on the fallen-state protocol (N = 50). The three percentage reductions are relative to Passive Fall; time-to-stand and recovery success rate are absolute Ukemi-SafeFall results.
Humanoid robots remain vulnerable to falls when balance recovery fails. Controllers that treat falling as terminal failure leave an open safety problem: how to reduce mechanical damage during unavoidable falls while preserving the option to stand up afterward.
This paper proposes Ukemi-SafeFall, an ukemi-inspired, injury-aware falling and recovery framework that shapes humanoid fall–recovery behavior through condition-gated rolling and impact-mitigation rewards. In simulation, Ukemi-SafeFall reduces injury-risk surrogates of normalized peak contact force, cumulative peak contact force, and unsafe-contact events by 32%, 78%, and 70%, respectively, compared with passive falling. Experiments further demonstrate that Ukemi-SafeFall transfers to a humanoid platform across a range of falling scenarios.
On a fallen-state evaluation protocol (50 simulated episodes per method), Ukemi-SafeFall reaches a 98% recovery success rate with time-to-stand 1.04±0.27 s. The same protocol is used for matched ablations and baselines described below.
Rewards activate from torso tilt and head height to encourage roll initiation, roll continuation, soft landing, and standing recovery, without a discrete phase estimator.
Training limits trunk and arm loading during contact so normalized peak contact force, cumulative peak contact force, and unsafe-contact events are reduced relative to uncontrolled falling.
Recovery rewards encourage upright posture and head-height progress. A frozen Score-Matching Motion Prior multiplies the task reward to keep motions consistent with human reference behavior.
We report injury-risk surrogates (simulation contact quantities, not direct hardware damage) and recovery surrogates under the fallen-state protocol. Names match the paper definitions.
We train a reinforcement learning policy with Proximal Policy Optimization for humanoid falling and recovery. The policy maps partial observations to target joint offsets. Training uses a multiplicative reward that combines condition-gated task rewards with a diffusion-based Score-Matching Motion Prior reward. Generative State Initialization samples diverse fallen starting poses from the same prior, and domain randomization varies physical and sensing parameters so the policy sees a wide range of conditions.
Using the same fallen-state evaluation protocol, we evaluate Ukemi-SafeFall under three ablations: Ukemi-SafeFall: Recovery Only, Ukemi-SafeFall without condition gating, and Ukemi-SafeFall without Score-Matching Motion Prior reward gating. We also compare against Passive Fall and, as a cross-protocol reference, SafeFall Triangle Proximal Policy Optimization.
Table. Ablation and comparison methods.
| Method | Ukemi task reward | Impact-mitigation | Score-Matching Motion Prior reward |
|---|---|---|---|
| Passive Fall | × | × | × |
| Ukemi-SafeFall: Recovery Only | × | × | ✓ |
| Ukemi-SafeFall without Condition Gating | Always active | ✓ | ✓ |
| Ukemi-SafeFall without Score-Matching Motion Prior Reward Gating | ✓ | ✓ | × |
| SafeFall Triangle Proximal Policy Optimization (*) | × | ✓ | × |
| Ukemi-SafeFall (Ours) | ✓ | ✓ | ✓ |
(*) Evaluated and reported as a cross-protocol reference
Illustrative simulation rollouts for each comparison method and ablation.
The controller outputs zero action after Generative State Initialization and is subjected to the evaluation disturbances. This baseline characterizes uncontrolled falling.
This ablation uses the same multiplicative Score-Matching Motion Prior reward gate and Generative State Initialization procedure as Ukemi-SafeFall, but it is trained only with the upward-motion and head-height recovery rewards.
This ablation retains the proposed ukemi-inspired rewards, contact-force shaping, and Score-Matching Motion Prior gate, but disables the activation conditions; the corresponding reward terms remain active throughout the episode.
This ablation retains the condition-gated task rewards and contact-force shaping, but removes the multiplicative Score-Matching Motion Prior reward gate. Generative State Initialization continues to sample initial states from the Score-Matching Motion Prior.
We re-implement a SafeFall method to show how a protective falling policy performs without an explicit recovery objective. It is adapted to the same humanoid but uses a different training protocol; therefore, we report it as a cross-protocol reference rather than a matched ablation.
The complete method combines condition-gated ukemi-inspired rewards, contact-force shaping, and multiplicative Score-Matching Motion Prior reward gating.
We validate Ukemi-SafeFall across falling scenarios including locomotion over obstacles, stair descent, and platform jumps. In a separate MuJoCo-based sim-to-sim environment, the controller exhibits rolling, distributed limb contact, and recovery toward standing. Hardware demonstrations are coming soon.
Sim-to-sim results: (a) walking and (b) running over 0.08 m-high obstacles, (c) descending stairs, and (d) jumping from a 0.6 m-high platform.
@inproceedings{anonymous2026ukemisafefall,
title = {Ukemi-SafeFall: An Ukemi-inspired Injury-Aware Falling and Recovery},
author = {Anonymous Authors},
booktitle = {Under Review},
year = {2026}
}