microduck-chimney-climb
A MicroDuck policy that climbs a narrow vertical corridor by bracing between two walls. The climbing policy is the main model; companion enter and exit policies are included to connect the climb to the floor and a top platform.
Status: simulation-only research release; not validated on hardware. The video is a selected successful enter–climb–exit–get-up sequence, not a robustness benchmark or an autonomous perception demonstration.
Open the final MP4. The ending uses Pollen Robotics' unchanged
official alpha_stand.onnx recovery policy, not a fourth policy trained here.
Policies
| Role | Files | Selected checkpoint |
|---|---|---|
| Main: climb | policy.onnx, checkpoint.pt |
w17; stored iteration99,250 |
| Enter the corridor | enter/policy.onnx, enter/checkpoint.pt |
a11; stored iteration25,500 |
| Exit onto the platform | exit/policy.onnx, exit/checkpoint.pt |
x19; stored iteration73,750 |
Checkpoints retain actor, critic, optimizer, and normalizer state. Stored iteration counters include continuation history; they are not standalone training-budget claims. Load PyTorch checkpoints only from trusted revisions.
Simulation and source
Source:
Vottivott/microduck-playground@c594e4457983ed1c1c36193e3df60e16de66b819.
git clone https://github.com/Vottivott/microduck-playground.git
cd microduck-playground
git checkout c594e4457983ed1c1c36193e3df60e16de66b819
uv sync
hf download HannesVonEssen/microduck-chimney-climb --local-dir policies/chimney
See USAGE.md for the continuous-chain command, exact handoff
conventions, rendering, export and verification. manifest.json
pins the official recovery dependency and its SHA256. It is fetched separately
with scripts/fetch_chimney_getup.py; the three released policies do not include
or modify its weights.
Contract
- Input
obs: float32[1,61]; outputactions: float32[1,14];50Hz. - Layout: angular velocity(3), projected gravity(3), relative joint positions(14), joint velocities(14), previous bounded actions(14), twist(3), head(4), body(6).
- Normalizer and per-joint output bounds baked into each ONNX via
scripts/export.py. - Action scale1.0; target = policy-specific offset + output. Climb/exit use the brace offset, not HOME. Enter uses the standing offset.
- No action low-pass filter. Joint order, exact offsets and bounds are in the
corresponding
config.jsonat root,enter/, andexit/.
Twist and head command slots are zero. Body command slots are:
| Policy | Six body-command values |
|---|---|
| Climb | [(width_m-.13)/.02, 0, 0, 0, 0, 0] |
| Enter | [dx_body, dy_body, sin(yaw), cos(yaw), (width_m-.13)/.02, 0] |
| Exit | [goal_y-root_y, goal_z-root_z, goal_x-root_x, 0, (width_m-.13)/.02, 0] |
Enter needs corridor-relative localization; exit needs platform-relative
position. Climb needs a measured corridor width, not vision. The video uses
a mirrored exit adapter to leave toward negative y while preserving the
robot's heading; this reflection is not baked into exit/policy.onnx.
Same-shaped inputs do not make these files drop-in walking-policy replacements.
Demonstrated geometry and evaluation
The video uses a12.5cm corridor gap and a platform3m above the floor. Platform and walls are6cm thick. The shelf is50cm wide, overlaps the upper walls by45cm, and extends another40cm into open space. Walls end40cm above the platform. Geometry is parametric in the accompanying source.
Selected seed29 switched enter→climb at1.74s, climb→exit at25.08s, and
exit→official get-up at27.80s. In the original36s simulation, the final3s had
both feet contacting the platform, tilt≤0.725degrees, and root speed≤0.00632m/s.
The original trajectory and contact audit are included under eval/.
Rendering replays that continuous trajectory; no physical resets or teleport
to standing occur during the demonstrated chain.
Two of four exploratory seeds completed the chain. A fresh seed29 run from the prepared source reached the exit but did not complete recovery; the full-chain behavior is sensitive and is not reliably reproducible from the seed alone. The saved trajectory makes the selected demonstration inspectable, but is not an independent fresh-policy success. No broader robustness claim is made.
Each ONNX was checked against its checkpoint with100 zero/random and100 live
observations. Maximum bounded-action discrepancy was below4.1e-5; all300 live
policy steps across the three tasks remained finite. See
eval/checkpoint-onnx-parity.json.
The source's76 chimney/enter/exit regression tests passed.
Sim-to-real boundary
Training retains BAM XL330 actuator dynamics, control/sensor delays, encoder and IMU noise, mass/friction/CoM randomization, bounded joint commands, and augmented body collision hulls. These are ordinary non-backlash policies. Neither the complete chain nor any individual released skill is hardware validated. Real wall friction, compliance, collision geometry tolerances, servo thermal/current loading, localization and handoff reliability remain open issues. Simulation wall contact is not evidence of a safe real3m climb.
For future hardware research, begin low with a catch/support rig and an independent torque-off path. Do not start by reproducing the3m demonstration.
Files
- Root
policy.onnx,checkpoint.pt,config.json: main climbing policy. enter/,exit/: companion ONNX, checkpoint and exact contracts.media/preview.mp4: final user-provided video, unchanged.manifest.json: roles, source provenance, geometry, dependency and hashes.USAGE.md: simulation and handoff instructions.eval/: parity checks, original trajectory/audit, fresh-chain result.SHA256SUMS: artifact integrity hashes.
Continue training
See TRAINING.md for a tested public continuation workflow, including the recovered 645-state handover bank, explicit wall-friction randomization (0.5–1.1; optional 0.4–1.1), and resume commands. No private archive is needed. A fresh 64-environment/five-iteration smoke test restored model/normalizer/optimizer/counter state and passed; 78 regression tests passed. The recipe is reconstructed from a preceding launcher, not an exact recovered w17 launch configuration. Original release checkpoints, policies and video are unchanged.
Enter and exit continuation
COMPANION_TRAINING.md provides tested full-checkpoint resume commands for enter and exit, plus an explicitly optional recovered 18-state arrival bank. Each path passed a 64-environment/five-update resume and normalized ONNX parity check. x19 originally used procedural starts, not this bank. These are documented continuations on the published source, not exact historical replay. Original policies, checkpoints and media are unchanged.
- Downloads last month
- 45