microduck-chimney-climb

A MicroDuck policy that climbs a narrow vertical corridor by bracing between two walls. The climbing policy is the main model; companion enter and exit policies are included to connect the climb to the floor and a top platform.

Status: simulation-only research release; not validated on hardware. The video is a selected successful enter–climb–exit–get-up sequence, not a robustness benchmark or an autonomous perception demonstration.

Open the final MP4. The ending uses Pollen Robotics' unchanged official alpha_stand.onnx recovery policy, not a fourth policy trained here.

Policies

Role Files Selected checkpoint
Main: climb policy.onnx, checkpoint.pt w17; stored iteration99,250
Enter the corridor enter/policy.onnx, enter/checkpoint.pt a11; stored iteration25,500
Exit onto the platform exit/policy.onnx, exit/checkpoint.pt x19; stored iteration73,750

Checkpoints retain actor, critic, optimizer, and normalizer state. Stored iteration counters include continuation history; they are not standalone training-budget claims. Load PyTorch checkpoints only from trusted revisions.

Simulation and source

Source: Vottivott/microduck-playground@c594e4457983ed1c1c36193e3df60e16de66b819.

git clone https://github.com/Vottivott/microduck-playground.git
cd microduck-playground
git checkout c594e4457983ed1c1c36193e3df60e16de66b819
uv sync
hf download HannesVonEssen/microduck-chimney-climb --local-dir policies/chimney

See USAGE.md for the continuous-chain command, exact handoff conventions, rendering, export and verification. manifest.json pins the official recovery dependency and its SHA256. It is fetched separately with scripts/fetch_chimney_getup.py; the three released policies do not include or modify its weights.

Contract

  • Input obs: float32 [1,61]; output actions: float32 [1,14];50Hz.
  • Layout: angular velocity(3), projected gravity(3), relative joint positions(14), joint velocities(14), previous bounded actions(14), twist(3), head(4), body(6).
  • Normalizer and per-joint output bounds baked into each ONNX via scripts/export.py.
  • Action scale1.0; target = policy-specific offset + output. Climb/exit use the brace offset, not HOME. Enter uses the standing offset.
  • No action low-pass filter. Joint order, exact offsets and bounds are in the corresponding config.json at root, enter/, and exit/.

Twist and head command slots are zero. Body command slots are:

Policy Six body-command values
Climb [(width_m-.13)/.02, 0, 0, 0, 0, 0]
Enter [dx_body, dy_body, sin(yaw), cos(yaw), (width_m-.13)/.02, 0]
Exit [goal_y-root_y, goal_z-root_z, goal_x-root_x, 0, (width_m-.13)/.02, 0]

Enter needs corridor-relative localization; exit needs platform-relative position. Climb needs a measured corridor width, not vision. The video uses a mirrored exit adapter to leave toward negative y while preserving the robot's heading; this reflection is not baked into exit/policy.onnx. Same-shaped inputs do not make these files drop-in walking-policy replacements.

Demonstrated geometry and evaluation

The video uses a12.5cm corridor gap and a platform3m above the floor. Platform and walls are6cm thick. The shelf is50cm wide, overlaps the upper walls by45cm, and extends another40cm into open space. Walls end40cm above the platform. Geometry is parametric in the accompanying source.

Selected seed29 switched enter→climb at1.74s, climb→exit at25.08s, and exit→official get-up at27.80s. In the original36s simulation, the final3s had both feet contacting the platform, tilt≤0.725degrees, and root speed≤0.00632m/s. The original trajectory and contact audit are included under eval/. Rendering replays that continuous trajectory; no physical resets or teleport to standing occur during the demonstrated chain.

Two of four exploratory seeds completed the chain. A fresh seed29 run from the prepared source reached the exit but did not complete recovery; the full-chain behavior is sensitive and is not reliably reproducible from the seed alone. The saved trajectory makes the selected demonstration inspectable, but is not an independent fresh-policy success. No broader robustness claim is made.

Each ONNX was checked against its checkpoint with100 zero/random and100 live observations. Maximum bounded-action discrepancy was below4.1e-5; all300 live policy steps across the three tasks remained finite. See eval/checkpoint-onnx-parity.json. The source's76 chimney/enter/exit regression tests passed.

Sim-to-real boundary

Training retains BAM XL330 actuator dynamics, control/sensor delays, encoder and IMU noise, mass/friction/CoM randomization, bounded joint commands, and augmented body collision hulls. These are ordinary non-backlash policies. Neither the complete chain nor any individual released skill is hardware validated. Real wall friction, compliance, collision geometry tolerances, servo thermal/current loading, localization and handoff reliability remain open issues. Simulation wall contact is not evidence of a safe real3m climb.

For future hardware research, begin low with a catch/support rig and an independent torque-off path. Do not start by reproducing the3m demonstration.

Files

  • Root policy.onnx, checkpoint.pt, config.json: main climbing policy.
  • enter/, exit/: companion ONNX, checkpoint and exact contracts.
  • media/preview.mp4: final user-provided video, unchanged.
  • manifest.json: roles, source provenance, geometry, dependency and hashes.
  • USAGE.md: simulation and handoff instructions.
  • eval/: parity checks, original trajectory/audit, fresh-chain result.
  • SHA256SUMS: artifact integrity hashes.

Continue training

See TRAINING.md for a tested public continuation workflow, including the recovered 645-state handover bank, explicit wall-friction randomization (0.5–1.1; optional 0.4–1.1), and resume commands. No private archive is needed. A fresh 64-environment/five-iteration smoke test restored model/normalizer/optimizer/counter state and passed; 78 regression tests passed. The recipe is reconstructed from a preceding launcher, not an exact recovered w17 launch configuration. Original release checkpoints, policies and video are unchanged.

Enter and exit continuation

COMPANION_TRAINING.md provides tested full-checkpoint resume commands for enter and exit, plus an explicitly optional recovered 18-state arrival bank. Each path passed a 64-environment/five-update resume and normalized ONNX parity check. x19 originally used procedural starts, not this bank. These are documented continuations on the published source, not exact historical replay. Original policies, checkpoints and media are unchanged.

Architecture graph for HannesVonEssen/microduck-chimney-climb. Open in hfviewer
Downloads last month
45
Video Preview
loading

Space using HannesVonEssen/microduck-chimney-climb 1

Collection including HannesVonEssen/microduck-chimney-climb