ETH Zürich · 3D Vision course project

Real-to-Sim 6-DoF Object Pose

Depth-sensor-free 6-DoF object pose from a single egocentric RGB stream.

Team 22
ETH Zürich — 3D Vision

Real — tracked 6-DoF pose on egocentric RGB

🧊
Sim twin
loading interactive 3D preview…

Simulation — interactive placeholder for the Isaac Sim replay
final replay: 326 frames @ 50 fps (6.52 s) for an equal side-by-side cut

Tracked CAD pose on egocentric RGB — grasp → carry → drop → settle. No depth sensor.


Abstract

Replaying a teleoperated demo in simulation needs the object's per-frame 6-DoF trajectory, but we have only a single egocentric RGB stream — no depth sensor — plus the robot/gripper joint log. The objects are near-symmetric and intermittently occluded, so we anchor one high-quality pose and propagate it with a model-based tracker, corrected by the joint kinematics and the object mask.


Method

Anchor once, then correct per phase from the most trustworthy signal — split at the ungrasp.

1

Front end

Object segmentation + scale-calibrated mono-depth. No sensor.

2

Base + anchor

Base 6D track; a refined frame-0 orientation fixes the ambiguous start.

3

In-hand

Position from the mask, depth from gripper kinematics, rotation from 6D refinement.

4

Drop → rest

Released & unoccluded: full 6D refinement; outliers dropped, smoothed.

5

Compose

Concatenate into the per-frame 4×4 trajectory for the sim replay.


Method progression

Each cue added in turn, on a teleoperated pick-and-place: the raw 6D track is jittery; a single refined frame-0 anchor + gripper kinematics stabilise the in-hand pose (no per-frame refinement); the final folds in windowed 6D refinement with light smoothing.

1 · base 6D + scaled depth
raw — jittery

2 · refined frame-0 + kinematics
single anchor, then rigid deltas

3 · final
+ 6D refinement + smoothing


Datasets

Egoverse duck

Teleoperated pick-and-place: grasp → carry → drop → settle. (Re-grasp out of the bowl is cut.)

Sequence 1

Sequence 2

Sequence 3

Cross-dataset: pan generalization

Transfers to another camera and a near-circular pan by swapping only the segmentation prompt. For larger, less-occluded objects the kinematics step is skipped — pose fully from vision.

Pan 1

Pan 2

Pan 3


3D reconstructions coming soon

🧊

Isaac Sim replays and interactive 3D of the full reconstruction will land here — the sim-twin slot in the teaser already shows the interactive viewer format.