Tracked CAD pose on egocentric RGB — grasp → carry → drop → settle. No depth sensor.
Replaying a teleoperated demo in simulation needs the object's per-frame 6-DoF trajectory, but we have only a single egocentric RGB stream — no depth sensor — plus the robot/gripper joint log. The objects are near-symmetric and intermittently occluded, so we anchor one high-quality pose and propagate it with a model-based tracker, corrected by the joint kinematics and the object mask.
Anchor once, then correct per phase from the most trustworthy signal — split at the ungrasp.
Object segmentation + scale-calibrated mono-depth. No sensor.
Base 6D track; a refined frame-0 orientation fixes the ambiguous start.
Position from the mask, depth from gripper kinematics, rotation from 6D refinement.
Released & unoccluded: full 6D refinement; outliers dropped, smoothed.
Concatenate into the per-frame 4×4 trajectory for the sim replay.
Each cue added in turn, on a teleoperated pick-and-place: the raw 6D track is jittery; a single refined frame-0 anchor + gripper kinematics stabilise the in-hand pose (no per-frame refinement); the final folds in windowed 6D refinement with light smoothing.
1 · base 6D + scaled depth
raw — jittery
2 · refined frame-0 + kinematics
single anchor, then rigid deltas
3 · final
+ 6D refinement + smoothing
Teleoperated pick-and-place: grasp → carry → drop → settle. (Re-grasp out of the bowl is cut.)
Sequence 1
Sequence 2
Sequence 3
Transfers to another camera and a near-circular pan by swapping only the segmentation prompt. For larger, less-occluded objects the kinematics step is skipped — pose fully from vision.
Pan 1
Pan 2
Pan 3
Isaac Sim replays and interactive 3D of the full reconstruction will land here — the sim-twin slot in the teaser already shows the interactive viewer format.