OmniEgoCap:
Camera-Agnostic Sequence-Level Egocentric Motion
Reconstruction
Abstract
Commercial egocentric devices capture human behavior in everyday settings, yet full-body motion reconstruction must generalize across diverse cameras and mountings. This is challenging because head trajectories under-determine body motion, while hand observations are intermittent and camera-dependent: a hand may disappear simply because it has left the camera’s field of view. Existing methods treat such absences as missing data, relying on fixed visibility regimes or post-hoc optimization, and fail to generalize across devices. We present OmniEgoCap, a sequence-level diffusion framework for device-agnostic egocentric full-body motion reconstruction. Our key insight is that egocentric videos contain sequence-level evidence about both the wearer and the camera-induced visibility structure, enabling OmniEgoCap to infer consistent body shape and coherent motion from sparse head and hand cues. To prevent overfitting to a single camera setup, we further introduce geometry-aware visibility augmentation, which synthesizes visibility patterns from realistic variations in camera geometry. We also present OmniEgoDB, the first mocap benchmark with ground-truth motion captured across multiple consumer egocentric devices. Experiments on synthetic, real-device, and in-the-wild settings demonstrate superior motion reconstruction, robust cross-device generalization, and coherent in-the-wild results.
In-the-Wild Demo
Comparison with Baseline
Method Overview
Our framework uses a conditional diffusion model to reconstruct full-body motion from head trajectory and intermittently visible hands in egocentric video. Built on an encoder-decoder transformer with local attention, the model efficiently handles long motion sequences while preserving local motion details. Beyond short-term estimation, we perform sequence-level inference to capture a more global perspective, including a single body shape and camera-dependent visibility cues, enabling temporally coherent reconstructions across diverse egocentric device configurations.
To enable device-agnostic generalization, we introduce a stochastic geometry-aware augmentation strategy that exposes the model to diverse egocentric camera geometries during training. By varying field-of-view, camera tilt, aspect ratio, and the resulting hand visibility patterns, the model learns to interpret intermittent hand cues under different viewing conditions. This enables robust reconstruction on unseen devices with different camera setups, without device-specific retraining.
BibTeX
@inproceedings{cho2026omniegocap,
title={OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction},
author={Cho, Kyungwon and Na, Jeonghyeon and Joo, Hanbyul},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026},
}