Joint camera and scene prediction
Camera parameters and Gaussian scenes are jointly predicted from unposed endoscopic video.
TL;DR: Endo4DLR jointly estimates cameras and reconstructs language-embedded Gaussians from unposed endoscopic streams, enabling online reconstruction and language grounding without per-scene optimization.
Drag the divider to compare RGB and depth.
Orbit the reconstructed point cloud and play the sequence.
Instrument, action, and location queries on the rendered language representation.

Camera parameters and Gaussian scenes are jointly predicted from unposed endoscopic video.
Gaussian language features support queries about instruments, their actions, and locations.
Anchors retrieve historical context for attribute prediction while Gaussian centers follow current geometry.
Queries adapt to local image details and form spatially local groups through Z-order serialization.