Endo4DLR: Online Language-Aligned 4D
Reconstruction with Persistent Gaussian Anchors

TL;DR: Endo4DLR jointly estimates cameras and reconstructs language-embedded Gaussians from unposed endoscopic streams, enabling online reconstruction and language grounding without per-scene optimization.

Abstract

Online RGB-D Reconstruction

Drag the divider to compare RGB and depth.

Input
Endo4DLRStream3R
Selected frames from the paper

Explore the Reconstruction in 3D

Orbit the reconstructed point cloud and play the sequence.

Reconstructed point cloud preview
Loads on demand
Reconstructed point cloud
Frame 1 / 3Full sequence

Online Language Grounding

Instrument, action, and location queries on the rendered language representation.

Input
Reference
Endo4DLR
Stream3R
ZipMap
CUT3R

3D Gaussian activation

Input Video with GT query localization
0:00
Activated Gaussian

Drag to orbit · Scroll to zoom · Shift-drag to pan

Method Overview

Endo4DLR method: causal reconstruction of language-aligned Gaussians, then online language grounding.

Joint camera and scene prediction

Camera parameters and Gaussian scenes are jointly predicted from unposed endoscopic video.

Language-embedded Gaussians

Gaussian language features support queries about instruments, their actions, and locations.

Persistent anchor memory

Anchors retrieve historical context for attribute prediction while Gaussian centers follow current geometry.

Adaptive Gaussian queries

Queries adapt to local image details and form spatially local groups through Z-order serialization.