SurgAtt-Tracker: Online Surgical Attention Tracking via Temporal Proposal Reranking and Motion-Aware Refinement

Abstract

Visualization Figure

Accurate and stable field-of-view (FoV) guidance is critical for safe and efficient minimally invasive surgery, yet existing approaches often conflate visual attention estimation with downstream camera control or rely on direct object-centric assumptions. In this work, we formulate surgical attention tracking as a spatio-temporal learning problem and model surgeon focus as a dense attention heatmap, enabling continuous and interpretable frame-wise FoV guidance. We propose SurgAtt-Tracker, a holistic framework that robustly tracks surgical attention by exploiting temporal coherence through proposal-level reranking and motion-aware refinement, rather than direct regression. To support systematic training and evaluation, we introduce SurgAtt-1.16M, a large-scale benchmark with a clinically grounded annotation protocol that enables comprehensive heatmap-based attention analysis across procedures and institutions. Extensive experiments on multiple surgical datasets demonstrate that SurgAtt-Tracker consistently achieves state-of-the-art performance and strong robustness under occlusion, multi-instrument interference, and cross-domain settings. Beyond attention tracking, our approach provides a frame-wise FoV guidance signal that can directly support downstream robotic FoV planning and automatic camera control.

Surgical Attention Dataset: SurgAtt-1.16M

Dataset Visualization
We propose SurgAtt-1.16M, a large-scale benchmark for surgical attention tracking spanning diverse laparoscopic procedures. Built upon a clinically grounded annotation protocol that bridges discrete expert attention regions and continuous heatmap representations, SurgAtt-1.16M enables systematic learning and evaluation of temporally coherent surgical attention.
Dataset Visualization
Raw laparoscopic videos are curated into high-quality surgical clips via optical-flow–based operation analysis and expert screening. Videos are sampled at 25 fps and grouped into five representative surgical scenes. During annotation, surgeons mark attention regions with bounding boxes, which are converted into continuous attention heatmaps for supervision. The resulting dataset provides dense, high-fidelity attention annotations across diverse surgical scenarios.

SurgAtt-SZPH Result

Qualitative comparison of attention heatmap predictions on SurgAtt-SZPH across five surgical scenarios. SurgAtt-Tracker produces sharper and more stable attention aligned with clinically relevant regions compared with representative SOTA baselines.

SurgAtt-Hamlyn Result

Qualitative comparison on SurgAtt-Hamlyn under zero-shot and fine-tuning settings. SurgAtt-Tracker shows more stable and better-localized attention than RT-DETRv2 and YOLOv12 across both regimes. (zs = zero-shot; ft = fine-tuning).

SurgAtt-AutoLaparo Result

Qualitative comparison on SurgAtt-AutoLaparo under zero-shot and fine-tuning settings. Our method maintains compact and consistent attention maps under domain shift and further improves after fine-tuning. (zs = zero-shot; ft = fine-tuning).

Licensing

The original dataset and annotations of SurgAtt-Tracker cannot be used for commercial purposes.