Preview

Journal of Instrument Engineering

Advanced search
Open Access Open Access  Restricted Access Subscription Access

SAM3R: Object-Centric 3D Mapping via Foundation-Model-Guided Data Association in Changing Scenes

https://doi.org/10.17586/0021-3454-2026-69-6-485-496

Abstract

For embodied agents to navigate and reason indoor spaces, they need object-level 3D representations that stay consistent over time as new frames arrive from a monocular camera. Current online 3D instance segmentation methods either depend on posed RGB-D input with ground-truth depth or couple tightly to the internal representations of specific foundation models, sacrificing modularity. We observe that appearance-based and geometry-based object matching exhibit complementary failure modes: appearance is ambiguous among spatially separated duplicates, while geometry is unreliable for visually distinct objects at similar locations. This motivates SAM3R, a training-free pipeline that fuses spatial overlap, 3D centroid displacement, and visual-semantic similarity into a single assignment cost solved via bipartite matching. The cost is constructed entirely from the outputs of frozen foundation models without accessing internal representations. Object tracks are classified through a cascaded decision tree that detects scene changes via field-of-view gated temporal voting. On ScanNet200 and Replica, SAM3R performs competitively with methods that require architecture-specific features or additional training, despite operating in a fully online, monocular setting. Qualitative evaluation on the Aria Digital Twin dataset further demonstrates that the pipeline maintains correct object identities through physical object manipulation, including hand occlusion and large spatial displacement.

About the Authors

M. Mohrat
ITMO University; Sber Robotics Center
Russian Federation

Malik Mohrat — PhD student, Faculty of Control Systems and Robotics; Leading Engineer-Developer

Saint Petersburg; Moscow



E. E. Derevyanka
Sber Robotics Center
Russian Federation

Ekaterina Е. Derevyanka — PhD, Executive Director

Moscow



I. I. Obrubov
Sber Robotics Center
Russian Federation

Ilya I. Obrubov — Leading Engineer

Moscow



I. А. Sosin
Sber Robotics Center
Russian Federation

Ivan А. Sosin — Executive Director

Moscow



S. A. Kolyubin
ITMO University
Russian Federation

Sergey A. Kolyubin — Dr. Sci., Professor, Faculty of Control Systems and Robotics

Saint Petersburg



References

1. Zhang J., Dai L., Meng F., Fan Q. et al. 3D-aware object goal navigation via simultaneous exploration and identification // CVPR. 2023. P. 6672–6682. DOI:10.1109/CVPR52729.2023.00645.

2. Mur-Artal R., Tardos J. D. ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras // IEEE Transactions on Robotics. 2017. Vol. 33, N 5. P. 1255–1262.

3. Kirillov A., Mintun E., Ravi N., Mao H. et al. Segment Anything // ICCV. 2023. P. 4015–4026. DOI: 10.1109/ICCV51070.2023.00371.

4. Oquab M., Darcet T., Moutakanni Th., Vo H. V. et al. DINOv2: Learning Robust Visual Features without Supervision // Transactions on Machine Learning Research. 2024. Vol. 2024. https://openreview.net/pdf?id=a68SUt6zFt.

5. Radford A., Kim J. W., Hallacy Ch., Ramesh A. et al. Learning Transferable Visual Models from Natural Language Supervision // ICML. 2021. P. 8748–8763. https://dblp.org/rec/conf/icml/RadfordKHRGASAM21.html.

6. Wang Sh., Leroy V., Cabon Y., Chidlovskii B. et al. DUSt3R: Geometric 3D Vision Made Easy // CVPR. 2024. P. 20697–20709. DOI: 10.1109/CVPR52733.2024.01956.

7. Zhang J., Herrmann Ch., Hur J., Jampani V. et al. MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion // ICLR. 2025. https://openreview.net/pdf?id=lJpqxFgWCM.

8. Wojke N., Bewley A., Paulus D. Simple Online and Realtime Tracking with a Deep Association Metric // ICIP. 2017. P. 3645–3649. DOI: 10.1109/ICIP.2017.8296962.

9. Du Zh., Danier D., Lenssen J. E., Bilen H. MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors // arXiv preprint arXiv:2512.15577. 2025.

10. Tang Y., Zhang J., Lan Y., Guo Y. et al. Onlineanyseg: Online zero-shot 3d segmentation by visual foundation model guided 2d mask merging // arXiv:2503.01309v3 [cs.CV]. 2025. DOI:10.48550/arXiv.2503.01309.

11. Sun Y.-Ch., Tseng Y.-H., Ho Y.-H., Liu Y.-L. 3AM: Segment Anything with Geometric Consistency in Videos // arXiv preprint arXiv:2601.08831. 2025.

12. Kuhn H. W. The Hungarian Method for the Assignment Problem // Naval Research Logistics Quarterly. 1955. Vol. 2, N 1–2. P. 83–97.

13. Ravi N., Gabeur V., Hu Y.-T., Hu R. et al. SAM~2: Segment Anything in Images and Videos // arXiv preprint arXiv:2408.00714. 2024.

14. Cheng H. K., Cho S. W., Schwing A. G. SAM2Long: Enhancing SAM\,2 for Long Video Segmentation with a TrainingFree Memory Tree // arXiv preprint arXiv:2410.16268. 2024.

15. Videnović J., Kristan M. & Lukežič A. Distractor-Aware Memory-Based Visual Object Tracking // International Journal of Computer Vision. 2026. Vol. 134. Art. no. 211. https://doi.org/10.1007/s11263-026-02790-7.

16. Zhang Ch., Han D., Zheng Sh., Choi J. et al. Mobilesamv2: Faster segment anything to everything // arXiv preprint arXiv:2312.09579. 2023.

17. Yang Y., Wu X., He T., Zhao H. et al. SAM3D: Segment Anything in 3D Scenes // ICCVarXiv:2306.03908v1 [cs.CV]. 2023. https://doi.org/10.48550/arXiv.2306.03908.

18. Johnson J., Douze M., Jegou H. DINOv3 // IEEE Transactions on Big Data. 2021. Vol. 7, N 3. P. 535–547.

19. Bewley A., Ge Z., Ott L., Ramos F., and Upcroft B. Simple online and realtime tracking // 2016 IEEE Intern. Conf. on Image Processing (ICIP). Phoenix, AZ, USA, 2016. P. 3464–3468. DOI: 10.1109/ICIP.2016.7533003.

20. Wen B., Yang W., Kautz J., and Birchfield S. FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects // 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle, WA, USA, 2024. Р. 17868–17879. DOI: 10.1109/CVPR52733.2024.01692.

21. Lu Sh., Chang H., Jing E. P., Boularias A. et al. OVIR-3D: Open-Vocabulary 3D Instance Retrieval Without Training on 3D Data // CoRL. 2023. P. 1610–1620. https://proceedings.mlr.press/v229/lu23a/lu23a.pdf.

22. Xu X., Chen H., Zhao L., Wang Zh. et al. EmbodiedSAM: Online Segment Any 3D Thing in Real Time // arXiv:2408.11811v3 [cs.CV]. 2025. https://arxiv.org/html/2408.11811v3.

23. Ester M., Kriegel H.-P., Sander J., Xu X. A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise // KDD. 1996. P. 226–231. https://www.cs.sfu.ca/~ester/papers/kdd_96.pdf.

24. Rozenberszki D., Litany O., Dai A. Language-Grounded Indoor 3D Semantic Segmentation in the Wild // Computer Vision – ECCV 2022. Lecture Notes in Computer Science. 2022. Vol. 13693. Springer, Cham. https://doi.org/10.1007/978-3-031-19827-4_8.

25. Yan M., Zhang J., Zhu Y., Wang H. MaskClustering: View Consensus Based Mask Graph Clustering for OpenVocabulary 3D Instance Segmentation // CVPR. 2024. P. 28274–28284. DOI: 10.1109/CVPR52733.2024.02671.

26. Straub J., Whelan Th., Ma L., Chen Y. et al. The Replica Dataset: A Digital Replica of Indoor Spaces // arXiv preprint arXiv:1906.05797. 2019.

27. Pan X., Charron N., Yang Y., Peters S. et al. Aria digital twin: A new benchmark dataset for egocentric 3d machine perception // ICCV. 2023. P. 20133–20143. DOI:10.1109/ICCV51070.2023.01842.


Review

For citations:


Mohrat M., Derevyanka E.E., Obrubov I.I., Sosin I.А., Kolyubin S.A. SAM3R: Object-Centric 3D Mapping via Foundation-Model-Guided Data Association in Changing Scenes. Journal of Instrument Engineering. 2026;69(6):485-496. (In Russ.) https://doi.org/10.17586/0021-3454-2026-69-6-485-496

Views: 210

JATS XML

ISSN 0021-3454 (Print)
ISSN 2500-0381 (Online)