Probing the Augmented Reality Scene Analysis Capabilities of Large Multimodal Models: Toward Reliable Real-Time Assessment Solutions
Augmented reality (AR) is transforming everyday experiences across domains like education, entertainment, and health care. As AR technologies become increasingly widespread, human-aligned, and scalable, AR-quality evaluation is critical for optimizing immersive user experiences. This article investigates the potential of large multimodal models (LMMs) for automating AR quality assessment. We curate DiverseAR+, a new dataset of 1405 scenes collected from diverse sources and environments, and use it to evaluate four commercial LMMs. Our results demonstrate that LMMs can perceive, describe, and judge AR content with promising accuracy. To deliver real-time, robust, and scalable AR-quality evaluation under diverse network conditions, we propose a hybrid cloud–edge architecture that combines LMMs with traditional machine learning models. We argue that task-tailored AR-LMM systems can make AR experience assessment more efficient, adaptive, and user-centered.
Duke Scholars
Altmetric Attention Stats
Dimensions Citation Stats
Published In
DOI
EISSN
ISSN
Publication Date
Volume
Issue
Start / End Page
Related Subject Headings
- Networking & Telecommunications
- 4606 Distributed computing and systems software
- 4009 Electronics, sensors and digital hardware
Citation
Published In
DOI
EISSN
ISSN
Publication Date
Volume
Issue
Start / End Page
Related Subject Headings
- Networking & Telecommunications
- 4606 Distributed computing and systems software
- 4009 Electronics, sensors and digital hardware