Skip to main content

Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality

Conferences
Xiu, Y; Chilukuri, J; Sen, S; Gorlatova, M
Published in: Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025
January 1, 2025

As augmented reality (AR) applications increasingly require 3D content, generative pipelines driven by natural input such as speech offer an alternative to manual asset creation. In this work, we design a modular, edge-assisted architecture that supports both direct text-to-3D and text-image-to-3D pathways, enabling interchangeable integration of state-of-the-art components and systematic comparison of their performance in AR settings. Using this architecture, we implement and evaluate four representative pipelines through an IRB-approved user study with 11 participants, assessing six perceptual and usability metrics across three object prompts. Overall, text-image-to-3D pipelines deliver higher generation quality: the best-performing pipeline, which used FLUX for image generation and Trellis for 3D generation, achieved an average satisfaction score of 4.55 out of 5 and an intent alignment score of 4.82 out of 5. In contrast, direct text-to-3D pipelines excel in speed, with the fastest, Shap-E, completing generation in about 20 seconds. Our results suggest that perceptual quality has a greater impact on user satisfaction than latency, with users tolerating longer generation times when output quality aligns with expectations. We complement subjective ratings with system-level metrics and visual analysis, providing practical insights into the trade-offs of current 3D generation methods for real-world AR deployment.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025

DOI

Publication Date

January 1, 2025

Start / End Page

320 / 326
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Xiu, Y., Chilukuri, J., Sen, S., & Gorlatova, M. (2025). Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality. In Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025 (pp. 320–326). https://doi.org/10.1109/ISMAR-Adjunct68609.2025.00069
Xiu, Y., J. Chilukuri, S. Sen, and M. Gorlatova. “Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality.” In Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025, 320–26, 2025. https://doi.org/10.1109/ISMAR-Adjunct68609.2025.00069.
Xiu Y, Chilukuri J, Sen S, Gorlatova M. Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality. In: Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025. 2025. p. 320–6.
Xiu, Y., et al. “Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality.” Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025, 2025, pp. 320–26. Scopus, doi:10.1109/ISMAR-Adjunct68609.2025.00069.
Xiu Y, Chilukuri J, Sen S, Gorlatova M. Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality. Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025. 2025. p. 320–326.

Published In

Proceedings 2025 IEEE International Symposium on Mixed and Augmented Reality Adjunct Ismar Adjunct 2025

DOI

Publication Date

January 1, 2025

Start / End Page

320 / 326