Jessica Ezemba, Jason Pohl, Conrad Tucker, Christopher McComb · Communications Engineering 2026 · 2026
DOI: 10.1038/s44172-026-00786-2
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Engineering simulation interpretation is a major bottleneck in design cycles, requiring expensive domain expertise to validate complex outputs and ensure safety and performance. While modern large language models (LLMs) may assist in interpretation, they face fundamental scalability limitations, as even modest simulations exceed the context windows of best-in-class LLMs. Vision-language models (VLMs), having demonstrated success across technical visual reasoning domains from medical imaging to materials characterization, represent a promising alternative for processing simulation visualizations as compressed representations. However, their effectiveness for engineering simulation interpretation remains unknown, constrained by the absence of large-scale evaluation frameworks and prohibitive expert annotation costs. We introduce OpenSeeSimE, a large-scale benchmark consisting of 200,000+ question-answer pairs across 10,000 parametrically-varied simulations. This 850 × scale increase, enables statistically robust evaluation across diverse simulation configurations and question types. Evaluation of ten state-of-the-art VLMs reveals that models demonstrating strong performance on general visual reasoning benchmarks perform at random chance levels (29-47%) on engineering simulations with negligible effect sizes, establishing critical baselines for domain-specific model development. These findings indicate that deploying VLMs for simulation interpretation will require domain-specific training rather than reliance on general-purpose models, and the benchmark provides a reusable framework with sufficient statistical power to measure incremental progress.
No comments yet — start the discussion below.