A large-scale benchmark to assess vision-language model question answering capabilities in engineering simulations
In our paper, we introduce a large-scale benchmark called OpenSeeSimE that asks the question of how well AI models can currently reason about engineering simulations.