A large-scale benchmark to assess vision-language model question answering capabilities in engineering simulations
In our paper, we introduce a large-scale benchmark called OpenSeeSimE that asks the question of how well AI models can currently reason about engineering simulations.
Published in Computational Sciences, Mathematics, and Mechanical Engineering
Engineering design usually takes the form of multiple stages over multiple experts and disciplines. There have been multiple abstractions for this process of engineering design, and one of them is Asimow's design cycle, where he breaks up the design process into three stages: Analysis, Synthesis, and Evaluation. This paper concerns the last stage, which deals with testing candidate designs against the requirements, then feeding the results back into the design.

Because it is difficult to build and manufacture a potential design and test it physically, especially in the early stages of the design when you are still trying to get the final specifications, most engineers run these tests using engineering simulation tools. So, for example, if you are designing a chair and you are still trying to determine if multiple candidate designs could work, or your design limit may be a baby bear (300 lbs), then instead of building each chair and dropping 300 lbs on it and seeing if it breaks, which is very expensive, you can run the simulation below to determine if the chair will break or not.
This is something that is common enough that there are multiple companies that have built tools to do these evaluative tasks, with experts being at the forefront to determine what is going on in these simulations (e.g., if the chair will break or not). Recently, there has been a paradigm shift to see if we can automate the engineering design cycle, with papers automating multiple different parts of the stages, but the evaluation stage still remains a bottleneck because humans typically make the determination of whether it passes or not though some papers are starting to tackle it. When we first started this work, we asked the question: what will it take to get an AI model to be able to take on these evaluative tasks and actually use them to inform the design loop? For example, looking at the simulation below to not just determine if it will break, but also determine if it meets the displacement criteria, or whether it moves a lot if the baby bear is placed on it. As anyone would know, if you sit on a chair and it deforms a lot but doesn't break, you probably will not want to sit on it again. Slight discrepancies like this are things experts know to design for, so we wanted to ask if AI models can do this.
In order to test that, we had to create a controlled experiment where we could measure the performance of these models on evaluative tasks. This is often called benchmarking, where you have a fixed set of questions and answers, and you use it to evaluate multiple different methods, and you can tell which method is better by looking at who got the most answers correct. This way, you can think of it like an exam, but for AI models. So we set out to create a benchmark for evaluative tasks in engineering design, because before this time, this had not been done. The test was expertly designed so that we would get up to 200,000 questions and answers like the ones below.
We used it to compare AI models available at the time, and as you can see below, most models scored around 50%, which shows the level of models at that time. This benchmark not only helps to evaluate where current models fall, but also where future models fall, because this test can be given to future AI models as well. Through this work, we see the distribution with the effect of binary questions, multiple-choice questions, or spatial questions below.
The results showed that models struggle with spatial and multiple-choice questions, and between domains (structural analysis and fluid dynamics), models struggle more with fluid dynamics than structural analysis. If you also look at the images and videos, there is not much difference in accuracy, even though videos tend to encode more information, such as flow development or loading direction. This difference also does not change much when looking at closed-source models and open-source models, showing that the models tested at the time still struggle with these evaluative tasks.
Looking qualitatively, we also saw that some errors are specific to models. For example, the smallest Qwen models had explicit refusals, where they had access to the information needed to answer the question but decided the information was insufficient. There are also errors where the model confused the direction of the flow, which also shows a lack of spatial understanding. In summary, the OpenSeeSimE benchmark allows us to test current and future models on evaluative tasks, with the goal that being able to automate this stage could lead to larger initial design sweeps and potentially the exploration of vast design spaces.
Follow the Topic
Engineering Design
Technology and Engineering > Mechanical Engineering > Engineering Design
Artificial Intelligence
Mathematics and Computing > Computer Science > Artificial Intelligence
Computational Science and Engineering
Mathematics and Computing > Mathematics > Computational Mathematics and Numerical Analysis > Computational Science and Engineering
-
Communications Engineering
A selective open access journal from Nature Portfolio publishing high-quality research, reviews and commentary in all areas of engineering.
Related Collections
With Collections, you can get published faster and increase your visibility.
Generative AI for mechanical engineering design and optimization
In this collection we aim to publish exciting advances in the capability of generative AI methods and their application directions within a broad mechanical engineering scope.
Publishing Model: Open Access
Deadline: Dec 31, 2026
Industry Showcase 2026: Solutions for Sustainability
In our Industry Showcase collection for 2026, we would like to celebrate the research contributions from industry researchers advancing engineering solutions to address issues of sustainability.
Publishing Model: Open Access
Deadline: Oct 28, 2026