Abstract for: Performance of Large Language Models in Graduate-Level System Dynamics Exams

Large language models (LLMs) are attracting growing attention in System Dynamics (SD) research and practice, particularly for tasks such as conceptual modeling, parameter calibration, and interpretation of simulation outputs. On the other hand, there is very limited evidence on the performance of LLMS in core SD problem-solving. This study evaluates the current capabilities of three widely used LLMs through an exam-based benchmarking in graduate-level SD education. Specifically, the models were asked to answer real midterm and final examination questions under a standardized, minimal prompt setting intended to provide an estimate of performance. Their responses were graded by the course instructor and compared directly with the performance of the student cohort. The findings indicate that ChatGPT 5.2 and Gemini 3 performed above the class average on both examinations, while DeepSeek performed closer to the class average overall. Across question types, the models were relatively strong in short answer questions, stock-flow identification, equation formulation, and qualitative interpretation of dynamic behavior, but weaker in tasks requiring explicit graphical representation of dynamic trajectories. By comparing LLMs with human graduate students in an authentic course setting, this study hopes to contribute to benchmarks for assessing LLM competence in SD. The results suggest that contemporary LLMs can already demonstrate substantial proficiency in several foundational SD skills, while also revealing persistent limitations in formal and visual dynamic representation. These findings can contribute to the future research on the effective integration of LLMs into SD education and modeling practice. Experiments are done by using LLMs