Abstract for: Do LLM-based AI agents missteer accumulation processes like humans?

With the rapid advancement of large language models (LLMs) such as ChatGPT, Claude or Gemini and the rise of agentic artificial intelligence (AI), a pertinent question has emerged: Can AI agents outperform human decision makers in the operational control of dynamic systems, or are they similarly susceptible to misunderstanding of accumulations? The research question is addressed following the well-established experimental paradigm for investigating dynamic decision-making. This pilot study uses a web-based simulation environment that requests participants to assume the role of a production manager and steer the input rate to a two-stock process over 25 days with one decision per day so that the output rate matches a target rate ChatGPT agents surpass human decision makers when managing a two-stock, three-flow production system. Yet, both perform poorly compared to the benchmark heuristic that properly accounts for Little’s Law. Although LLM AI agents are easily accessible and usable, it might be the better option to search for more specialized agentic tools that are better tailored to dynamic operational systems. gAI was extensively used in the experiment, but also to assist with editing and evaluation.