Abstract for: Evaluating the Capabilities of AI-Enabled Tools to Build and Edit Models with Modules
Despite rapid progress in AI assisted SD modeling, an important gap remains: there is currently no systematic way to evaluate the extent to which AI-enabled modeling tools can correctly construct and modify SFD models in a modular manner. To evaluate the ability of AI tools to correctly understand, create, and iterate on modular SFD models, three new evaluation categories were created and added to BEAMS Initiative’s evaluation framework. These evaluations test the ability of AI tools to reason, translate, and modify modular SFD models. AI tools are able to competently use modules in SFD model construction and editing. Their performance wanes as the tasks grow in complexity. The tools excel on translation and modification tasks, and perform less well on reasoning tasks. The main challenge identified was with cross-level ghost creation, future AI tools should find a better way to emphasize cross-module relationships in their design. Reasoning tasks performance can likely be enhanced through iteration and multiple applications of existing tools in series. As subject of study