Abstract for: Evaluation and Review of Causal Loop Diagrams​ using Large Language Models
Due to the qualitative nature of causal loop diagrams, models quality evaluation methods generally consist of expert review and stakeholder interviews. These methods can be expensive to conduct, and practical limitations often leads to insufficient coverage of different stakeholder viewpoints. Due to their wide breadth of prior knowledge, Large Language Models can serve as a preliminary evaluation method to allow for feedback and review prior to engaging human reviewers. We develop a set of methods using Large Language Models (LLMs) to evaluate the causal structure and coverage of topics in a causal loop diagram. Our approach prompts LLMs to generate a causal loop diagram (CLD) either given a system description or an existing set of nodes. The LLM is prompted repeatedly, resulting in an ensemble of CLDs which can be compared against a CLD generated by human experts. We compare a CLD describing causal pathways of grid outage risk to an ensemble of LLM-generated CLDs. We report classification scores, confidence scores for each edge based on percentage occurrence in the ensemble, and suggestions for additional variables for inclusion based on clustering the LLM-generated variable names. LLM-generated justifications provides insight into the model’s decision making process when generating the CLD. We explored methods for LLMs to evaluate qualitative models and provide qualitative review and feedback to the CLD modelling team. Given the relatively low barriers to integrating automated analysis, we see this approach as valuable for generating feedback that can be used for initial evaluation prior to conducting human expert review. AI was used to evaluate causal loop diagrams