Can AI agents simulate a firm?
Abstract
AI agents built on large language models can now plan, remember, and interact, making it possible, for the first time, to simulate an organisation populated by human-like individuals. But before such simulations can be trusted, one question has to be answered: do these agents genuinely behave like people, or do they just repeat patterns from the research they were trained on?
My PhD develops a validation protocol to tell the two apart — testing AI agents at the level of the individual, the group, and ultimately the firm. The goal is to establish whether AI-based firm simulation can become a credible instrument for management research and, eventually, a practical tool for organisations to rehearse their decisions before making them.
Why simulate organisations at all?
Firms are complicated. Decisions made in one corner of an organisation ripple through hierarchies, incentives, and team dynamics in ways that are hard to see and even harder to predict. Management researchers use computer simulation to make sense of this complexity — building simplified models of firms to understand how they search for new strategies, adapt to change, and sometimes get stuck. Beyond theory, simulation also holds a practical promise: letting a firm test a strategy, a reorganisation, or a what-if scenario before committing to it in the real world.
However, to keep simulations tractable, the "people" inside them had to be radically simplified. This meant reducing the behavioural richness that makes organisations interesting — and potentially limiting how well these models can predict what real firms do.
What large language models change
Large language models (LLMs) are trained on vast amounts of human-generated text. Because of this, researchers have argued that they are, in effect, implicit models of human behaviour. Early evidence is encouraging: when LLM-based agents are placed in classic experiments from psychology, economics, and management, they tend to reproduce (to a certain extent) the choices and patterns that real humans show.
The technology is also moving fast. LLMs have evolved from chatbots into agents: systems that can plan, remember, use tools, and act in an environment. Researchers have built virtual villages where 25 AI agents autonomously organised a Valentine's Day party, simulated societies of 100 agents living out ten years of social life, and even virtual software companies where AI agents in roles like CEO, programmer, and tester collaborate to ship working code.
For the first time, it seems feasible to populate a simulated organisation with individuals that behave in genuinely human-like ways.
The catch: imitation is not understanding
Building such systems to model organisations is not this straightforward, however, and poses many challenges. A major one: experiments used to test whether LLMs behave like humans are often published in the very literature the models were trained on. So when an AI agent "replicates" a famous finding, there are two very different explanations:
- It genuinely reproduces the mechanism — the underlying behavioural process that generates the outcome.
- It merely recites the outcome it has read about, like a student who memorised the answer key without understanding the material.
A simulation that recites can only tell you what research has already found. A simulation that reproduces mechanisms could tell you something new — how a firm might behave in a situation no one has studied yet.
Whether LLM-based simulations reproduce mechanisms or merely recite known outcomes has never been tested at the level of the firm.
My research
I argue that the binding constraint on AI-based firm simulation is no longer building such simulations, it is validating them. Before multi-agent systems can serve as instruments for strategy and management research, we need a way to tell a simulation that genuinely behaves like an organisation apart from one that convincingly parrots the literature.
My PhD develops that validation protocol in three steps:
- 1
The individual
Do single AI agents reproduce the behavioural mechanisms of search and decision-making, or do they recite documented outcomes? I confront agents with classic strategy tasks under conditions specifically designed to separate the two.
- 2
The group
Extending the test to interacting agents: do teams of AI agents reproduce the dynamics of group search and decision-making, rather than just its average outcomes?
- 3
The firm
Finally, composition: when individually validated agents are embedded in an organisational structure — hierarchy, incentives, interdependent tasks — does the emergent firm-level behaviour remain faithful?
Why it matters
If validated, LLM-based firm simulation could offer what earlier methods never could: organisational models populated by behaviourally rich individuals, with credibility from the individual level up to the firm. That would open a new instrument for management research and, eventually, a practical tool for organisations to rehearse their most consequential decisions before making them.
If not, the field will have learned something equally important: where the line runs between artificial imitation and genuine behavioural replication, and how to test for it.