AgentWorld Benchmark: What AI Teams Playing a Game Can Teach Us About Team Intelligence
As large language models become capable of reasoning, programming, and using tools, a natural question follows: can we organize them into teams that accomplish work beyond what an individual agent can handle?
That possibility is a major part of the appeal of multi-agent systems. One member investigates a problem, another carries out a plan, and a third checks the result. They can explore different directions at once or contribute complementary capabilities. From software development to scientific research, many complex projects already depend on this kind of division of labor.
But turning individual capability into team performance requires coordination. Human teams need a shared goal, an understanding of one another's responsibilities, and the ability to revise their actions as new information arrives. Those abilities cannot simply be assumed for teams of LLM-powered agents.
AgentWorld is a benchmark for agent collaboration. It asks whether multiple agents, with different roles and resources, can work together over many rounds to complete a shared task.
This connects foundation-model research with practical system design. It raises questions about what future models should learn, how we should organize the models available to us, and how we can tell whether a multi-agent system is actually improving.

Figure 1. A staged team scene in the AgentWorld engine, illustrating different roles. Characters and dialogue were arranged by a demonstration script; this is not an autonomous model rollout.
1. What turns a collection of agents into a collaborating team?
Multiple agents can provide several different kinds of benefit.
The most immediate is broader exploration: several agents investigate different questions in parallel, then combine their findings. Another is complementary capability: members take on responsibilities suited to their tools, information, or expertise. A more demanding form is sustained coordination, where one member's actions change what other members should do next.
These situations make different demands on a team. Reading ten separate papers and producing a combined report allows relatively independent assignments. Developing a changing software system requires members to manage interface changes, dependencies between components, and revisions to the implementation plan.
This leads to a deeper question: when work cannot be fully separated into independent pieces, can multiple agents still deliver reliable gains?
A team then needs more than a list of each member's abilities. It must understand who is waiting for whom, which information has become outdated, and whether local objectives still support the shared goal. A member can execute its assigned step correctly while leaving the overall project blocked because a necessary handoff never occurred.
AgentWorld puts these dependencies at the center of evaluation. It creates situations in which the usefulness of one agent's work depends on how that work connects to the rest of the team.
2. How AgentWorld makes collaboration observable
AgentWorld places agents in an MMORPG sandbox with 100 human-authored tasks and 100 augmented variants. Each task involves 3 to 20 agents, with activities such as gathering materials, crafting equipment, allocating resources, and completing combat objectives. Long-horizon settings extend beyond 50 interaction rounds.
The value of this environment is that coordination has consequences in the shared world. If materials are not transferred, the next crafting step cannot begin. If an update does not reach a teammate, that teammate may continue following an outdated plan. Across many rounds, the team has to connect its members' local actions into a coherent sequence.
High-level tools handle operations such as navigation and gathering, reducing the demands of low-level game control. Agents cannot directly inspect one another's internal reasoning. They coordinate through messages and observable actions, while programmatic verifiers check whether task objectives have been achieved.
This makes collaboration easier to study, while retaining the planning and tool-use challenges that accompany it. Results describe team performance within this setting; they do not isolate a single, context-free measure of collaborative intelligence.
A five-member team turns distributed resources into three items
To make the dependencies concrete, consider the mining and crafting production goal described in Task 46: an iron supplier, a coal supplier, a smelter, a smith, and a courier must connect their work to produce a pickaxe, an axe, and a heavy sword.
The four scenes below combine actual game-engine screenshots, item artwork, and saved inventory records. Resources and character positions were preset so the demonstration could focus on handoffs and production, omitting gathering and long-distance travel. Transfers, smelting, and crafting were executed in the engine.
This is a scripted demonstration of the production chain, not an autonomous LLM rollout or a formal benchmark score.
Scene 1: understand the shared goal and the resources each member holds
The iron supplier holds ore, the coal supplier holds fuel, and the courier holds wood. The smith already has a sword hilt. The smelter and smith perform different processing steps. No member begins with all the conditions needed to complete the production goal.
The first coordination problem is therefore to connect available resources with the people who can use them. The team needs to understand both what exists and where it needs to go. A resource sitting in one member's inventory does not automatically help another member execute a recipe.

Figure 2a. The initial state distributes ore, coal, and wood across three suppliers. The processing roles depend on those inputs reaching them. Quantities come from saved inventory snapshots.
Scene 2: bring the supply lines together
Ore and coal are transferred to the smelter, while wood goes directly to the smith. The smelter now has the inputs required to begin processing. The smith has made progress toward readiness but still needs iron bars.
This distinction matters in any collaborative system: the team possessing something and the responsible member being able to use it are different states. A completed handoff opens up the next action. Announcing that resources are available does not, by itself, establish that the next member can proceed.

Figure 2b. Solid arrows show completed transfers. The dashed arrow shows the iron-bar handoff that still needs to happen. Crops show characters at this stage; arrows annotate the recorded resource flow.
Scene 3: the first finished item is only partial progress
The smelter delivers the first five iron bars. The smith combines them with one log to make a pickaxe. That completes one of the three requested outputs.
The team must continue. The axe needs three more bars and one log; the heavy sword needs two bars and the hilt. The smelter therefore has to supply another five bars, and the smith has to preserve the remaining components for their intended uses.
This is where a shared goal has to remain active across multiple local successes. Completing one useful action is not sufficient evidence that the entire task is finished. Members need to track the remaining dependencies as the world changes.

Figure 2c. Five bars and a log become the first pickaxe. The inventory confirms the item, while the team still needs an axe and a sword.
Scene 4: verify the common objective through actual outputs
After the second batch of five bars arrives, the smith finishes the axe and heavy sword. A final inventory check confirms that all three newly crafted items exist.
The effective collaboration in this sequence is visible in the flow of resources across roles, the ordering of processing steps, and the team's continued attention to the complete goal. Completion is grounded in the resulting state of the world.

Figure 2d. A final-stage screenshot and the three outputs confirmed in the smith's inventory. This checks the demonstration's production goal and is not a formal task score.
The example gives team intelligence a concrete meaning. Can members recognize dependencies, supply the right teammate, and distinguish local progress from overall completion?
In software or research work, the object being handed over might be a usable module, experimental data, or a validated finding. The common requirement is that one member's work becomes something another can actually use. Whether improvements in this sandbox transfer to those domains is a question for further evaluation.
3. What the experiments reveal about individual capability and team performance
The paper evaluates teams powered by Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B. On the main tasks, they achieve success rates of 52.0%, 45.0%, 36.0%, and 20.0%, respectively. On the augmented tasks, their success rates are 24.0%, 26.0%, 21.0%, and 10.0%.
Figure 3. Results for the models and configurations evaluated in the paper. They describe performance under those settings, rather than an upper bound on model capability or on multi-agent architectures.
These results show substantial difficulty in sustaining effective team performance in the evaluated setting. Understanding a task, planning actions, and using tools are all relevant abilities, but their translation into a team's final delivery has to be measured directly.
Communication volume cannot answer that question on its own. In the main task set, GPT-5 Mini averages 44.1 messages, compared with Gemini 3 Flash's 11.0, while Gemini achieves a higher success rate. This comparison does not demonstrate that more messages cause worse results. It suggests looking at what communication changes in the team's behavior, alongside how much communication takes place.
For model developers and application teams, an important research question follows: which coordination failures can stronger reasoning resolve, and which require changes to task allocation, shared information, or execution feedback?
AgentWorld provides an executable environment for investigating these factors. Controlled comparisons could hold the model fixed while changing the team structure, or hold the structure fixed while changing the model. Such comparisons are needed before attributing a result to a particular source of improvement.
4. CCE: how much of the team's activity contributes to the result?
Success rate tells us whether a task was completed. It gives a less detailed account of how the team got there. Two successful teams may have followed very different paths, with one coordinating smoothly and another repeating work or pursuing unproductive actions.
The paper introduces Causal Collaboration Effectiveness (CCE) to examine this process. It starts from actions that directly achieve the task and traces backward through relevant earlier actions and messages, identifying a chain of contributions to the final result.
In the crafting example, the finished equipment depends on iron bars arriving at the smith. Those bars depend on the smelter receiving ore and fuel. A message that prompts a necessary delivery may also become part of the contribution chain. CCE asks how the team's activity connects to the outcome.

Figure 4. In this illustrative trajectory, 12 of 27 actions count as contributing, giving approximately 44.4%. This number explains the pictured example, not the benchmark's aggregate performance.
For successful tasks, CCE is the fraction of actions judged to contribute. Failed tasks receive zero under the paper's definition. This adds a process perspective to evaluation: did messages and actions form useful connections leading toward the result?
The contribution decisions are LLM-assisted binary judgments. In a small validation study covering 84 judgments across 11 tasks and two agent models, one human annotator agrees with GPT-4.1 on 69 judgments, approximately 82%, with Cohen's κ of 0.64. GPT-4.1 and Claude Sonnet 4 agree on 84% of 7,298 decisions, with κ of 0.66. Human-judge agreement is similar to cross-judge agreement in these samples.
That provides initial support for the method, while leaving limitations to investigate. The human sample is small, and an inclusive contribution policy may count weak contributions. The inferred graph is not a substitute for causal intervention experiments. Since failed tasks score zero, CCE also partly reflects task success.
CCE is therefore useful alongside success rate, time, and cost. It offers a way to discuss which handoffs helped, which exchanges changed subsequent actions, and which activities may have failed to advance the goal. Those questions are more informative than judging collaboration by how fluent a conversation sounds.
5. Paths from capable models to capable teams
From a system-design perspective, collaboration can be organized in several ways, with different levels of interdependence. Three useful patterns are parallel work, coordinator-led execution, and continuous adaptation among teammates. They suit different tasks and can be combined.
Parallel work: broaden exploration and combine the results
Multiple agents can explore separate directions before a member brings their findings together. This makes direct use of models' existing ability to solve problems independently.
The design challenge is to define useful task boundaries, reduce duplicated effort, and reconcile conflicting findings. Where assignments are relatively independent, parallelism can provide a straightforward structure. As dependencies increase, the team needs a way to handle discoveries that change someone else's assignment before the final synthesis.
Coordinator-led teams: maintain a shared execution plan
A coordinating agent can break down the objective, assign responsibilities, and revise the plan as results arrive. For tasks with many dependencies, this offers a common view of what should happen next and which members are blocked.
That structure also concentrates responsibility. The coordinator's judgment, information-processing capacity, and understanding of the shared state can constrain the team. A useful evaluation should examine how it responds to missing updates, mistaken assumptions, or a plan that no longer fits the situation.
Adaptive teams: adjust while the work is in progress
Members can also update their behavior in response to teammates' progress, needs, and changes in the environment. Collaboration then includes knowing when to request help, when to proactively share information, and when to revise the original division of work.
AgentWorld's long-horizon tasks, asymmetric roles, and blackbox interaction make these behaviors observable. An agent may need to notice that its teammate cannot proceed and change its own next action, even if its original assignment has already been completed.
Figure 5. Three ways to organize collaboration. This is an analytical framework for the article, not a measured performance ranking or an inevitable sequence of technical progress.
These patterns do not imply that decentralization must outperform centralized coordination. The meaningful comparison is which organization succeeds more reliably on the same tasks, under comparable resource budgets, and which can recover when something goes wrong.
For industry readers, this means evaluating both model capability and organizational design. A model upgrade may improve reasoning and planning. A better team structure may make more effective use of capabilities already available. Their interaction needs controlled experiments and measurable task outcomes.
6. What comes next: learning to collaborate and improving the system
One direction is to make collaboration itself part of what models learn.
If training feedback focuses only on whether an individual's output is correct, it may provide limited guidance about when to communicate, how to understand a teammate's needs, or when to change a local action for the shared goal. Team outcomes and process feedback could help models learn those decisions through interaction.
CCE-like measures may support analysis of that process. Whether they also make effective training signals needs separate validation. A useful diagnostic measure does not automatically become an objective that produces better behavior when optimized.
A second direction is to make team organization depend on the task. Independent parts of a project may benefit from parallel exploration. Steps with tight dependencies may need coordination. Uncertain results may call for additional verification before they become inputs to other members.
Future systems could investigate how to choose team size, assign responsibilities, and revise those arrangements during execution. This raises a practical resource question: when does adding a member contribute enough value to justify the extra communication, waiting, and coordination it creates?
Evaluation also needs to become more demanding. Does the team remain reliable when an objective changes, a member becomes unavailable, or agents powered by different models work together? Can it recover without restarting the entire task? Does an apparent improvement persist under comparable time and compute budgets?
These are forward-looking research questions, rather than capabilities established by the current experiments. A shared environment helps make them testable by connecting system choices to observable behavior and final outcomes.
AgentWorld provides a common starting point for this work: executable tasks, observable collaboration processes, and comparable results.
As individual models continue to improve, we also need to understand how their capabilities combine. Reliable teams will depend on the joint development of models, organizational structures, and evaluation feedback. Studying those relationships is a practical step toward agents that can sustain complex work together.
About the research team
The paper has been accepted at COLM 2026 and is now available on arXiv.
AgentWorld was developed by researchers from OpenAgents, Columbia University, the University of Pennsylvania, Seoul National University, and Penn State. Raphael Shu, Yusen Zhang, and Young Min Cho are the co-first authors.
Project website: agentworld.io
Code and dataset: openagents-org/agentworld on GitHub
Paper: AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs