Which AI is better at collaborating with humans?
As part of our ongoing mission to enhance human+ AI collaboration, we ran an experiment using our model to see how AIs and humans coordinate in a simple game. In this case, we used a platformer where the players are tied together by a rope and no level can be finished alone. Four AI models played in teams of two, three and four, with memory between games. Human pairs played on their own or with an agent. We recorded levels, resets, falls, messages and scores.
Two things stood out, and both are about collaboration, not raw capability. Who's in charge changes everything: mixed teams did well when a human called the plan and badly when the agent did, same agent, same task, opposite result. And more agents made things worse, not better: every model's win rate dropped as the team grew from two to four. Understanding a teammate matters more than being smart alone. In good news for us humans everywhere, when paired together, humans performed better than the best models.
Four frontier models and dozens of human pairs ran the same eight levels, alone, together and mixed, while we logged every fall, message, reset and score, plus who led, who followed, and how the humans adapted. That's the same harness we expose to partner labs that want to know whether their model actually helps a person, not just how it performs alone.
Skillprint's technology instruments games to capture how people think, act and feel while they are playing, not just what actions they took. We think useful AI has to be human-context aware: it has to understand the person it's working with, especially in context of the task in front of it. This rope game is one instrument built to test the nature of coordination, for humans and models alike.

In one minute
Teams of two to four play an eight-level cooperative platformer that cannot be finished alone. Teams are made of AI models (Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Kimi K3), of humans, or of two humans and one agent. The question: who is the better teammate, and what happens when you mix them.
Every seat is its own player with its own screen, a shared chat box, and, for agents, a playbook that carries between games. A game is won when all eight levels are cleared inside the budget. We measure how far a team gets, how many resets it uses, how many times it fails to coordinate, how many decisions it takes, how much it talks, and its score.
–% of human pairs beat the game, –% of human-led mixed teams, –% of the best agent pair, –% of agent-led mixed teams, –% of agent teams overall. Humans win on coordination, planning and speed; the top models can win with memory and tries; more agents on the rope means fewer wins.
Findings
Who beats the game
A game is beaten when every player stands in the eighth exit. The rate below includes every run of that team type. Each bar shows the win rate and the number of games beaten.
Figure 1. Games beaten, by team type
Which agent pairs
Every pairing of the four models, – games each, from level 1 with memory on; the diagonal is a model with a copy of itself. The last row and column add the humans: two humans on their own, and two humans with one agent of that model, – games each, split by who called the plan.
Figure 2. Pairings, humans included: share of games beaten


Two, three, four on a rope
Homogeneous teams of each model, at three sizes. Adding teammates reduced the win rate for every model.
Figure 3. Games beaten by team size, homogeneous agent teams

Humans, and humans with an agent
Two humans with one agent, – games per model, half with a human calling the plan and half with the agent calling it and the humans following.
Figure 4. Coordination failures, resets and time, by team type

Who talks
Figure 5. Messages per 10 decisions, and the share that name a teammate
How it was measured
The same six numbers for every team, human or not: how far (levels cleared and game clock), attempts (team resets), coordination failures (rope hangs, falls with nobody anchored, bridges retracted under a player, wrong code digits, lift trips made empty), steps (decisions per seat; for humans, actions), communication (messages, and the share naming a teammate), and score. Everything is read from the room's event log, so a human room and an agent room are measured by the same code.