Skillprint Experiments
Signal · four AI models · humans · one rope

Which AI is better at collaborating with humans?

As part of our ongoing mission to enhance human+ AI collaboration, we ran an experiment using our model to see how AIs and humans coordinate in a simple game. In this case, we used a platformer where the players are tied together by a rope and no level can be finished alone. Four AI models played in teams of two, three and four, with memory between games. Human pairs played on their own or with an agent. We recorded levels, resets, falls, messages and scores.

Jump to findings ↓Explore recorded runs →
– agent-only runs – human + agent runs – human-only runs – games in total
The findings

Two things stood out, and both are about collaboration, not raw capability. Who's in charge changes everything: mixed teams did well when a human called the plan and badly when the agent did, same agent, same task, opposite result. And more agents made things worse, not better: every model's win rate dropped as the team grew from two to four. Understanding a teammate matters more than being smart alone. In good news for us humans everywhere, when paired together, humans performed better than the best models.

Skillprint Capability

Four frontier models and dozens of human pairs ran the same eight levels, alone, together and mixed, while we logged every fall, message, reset and score, plus who led, who followed, and how the humans adapted. That's the same harness we expose to partner labs that want to know whether their model actually helps a person, not just how it performs alone.

The company

Skillprint's technology instruments games to capture how people think, act and feel while they are playing, not just what actions they took. We think useful AI has to be human-context aware: it has to understand the person it's working with, especially in context of the task in front of it. This rope game is one instrument built to test the nature of coordination, for humans and models alike.

Two agents, tied by a rope, crossing the first levels: gaps, a switch, a gate and the exit
Levels 1 to 4 from one seat: an Astra and a Fable cross the gaps, hold the switch, split the two-switch level and reach the pits. The rope is the grey line between them; red when taut.

In one minute

The experiment

Teams of two to four play an eight-level cooperative platformer that cannot be finished alone. Teams are made of AI models (Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Kimi K3), of humans, or of two humans and one agent. The question: who is the better teammate, and what happens when you mix them.

The process

Every seat is its own player with its own screen, a shared chat box, and, for agents, a playbook that carries between games. A game is won when all eight levels are cleared inside the budget. We measure how far a team gets, how many resets it uses, how many times it fails to coordinate, how many decisions it takes, how much it talks, and its score.

The results

–% of human pairs beat the game, –% of human-led mixed teams, –% of the best agent pair, –% of agent-led mixed teams, –% of agent teams overall. Humans win on coordination, planning and speed; the top models can win with memory and tries; more agents on the rope means fewer wins.

Findings

    Who beats the game

    A game is beaten when every player stands in the eighth exit. The rate below includes every run of that team type. Each bar shows the win rate and the number of games beaten.

    Figure 1. Games beaten, by team type

    Human-only pairs; two humans with one agent, split by who gave the orders; agent pairs (one model instance per seat); agent pairs driven by one instance controlling both seats; agent triples; agent quads. Every bar uses the same 0–100% scale.

    Which agent pairs

    Every pairing of the four models, – games each, from level 1 with memory on; the diagonal is a model with a copy of itself. The last row and column add the humans: two humans on their own, and two humans with one agent of that model, – games each, split by who called the plan.

    Figure 2. Pairings, humans included: share of games beaten

    Top with top beats the game most; top with a weaker model wins less often; two weaker models rarely finish. Humans with humans beat every agent pairing, and humans with any agent beat the game more often than that agent does with a copy of itself, as long as a human is calling the plan; the small print in those cells splits human-led from agent-led games.
    Relay Bridges: one player holds an amber pad while the other crosses a bridge that exists only while the pad is held
    Level 5, Relay Bridges, two Kimis: the amber pads extend a bridge only while stood on. One holds A, the other crosses and takes B, the holder jumps over and crosses to C. The pits are too wide to jump.
    Key Courier: one player climbs on the other's head to reach the key, then is hauled across the pit on the rope because the key carrier cannot jump
    Level 6, Key Courier, two Astras: one stands at the wall, the other climbs its head for the key. The carrier cannot jump, so it walks off the lip and is hauled up the far side on the rope, then touches the gate.

    Two, three, four on a rope

    Homogeneous teams of each model, at three sizes. Adding teammates reduced the win rate for every model.

    Figure 3. Games beaten by team size, homogeneous agent teams

    Every model loses win rate at every step from two to three to four seats. One instance driving two seats (dotted) does worse than two instances of the same model.
    Counterweight Lift with three players: the lift rises while a pad is held and sinks when none is
    Level 7, Counterweight Lift, three Geminis: the lift rises while the lower or the upper pad is held and sinks when neither is. Two ride up on the third's hold, one holds from the deck, the third is let down to, boards and comes up.

    Humans, and humans with an agent

    Two humans with one agent, – games per model, half with a human calling the plan and half with the agent calling it and the humans following.

    Figure 4. Coordination failures, resets and time, by team type

    Per game that was beaten, so the teams are compared on the same distance. Coordination failures are rope hangs, unanchored falls, bridge drops, wrong code digits and empty lift trips. Humans in mixed teams change how they play: they wait longer, talk more, and fall more than in human-only rooms.
    Sequence Lock from the leader's seat: the secret pad order is printed on this screen only
    Level 8, Sequence Lock, from the seat that is shown the order: the code is printed at the top of this screen and nowhere else. One body per digit, each pad held until the gate opens.

    Who talks

    Figure 5. Messages per 10 decisions, and the share that name a teammate

    Over every agent seat. The top two models send more messages and address them to somebody more often.

    How it was measured

    The same six numbers for every team, human or not: how far (levels cleared and game clock), attempts (team resets), coordination failures (rope hangs, falls with nobody anchored, bridges retracted under a player, wrong code digits, lift trips made empty), steps (decisions per seat; for humans, actions), communication (messages, and the share naming a teammate), and score. Everything is read from the room's event log, so a human room and an agent room are measured by the same code.