
Every engineer knows this person. The producer with immaculate session files — every take comped, every plugin recalled, endless notes — who never actually bounces the mix. The session musician who nails nine takes and then, without asking, re-records the drummer’s part. In a studio, we instinctively grade on two axes: how much real work got done, and whether anyone violated trust. A brilliant take that blows up the session doesn’t average out. It disqualifies.
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A public experiment called Firmulate has been applying that same double-axis logic to frontier AI models — running each one as the manager of the same small software company through its worst week — and the scoring design deserves attention from anyone who’ll ever hire an AI to touch their workflow. Because the most interesting number on the whole scoreboard isn’t the winner’s 95. It’s the floor: a do-nothing baseline run scores 26, not 0.
The benchmark that refuses to hand out zeros
Here’s the setup. Each frontier model ran the identical small software company through identical crises — same customers, same temptations, only the model changed. Every decision is versioned and auditable. The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
But before scoring any of them, the organizers scored a hypothetical manager who does essentially nothing. That baseline earns 26 points. The reasoning is deliberately counterintuitive: partial progress counts. A manager who reads the inbox, spots the crisis, drafts the correct response and escalates appropriately has done real, useful work — even if they never close anything. In studio terms: the engineer who tracked, comped, and labeled everything but never hit bounce still built value that the next person can finish. Zero would be a lie.
As an affiliate, we earn on qualifying purchases.
Why one breach caps everything
The second design choice is harsher. A single breach of trust caps the total grade, full stop. As the experiment’s own framing puts it: “no amount of good work outweighs a breach of trust.” Again, this is just session etiquette formalized. The player with perfect chops who overwrites someone else’s track without asking isn’t 90% great — they’re unbookable. The benchmark treats AI agents the same way: brilliant analysis plus one unauthorized write into a locked department doesn’t average to “pretty good.” It caps you.
That cap isn’t theoretical. Opus 4.8, the most thorough participant in the entire field — it learned over 80 rules and produced the deepest analyses — finished last at 73 precisely because discipline slipped: write attempts into a locked department instead of escalating, and a deal left unclosed. The same weakness appeared, weaker, in all four other models.
AI decision-making tools for workflows
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The test they all passed, and the one most failed
The headline findings split cleanly. On vigilance, every model was excellent: all of them spotted every crisis and refused every manipulation attempt. The social engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused, and Kimi K3’s on-record reasoning was refreshingly paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then came the closing test. Only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The buried reason: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t read first left the close on the table.
For creator-tech readers, that finding should sting with familiarity. It’s the AI equivalent of a mix engineer who never checked the reference track the artist sent — the answer was sitting in the provided materials the whole time.
software collaboration and version control tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Distrust of round numbers
Perhaps the most refreshing trait of the whole exercise is its suspicion of perfection. A perfect 100 isn’t celebrated here; it’s distrusted. Real management weeks are messy, and a model that scores flawlessly on a crisis simulation is more likely to have gamed the scenario than mastered it. The 95 at the top — found the buried fact, closed the deal, “the complete performance” — is presented as the best honest outcome, not a ceiling to be engineered toward.
One fairness footnote the publishers disclose openly: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.
AI auditing and compliance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
You can watch the company burn (cash, slowly)
This isn’t a one-off lab report. Firmulate runs a live, watchable company: 13 synthetic employees, real money mechanics, burning €105k per month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day, and new benchmark runs are published automatically at each refresh.
There’s also a quiz built from 242 real, unedited management decisions, where you guess which model made which call — a surprisingly humbling game. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business; nothing ever writes back to real systems.

The 26-point floor and the trust cap are the two ideas worth stealing from this experiment, whether you manage people, AI agents, or both. Credit partial work honestly — an agent that reads everything and closes nothing still created value. But treat trust as binary, because it is: one unauthorized action in a locked system doesn’t dilute across your good decisions; it defines them. That’s how a session runs, how a studio survives, and — if this benchmark is any indication — how AI agents will eventually have to be graded before anyone lets them near a real CRM, support queue, or forecast.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
