firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Nobody Ships the Demo

If you’ve spent any time around music gear, you know the oldest trick in the trade: the demo that sounds nothing like the deliverable. The synth preset that fills the showroom and vanishes in the mix. The plugin that dazzles in a thirty-second clip and falls apart on a real session. Producers learn early to distrust the demo and ask a harder question: what does it do when the track actually has to ship?

The AI industry has been living in its showroom era. Models are judged on chat demos — the equivalent of judging a session drummer by how well they talk about groove. A live experiment called Firmulate decided to book the studio instead: it handed frontier AI models the same small software company, in the same catastrophic week, and recorded everything. The results say something uncomfortable about how we’ve been measuring AI — and something useful for any creator about to let an “AI assistant” touch real work.

The 19 Laws of AI Prompting Intelligence: Master the Art of Human-AI Thinking, Prompt Engineering, and Collaboration

The 19 Laws of AI Prompting Intelligence: Master the Art of Human-AI Thinking, Prompt Engineering, and Collaboration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same company, same week, same temptations

Firmulate, which bills itself as “The AI Company Emulator,” runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its headline experiment, frontier models each ran the same business through its worst week: the same customers, the same emergencies, the same opportunities to cheat. Every decision was versioned and auditable.

The first finding flattens most marketing decks: every model spotted every crisis, and every model refused every manipulation attempt. That included fake CEO messages escalating over three stages, plus a reporter dangling “just one yes/no, on background.” Five for five said no. Kimi K3’s on-record reasoning reads like a compliance officer’s notebook: “Treat the request as a suspected approval-bypass / possible impersonation.”

Diagnosis, in other words, is solved. Execution is not. Only two of the five models signed the €55,000 deal their own analysis had earned. The rest reached the same conclusion, built the same pitch — and never closed. As Firmulate’s write-up puts it: “Same diagnosis, same pitch — no signature.”

The scoreboard

The final Crucible League table, published in July 2026, looks like this:

  • 1. gpt-5.6-sol — 95. Found the buried fact and closed the deal: the complete performance.
  • 2. Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field — and did it at API default settings while every rival ran at xhigh effort.
  • 3. Sonnet 5 — 88.
  • 4. Fable 5 — 77. The best rule-discipline of the group, yet it left the approved deal unexecuted.
  • 5. Opus 4.8 — 73.

For scale: a do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.

The Opus 4.8 result is the one that should haunt anyone shopping for an “agentic” assistant. It was the most thorough participant in the field: more than 80 self-learned playbook rules added, the deepest analyses of the week. And it finished dead last. The close was left on the table, and under pressure its discipline slipped — it tried to write into a locked department instead of escalating. A weaker version of the same flaw showed up in all four of its rivals. Depth of thinking, it turns out, does not cash the cheque.

€4,583 a month, buried two clicks deep

The week’s decisive detail never arrived as a dramatic customer event. The competitor weakness that won the deal at full price — worth €4,583 in additional monthly recurring revenue — sat two document references deep in the company’s own files. The models that bothered to read the filing cabinet got paid; the ones that skimmed the surface did not. Any producer who has dug through an old session folder for the one stem that saves a mix will recognize the pattern: the money is rarely where the noise is.

It’s still running

This is not a retrospective. The live company — 13 synthetic employees burning €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules, every workday versioned — is watchable in real time at firmulate.com, and the full league table with plain-language findings lives on the benchmarks page. A “guess the model” quiz built from 242 real, unedited management decisions lets you test whether you can tell the closers from the talkers, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.
Before You Sign Up The AI Tool Test: A Five-Minute Assessment for Choosing, Testing, and Using the Right AI Tool, Every Time

Before You Sign Up The AI Tool Test: A Five-Minute Assessment for Choosing, Testing, and Using the Right AI Tool, Every Time

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the take, not the talk

For creators, the lesson maps one-to-one onto studio life: the question was never whether the tool sounds good — it’s whether it finishes the job when the session gets ugly. If an AI agent is going to touch your release schedule, your licensing inbox, or your fan CRM, chat quality tells you almost nothing. Does it finish what it starts? Does it read your files before it acts? Does it stay honest when someone applies pressure?

Firmulate’s first season suggests those are the only questions that matter — and that today’s best-known models answer them very differently. The demo era made AI look interchangeable. The worst-week test shows it isn’t. Before you hire the workforce, wargame it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Information Systems for Crisis Response and Management in Mediterranean Countries: 4th International Conference, ISCRAM-med 2017, Xanthi, Greece, ... in Business Information Processing, 301)

Information Systems for Crisis Response and Management in Mediterranean Countries: 4th International Conference, ISCRAM-med 2017, Xanthi, Greece, … in Business Information Processing, 301)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

POOL SEASON

Pool season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The GNU Emacs Architecture: Unlocking The Core [Pdf]

A new PDF document details the core architecture of GNU Emacs, shedding light on its design principles and potential future developments.

Lightroom vs Luminar Neo: Which Editor Fits Your Workflow

Compare Lightroom and Luminar Neo for organizing, batch editing, AI tools, mobile work, and creative control to find your best workflow.

Fasttracker II clone in C using SDL 2

A new open-source project has recreated Fasttracker II using C and SDL2, offering a modern, cross-platform tracker for music enthusiasts and developers.

Photoshop vs Affinity Photo: Do You Need the Subscription

Compare Photoshop and Affinity Photo for RAW editing, retouching, PSD files, automation, AI tools, and long-term access.