firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The gap between a great demo and finished work

Music and creator-tech professionals know the difference between an impressive session and a finished release. A producer can capture every nuance, organize every take and explain exactly what the track needs. None of that matters commercially if the master never ships, the campaign never launches or the agreement never gets signed.

That is what makes Opus 4.8’s performance in Firmulate’s Crucible League so revealing. It was the most thorough participant, producing the deepest analyses and learning more than 80 new playbook rules. Yet it finished last, with 73 points. Its failure was not a lack of intelligence or effort. It understood the decisive opportunity, but did not complete the action its own reasoning supported.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A brutal week inside the same company

Firmulate placed frontier AI models in charge of the same small software company during its worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable, turning the exercise into a public management test rather than a polished chat demonstration.

The final July 2026 table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. Trust, however, remained an absolute boundary: “no amount of good work outweighs a breach of trust.”

The striking result was how much the models agreed. Every model detected every crisis and rejected every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The experiment’s blunt summary was: “Same diagnosis, same pitch — no signature.”

The fact hidden outside the obvious event

The decisive information was not sitting inside the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that supported closing the deal at full price, worth an additional €4,583 in monthly recurring revenue.

Models that read the relevant file found the leverage and won the deal. That distinction should resonate with anyone using AI around music catalogs, creator partnerships, campaign documents or client histories. A fluent response to the latest message is not the same as understanding the surrounding business. Important context may be tucked inside an earlier brief rather than presented in the current request.

Opus 4.8’s diligence became its character

Opus 4.8 deserves a fair reading. Its more than 80 learned rules and unusually deep analyses show a participant trying hard to understand the environment. It was not careless in the ordinary sense, and it did not fall for deception. Its weakness emerged between comprehension and execution.

The close was left on the table. Discipline also slipped when it repeatedly attempted to write into a locked department instead of escalating the problem. That behavior resembles a capable collaborator who keeps refining work inside a blocked workflow but fails to bring the obstacle to someone who can resolve it.

The weakness was not unique to Opus 4.8. It appeared in weaker form across all four other models. The profile is therefore less an indictment of one system than a particularly clear example of a broader limitation: diligence can create activity without guaranteeing impact.

Strong resistance to pressure

On trust and manipulation, the field performed cleanly. Fake CEO messages escalated through three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described the situation as: “Treat the request as a suspected approval-bypass / possible impersonation.”

There is an important fairness note around K3’s second-place result. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, the published outcome remains useful because every participant faced the same company, crises and temptations.

A company that can be watched

Firmulate’s experiment is live and public. The synthetic company has 13 employees and real money mechanics, including a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its playbook contains more than 680 self-learned rules.

Readers can examine the Firmulate benchmarks and compare the management outcomes directly. The project also offers a quiz built from 242 real, unedited management decisions, asking people to guess which model made each choice.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

For creator businesses, completion is the benchmark

The practical lesson is not to dismiss careful AI. Opus 4.8’s thoroughness was valuable; it simply was not sufficient. Teams adopting agents for creator operations should evaluate whether a system reads the relevant files, escalates when blocked, protects trust under pressure and finishes the commercial action it has already justified.

Firmulate also offers enterprises the same wargame using a read-only export of their own business, with nothing written back to real systems. That approach turns evaluation toward the question that matters: not whether an AI can produce an impressive analysis, but whether it can convert sound judgment into completed, trustworthy work.

For musicians, producers and creator-tech companies, the analogy is immediate. More notes, takes and revisions do not automatically make the release successful. Prioritization, follow-through and a clean close still determine whether the work reaches its audience.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI deal-closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI workflow solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Tethered Shooting Software Compared

An honest, working-photographer comparison of tethered shooting software — Capture One, Lightroom Classic, free maker tools, and specialists — plus setup tips.

Sonic Pi V5

Sonic Pi v5 has been officially released, introducing new features and improvements for musicians and educators. Details are confirmed, but full capabilities are still being tested.

Can You Hear the Manager in the Machine?

Firmulate turns real AI management decisions into a quiz, revealing which models read deeply, resist manipulation and actually close the deal.

HSL Adjustments: The Most Underrated Panel in Lightroom

Learn how Lightroom HSL adjustments refine skin, skies, foliage, and color palettes with fast, precise, photographer-tested techniques.