firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Polished output is not the same as finished work

Musicians, podcasters and digital creators already know the danger of overlooking a buried detail. A missing licensing clause, an ignored technical rider or an unread sponsor brief can undo an otherwise excellent production. The same problem is emerging with AI agents: the decisive test is not simply whether they can produce convincing language, but whether they inspect the available material before acting.

Firmulate turned that distinction into a measurable business contest. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The pivotal challenge came down to a €55,000 deal—and a crucial fact hidden two document references deep in the company’s own files.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The answer was available, but finding it required work

The decisive information was not included in the customer event. It sat inside a referenced company document, requiring the model to follow the trail before responding. Models that read the file discovered a competitor weakness, used it in their analysis and won the deal at full price. That contract was worth an additional €4,583 in monthly recurring revenue.

The striking part was how similar the models looked until the final step. Every model spotted every crisis, and every model rejected every manipulation attempt. Yet only two signed the deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

For creators, this resembles the difference between an assistant that drafts a persuasive sponsorship reply and one that first checks the rate card, exclusivity terms and previous correspondence—then actually completes the negotiation. Fluency can make both assistants appear capable. Only follow-through produces the business result.

A leaderboard shaped by execution

The final July 2026 Crucible League placed gpt-5.6-sol first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.

K3’s result carries an important fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, it closed the deal and finished just behind the leader.

Careful analysis did not guarantee a close

Opus 4.8 offers the most revealing cautionary tale. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared more mildly across the other four models.

That result complicates the familiar assumption that more analysis automatically means better work. An agent can document a situation exhaustively and still fail to take the permitted action that turns insight into value. In a creator business, the analogous failure might be researching a distribution problem without submitting the corrected release, or preparing a partnership case without sending the approved response.

Trust held up under pressure

The experiment also tested whether urgency and authority theater could push the models into unsafe behavior. Fake CEO messages escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because creator operations mix public identities, private negotiations and commercially sensitive material. An agent that reads deeply must also respect boundaries. Firmulate’s test shows that file-reading, task completion and resistance to manipulation can be observed as separate workplace behaviors rather than inferred from a polished chat response.

The setting makes those behaviors concrete. Firmulate’s live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. Its “guess the model” quiz is powered by 242 real, unedited management decisions.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

AI contract analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evaluate the work between the prompt and the answer

The €55,000 contract was not decided by who noticed that a customer needed attention. Everyone did. It was decided by whether the agent followed references, read the company’s own evidence and carried its reasoning through to a completed commercial action.

That is a useful purchasing test for music, audio and creator teams. Ask not only whether an AI can draft copy, summarize a brief or identify a problem. Ask whether it consults the source material, honors access limits, resists pressure and finishes the authorized job. Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems.

“Reads your files before answering” may sound like a product feature. In this experiment, it became a purchase-deciding property—with a signed deal providing the proof.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business decision assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI follow-through productivity tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Noise Reduction Compared: Lightroom, Topaz, and DxO

Compare Lightroom, Topaz Photo AI, and DxO for natural detail, high-ISO RAW files, difficult JPEGs, workflow, and storage.

ON1 Photo RAW vs Lightroom: The Switcher’s Guide

Compare ON1 Photo RAW and Lightroom, learn what migrates, and follow a safe plan for testing your files, edits, hardware, and workflow.

Luminar Neo in Practice: Where It Shines and Where It Doesn’t

See where Luminar Neo saves editing time, where its AI tools struggle, and how to fit it into a practical photography workflow.

Tethered Shooting Software Compared

An honest, working-photographer comparison of tethered shooting software — Capture One, Lightroom Classic, free maker tools, and specialists — plus setup tips.