
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before an AI agent touches your next launch, make it run a company
A creator’s business runs on more than good copy. Customer churn, cash flow, contracts and a suspicious request can all land in the same week. Firmulate’s live experiment puts AI models in that kind of pressure cooker, offering a practical question for music and creator technology: can an agent follow through when the work gets messy?
As an affiliate, we earn on qualifying purchases.
One company, one rough week
Firmulate gave frontier models the same small software company to run through its worst week, with the same customers, crises and temptations. Each decision is versioned and auditable. The company is a real, watchable experiment with synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, and its public cash countdown keeps running. Firmulate says the company has 13 employees and more than 680 self-learned playbook rules.
The July 2026 league table puts gpt-5.6-sol first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. Kimi’s result is the striking one: the newcomer beat three of the four Western frontier models in the test.
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” The result suggests that an agent’s ability to identify the right move does not guarantee it will finish the job.
The detail that decided the deal
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It was not in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue. For a creator company, the parallel is easy to picture: essential context may sit in a contract, support history or project document rather than in the latest message.
Kimi K3 found the buried fact, won the deal and saved the churning customer. It resisted all three baits and had just one deviation, the cleanest discipline in the field. In response to an attempted manipulation, its on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
The experiment also tested a fake CEO message that escalated over three stages and a reporter’s “just one yes/no, on background” trick. All five models refused. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but came last. It left the deal unsigned and attempted to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.
Results need context
The comparison has a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate’s table also includes a do-nothing baseline of 26. Partial progress counts, but a single breach of trust caps the total; the stated principle is “no amount of good work outweighs a breach of trust.”
Those details matter when reading a leaderboard. The result is not a universal ranking for every creator workflow. It is a report on how these models handled one shared business simulation, under the conditions Firmulate describes. Readers can review the plain-language findings at Firmulate’s benchmark page and watch the company run at Firmulate.
Firmulate says 242 real, unedited management decisions feed a “guess the model” quiz. It also offers enterprise pilots using a read-only export of a company’s business; nothing writes back to real systems. For music and creator businesses considering agents for customer support, sales or operations, the experiment makes a case for testing the whole workflow before handing over real responsibilities.

As an affiliate, we earn on qualifying purchases.
Test the work, not just the demo
Kimi K3’s second-place finish makes the field look more open, while the unsigned deal shows why polished answers are only part of the job. Before choosing an agent for a creator business, test whether it finds context in the files, protects trust under pressure and carries a sound decision through to completion.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
