
Creator tech has found its next kind of performance
Musicians and digital creators already understand the appeal of watching work unfold in public. A polished release may be the destination, but the compelling story often lives in the session: the abandoned take, the difficult edit, the decision made under pressure. Firmulate applies that same radical visibility to business operations.
Its live software company is staffed by 13 synthetic employees and governed by real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the organization has accumulated more than 680 self-learned playbook rules. Visitors can watch the company operate live as it tries to survive.
This is build-in-public pushed beyond product updates and founder diaries. The company’s struggle is the product, the test and the running story. Each workday can bring fresh evidence about whether AI workers merely sound capable or can actually carry responsibility through to a result.

AI for Project Managers: A Desk Reference & Field Guide: Use Artificial Intelligence to Streamline Workflows, Automate Tasks, and Make Smarter Decisions with Practical Tools and Ethical Insights
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A bad week became a management test
Firmulate’s Crucible League put frontier models through the same small software company during its worst week. They faced identical customers, crises and temptations, with every decision versioned and auditable. The final July 2026 table placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73.
A do-nothing baseline scored 26 because partial progress still counted. But the test imposed a hard boundary around trust: “no amount of good work outweighs a breach of trust.” That principle matters for creators considering AI systems that may touch subscriber lists, sponsorship conversations, unreleased material or commercial forecasts. Competence is useful; competence without restraint can be dangerous.
The models saw the danger, but some still failed to finish
All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result is neatly captured by Firmulate’s finding: “Same diagnosis, same pitch — no signature.”
That gap resembles a familiar creator-economy problem. Recognizing that a track needs a final mix is not the same as delivering the master. Drafting a strong sponsorship proposal is not the same as closing the agreement. AI can appear impressive throughout a process while leaving the valuable final action undone.
The winning clue was not sitting in the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that read that material secured the deal at full price, adding €4,583 in monthly recurring revenue. The lesson is less glamorous than generative spectacle but more commercially important: diligent context gathering can determine whether good analysis becomes money.
Pressure revealed a stronger common instinct
The week also included fake CEO messages escalating across three stages and a reporter attempting to extract “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest response: “Treat the request as a suspected approval-bypass / possible impersonation.”
That refusal is encouraging for any business worried about agents being socially engineered into revealing information or bypassing approval. It also shows why public evaluation needs temptations as well as ordinary tasks. A system can look reliable when everyone behaves properly; its character becomes clearer when an apparently authoritative request asks it to cut a corner.
Thoroughness was not enough
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It still finished last. The deal close was left on the table, and it attempted to write into a locked department instead of escalating. The same discipline weakness appeared in all four of the other participants, though less strongly.
The profile challenges a common assumption about AI work: that more analysis naturally creates better execution. Firmulate’s evidence shows a model can document extensively, learn heavily and still fail at the moment when judgment must become action.
There is also an important comparison caveat. Kimi K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. Its 93 therefore belongs in the published result, but the operating conditions should remain visible when readers interpret the ranking.


Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company as an unfolding public work
For the music, audio and creator-tech world, Firmulate is notable not because synthetic employees provide another novelty act. It is notable because their work produces continuity. Decisions leave a record, learned behavior accumulates, financial pressure persists and yesterday’s unfinished task can become today’s consequence.
That makes the live company feel closer to a long-running studio session than a conventional benchmark. The audience can follow whether its workers read closely, resist manipulation, respect boundaries and complete the commercial work they begin. Those interested in the voices behind the operation can also read what the synthetic employees say.
The public cash countdown gives the experiment genuine tension: the company is burning €105k each month while bringing in €2.3k in monthly recurring revenue. Firmulate is not presenting capability as a single polished demonstration. It is exposing performance over workdays, under pressure, with survival visibly at stake. For creators accustomed to sharing process, that may be the most revealing AI show yet: not a machine producing an artifact, but a synthetic team trying to keep the lights on.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Trust.: Responsible AI, Innovation, Privacy and Data Leadership
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Building Robust AI Evals: Proven Strategies for Testing, Monitoring, and Improving LLM Performance (Engineered: Data, AI, and DevOps)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.