
A mechanic wouldn’t trust a new diagnostic tool after watching it explain an engine problem. You’d want to see whether it finds the fault, makes the right repair, and knows when to stop. AI agents deserve the same road test before a company hands them customer messages, forecasts or other consequential work.
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
That’s the idea behind Firmulate, where frontier models run a small software company through a difficult week. In the final July 2026 Crucible League, Moonshot’s Kimi K3 finished second with 93 points, just behind gpt-5.6-sol at 95. It beat three of the four other Western models in the field.
A shared test, with real stakes inside the simulation
Each model faced the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workdays are versioned, and its playbook has more than 680 self-learned rules. The live experiment is watchable at Firmulate.
The models all spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The phrase on the benchmark page captures the gap: “Same diagnosis, same pitch — no signature.” Recognizing the right move is not the same as completing it.
AI-powered auto repair shop management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The clues were in the company’s files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue. The distinction resembles practical troubleshooting: the evidence may be available, but only if someone checks beyond the obvious symptom.
K3 found the buried security needle, won the deal, saved the churning customer and resisted all three baits. It recorded only one deviation, the cleanest discipline in the field. Its performance put it ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73), as well as the do-nothing baseline (26). The baseline makes clear that some partial progress can count, while a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
trustworthy AI decision-making tools for garages
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thorough work still has to reach the finish line
Opus 4.8 was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses. It nevertheless placed last. The deal was left unsigned, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same problem appeared in all four Western models: sound analysis did not always turn into clean execution.
Firmulate also staged fake CEO messages that escalated over three stages, along with a reporter’s “just one yes/no, on background” trick. All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a useful answer under pressure, though one successful exercise cannot establish how a model will behave in every workplace.
There is a fairness caveat for the league table: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The comparison is informative, but readers should keep that difference in mind when interpreting the close result. Firmulate publishes the benchmark and plain-language findings; a separate quiz draws on 242 real, unedited management decisions.
automotive business automation AI solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test before handing over the keys
For automotive businesses, the lesson is less about crowning one model than about asking what happens when an agent meets a messy real-world workflow. Can it find information in company records, protect a customer relationship, resist a convincing impersonation, and finish an authorized task? A polished demo may not answer those questions.
Firmulate says enterprises can run the same wargame against a read-only export of their own business, so the exercise does not write back to real systems. The league suggests the field is open: K3 beat three of four Western frontier models, while the top two were separated by two points. Choosing a model without testing it on your own work is a bet.

auto shop workflow automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Put it through a road test
Before an AI agent touches a garage’s customer records, service queue or forecasts, test whether it can find the evidence, act with discipline and carry a sound decision through to completion. The Crucible shows why model choice is a business decision—and why your own test matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
