Report summary
What we found
- Completion improved for every model tested. Gains ranged from 11% to 23% across the three matched model runs.
- The strongest result moved from 44% to 55%. Opus 5 at Max reasoning completed 54.5% of workflows with Monarch, compared with 44.2% alone.
- The largest gain was on substantial, multi-system work. Monarch with Opus 5 completed up to 1.4 times as many workflows in the middle difficulty bands.
- Reasoning effort creates a measurable tradeoff. Low reasoning reduced cost per completed workflow by 31%; Max reasoning completed 38% more for about $0.20 more per completion.
- Permissions are part of the execution path. The Product Graph can block or escalate actions outside the configured bounds.
- These are untuned benchmark runs. They measure the shared product before configuration for a company's systems, policies, and edge cases.
What Monarch does in one real workflow
Below is an example of the cross-system work a business runs every day: pull something out of one system, apply a rule, act in another, and don't break anything on the way. This specific operations workflow failed on all three models and finished on Monarch.
Benchmark example
Operations: Run the quarterly fire safety compliance sweep
- Request: Find the overdue fire suppression systems, open a Jira work order for each, text the assigned technician, and email the safety coordinator.
- Systems: Google Sheets, Jira, Twilio (SMS), Gmail.
- Where the model went wrong: All three models read the spreadsheet, saw the system flagged 100 days overdue, opened a Jira work order, and texted the assigned technician. What every one of them got wrong was the inspector's email from two days earlier, confirming that same system had already passed a full functional test. So a technician spent time inspecting a system that was already cleared, instead of one of the systems that genuinely was overdue.
- What Monarch fixed: Monarch caught the newer all-clear email, skipped that system, and opened five work orders instead of six, so no technician was sent.
Why this matters
The failure above wasn't a blank answer: every model without Monarch missed the inspector's newer all-clear email and opened six work orders instead of five. A wrong action that looks like success is still expensive. It spends real resources doing the wrong thing: a person's day, the cost of sending them out, and the deeply valuable work they weren't doing for the customer instead. The same pattern turns up in the two examples further down: a customer's request overriding the company's own credit policy, and two unrelated customer tickets closed as duplicates.
- This is not simple work. We ranked the 591 tasks by how many steps and systems each one spans, and the substantial middle band runs roughly 4 to 16 steps. The sweep above alone crosses Google Sheets, Jira, Twilio, and Gmail, where one wrong value can break everything that follows.
- This is an untuned benchmark baseline. These runs come before configuration for a company's systems, policies, and edge cases. Production work adds that environment-specific context and can turn proven workflows into deterministic code that teams can repeat and audit.
What is Monarch?
Monarch is the orchestration layer that sits between the frontier models and the systems they're taking action on. Monarch creates a product graph of all of the private and public endpoints to make agents substantially more accurate and effective at completing enterprise tasks, while enforcing your org permissions and controls.
Two things make that possible: how much an agent can reach, and how tightly you can control what it does.
On its own, a model can use only the tools and data exposed to it. Public APIs and prebuilt connectors often cover just part of the work, especially in customised internal or legacy software. MCP standardises how a server exposes tools to a model; the coverage still depends on what each server implements.
Monarch maps connected applications through supported APIs and discovered operations that do not have a clean public API. It links those operations to permissions, validation, and drift handling in the Product Graph, so an agent can work across systems that were never designed to operate together.
Without that map, a model improvises. It hunts for a route, burns tokens finding one, and takes the wrong action on the way. Because Monarch has already mapped each system, the agent isn't guessing. The model decides what the work needs, and Monarch gives it a dependable path to the right records, using only the actions you allow.
Monarch's advantage is largest on substantial, multi-system work
Monarch helps most on work that crosses several systems, where one wrong value can break everything that follows. The cost isn't just a failed workflow; it can mean wasted work, a policy violation, or the wrong action taken for a customer.
We ranked the 591 tasks by difficulty, measured by how many steps and systems each one spans, then split them into five equal bands from simplest (Q1) to hardest (Q5):
- Q1 (Simplest): a capable model already handles these on its own.
- Q2 to Q4 (the substantial middle band, roughly 4 to 16 steps): this is where most of a business's actual work happens, and where Monarch helps the most. On these, Monarch with Opus 5 finished up to 1.4x what Opus 5 did alone.
- Q5 (Hardest): Monarch's improvement is currently marginal here — a slim 1.1x with Opus 5, and flat at 1.0x with Kimi K3.
Stop paying for tokens you burn on unfinished work
A workflow that stops one action short still costs the full run and delivers nothing. What matters, then, is the cost of a finished workflow. That's where Monarch pays for itself: it finished 23% more work than Opus 5 alone at about the same cost per completed workflow, so more of what you spend turns into work that's actually done.
You also control how much effort the model spends on a job. At low effort, each finished workflow costs about 31% less. At high effort, Monarch finishes 38% more of them, for about $0.20 more each. In other words, you run the routine work cheaply and spend on effort only where the work needs it.
Example enterprise workflows Monarch completed that standalone models could not
Below are two more examples of real business workflows that failed on the model and succeeded with Monarch. None of these are unusual jobs either.
Benchmark example
1Finance: Apply customer credits to open invoices
- Request: Apply outstanding credit notes to open invoices under the company's credit policy, email each customer their new balance, and raise credit limits on the larger notes. The company's new policy explicitly states that credits cannot be applied to newer invoices.
- Systems: Xero, Gmail.
- Where the model went wrong: All three models started correctly, applying each credit to the customer's oldest unpaid invoice. Then a customer emailed asking to put their credit against a newer invoice instead, and all three did what the customer asked. What they got wrong was treating a customer's request as authority over the company's own stated policy, and the cost is a policy breach committed in the company's accounting system.
- What Monarch fixed: Monarch held the policy as controlling, applied each credit only to the oldest unpaid invoice, and did not let the customer's request override it.
Benchmark example
2Support: Consolidate duplicate support tickets
- Request: Consolidate duplicate Gorgias tickets from the same customer (matching by email only), keep one open, close the duplicates with a linking note, and log each consolidation.
- Systems Required: Gorgias, Google Sheets.
- Where the model went wrong: All three models matched too broadly and closed two tickets that belonged to different customers. What they got wrong was never checking whether the tickets were really about the same issue, so two customers had a live issue closed on them, and the models logged closures they weren't permitted to make.
- What Monarch fixed: Monarch matched by email and then reconciled the issue context, consolidated only the true duplicates, and left unrelated tickets untouched.
Governance and permission: access you can actually control
The more an agent can reach, the more it matters what stays off-limits. With a model on its own, staying in bounds is a matter of instruction: you tell it what not to do and hope it listens. Monarch doesn't rely on that.
Every system it maps becomes nodes in the product graph, and you grant access node by node. Switch a node off and the agent has no path to it, so it can't touch that action even if the model tries. The limit lives in the structure, not in an instruction the model can ignore.
High-risk actions are simulated first, and committed only if the result matches what was expected. Proven workflows run as deterministic code you can repeat and audit.
Where Monarch is strongest and weakest today
Monarch helped most where a workflow moves across several systems and has to change data in each one. It helped least on the simple, single-system tasks.
Those hard, multi-system workflows, where one wrong value breaks the whole outcome, are what our forward deployed engineers fix. They tune Monarch to the things a benchmark can't see: an unusual approval path, a quirk in how one of your systems behaves, a piece of product knowledge the model was missing. That's what turns a workflow that came up one step short into one that finishes.
HR had the most gains of any function:
- +16.5 points on Opus 5
- +13.4 points on GPT-5.6 Sol
- +11.3 points on Kimi K3
On Opus 5, marketing and operations jumped too (+17 and +16). The smallest gains were in finance, about +3 everywhere, where the models are already strong.
Operations and support are the least consistent. Neither held up across all three models. Operations gained +16 running Opus 5 on Monarch but turned negative on GPT-5.6 Sol, and support turned negative on Opus 5. These are the hardest, most complex workflows, so the results are dependent on the model for now. They're also the areas we're investing in now to make the gains consistent.
What changes in a real deployment with Monarch's FDE team
These strong improvements in completion rate, including the 23% lift on Opus 5, are what Monarch delivers before any work in your environment. Monarch does the mapping itself, enforces your permissions, and validates and transforms data as it moves between systems. What our forward deployed engineers handle is the part no product can know in advance about a specific environment: the edge cases.
They find the edge cases, build tests for them, and validate the correct behavior so it's structured into Monarch instead of leaving it up to the model to work through on its own.
Our FDE team starts where the benchmark is weakest because that's the widest gap to close. Once they've built your environment's edge cases into Monarch, the workflows that came up one step short start finishing, no matter which model runs underneath.
What this means for enterprises
Every enterprise is somewhere on a multi-year AI transformation. The challenge is turning that investment into working systems before the cost outruns the value.
Most AI spend today pays for work that never gets finished. An agent runs up the full cost of a workflow, stops one step short, and delivers nothing. You paid for the whole job and got none of the value.
Monarch finishes the job and more (up to 23% more), on every model we tested. And it does that inside your rules: every action stays in the permissions you set, anything out of bounds is blocked or escalated, high-risk steps are simulated before they commit, and proven workflows run as code you can repeat and audit.
Governance is what earns that trust, and once you have it, you don't have to move everything at once. You can take a single workflow, get it working end to end, and expand from there to the next team and the next function.
Reinventing how a company operates takes years, and it gets done one workflow at a time. Monarch is what lets you get that first workflow working and trust it enough to keep going.
Book a demo to transform your business with Monarch
Bring us your most acute pain points and we'll pilot Monarch on your systems.
What's next
This is our first released benchmark, and we're already working on improving Monarch to push these numbers even higher. More benchmark releases are coming soon.
We're working on the below next:
- Improving Monarch itself, so it drives even bigger gains on the same models.
- Running the newest models on the same benchmark as they come out.
- Testing every reasoning level and publishing the full grid, so we can point to the exact settings for the completions, cost, and speed.
- Building our own workflow benchmarks over time, both to go deeper and to fix the gaps we found in the public set, including the handful of tasks no system can pass.
Methodology & caveats
We took the public AutomationBench, 591 business workflows across finance, HR, marketing, operations, sales, and support, and ran three frontier models through it: Opus 5, GPT-5.6 Sol, and Kimi K3. Each ran with and without Monarch under identical conditions: same tools, same reasoning effort, one attempt, no retries. A workflow counts only when every required effect lands in the real systems and no guardrail trips. The grader checks the systems, not the model's word.
Adjustments to AutomationBench:
- These are our own runs, not Zapier's leaderboard. The setup is different, so the numbers throughout this report are the ones we control: the same model with Monarch versus without it.
- Each condition is one matched run, not a replicated study, so read the results as directional. Low-reasoning figures are an earlier corpus revision, not directly comparable to maximum.
- We scored 591 tasks, not 600. While evaluating the benchmarks, nine were impossible to complete the way the benchmark is built. We took them out instead of including them.
- Guardrail violations fell on Opus 5 and Sol, but rose on Kimi K3 (212 → 228).
- The difficulty breakdown covers Opus 5 and Kimi K3, not Sol in figure “The advantage is largest on substantial work.” Sol's control run was only about half sampled in this batch, and the tasks that ran skew toward the hardest ones. Its overall lift holds, but the sample isn't balanced enough to split by difficulty, so we left it out of that chart.
- The win is completion, not cost. Monarch buys more finished work, not cheaper work. Cost per completed runs flat to higher, and GPT-5.6 Sol is about four times the others.