Report summary

What we found

  • Completion improved for every model tested. Gains ranged from 11% to 23% across the three matched model runs.
  • The strongest result moved from 44% to 55%. Opus 5 at Max reasoning completed 54.5% of workflows with Monarch, compared with 44.2% alone.
  • The largest gain was on substantial, multi-system work. Monarch with Opus 5 completed up to 1.4 times as many workflows in the middle difficulty bands.
  • Reasoning effort creates a measurable tradeoff. Low reasoning reduced cost per completed workflow by 31%; Max reasoning completed 38% more for about $0.20 more per completion.
  • Permissions are part of the execution path. The Product Graph can block or escalate actions outside the configured bounds.
  • These are untuned benchmark runs. They measure the shared product before configuration for a company's systems, policies, and edge cases.

What Monarch does in one real workflow

Below is an example of the cross-system work a business runs every day: pull something out of one system, apply a rule, act in another, and don't break anything on the way. This specific operations workflow failed on all three models and finished on Monarch.

Benchmark example

Quarterly fire safety compliance sweep Google Sheets shows one fire suppression system as 100 days overdue, but a newer Gmail message says it passed a full functional test two days ago. The workflow reconciles both signals, skips that cleared system, and acts on the five systems that remain overdue by opening Jira work orders, texting assigned technicians through Twilio, and emailing the safety coordinator. Three models without Monarch missed the newer email and created six work orders; with Monarch, five work orders were created and no technician was sent to the cleared system. OPERATIONS WORKFLOW Quarterly fire safety compliance sweep 1 · READ ALL SIGNALS Compliance tracker Suppression system 100 days overdue Inspector email Full functional test passed 2 days ago · newer evidence 2 · RECONCILE Which is current? NEWEST VERIFIED STATUS WINS 3 · DECIDE Still overdue? 5 systems: YES 1 SYSTEM CLEARED Skip · no technician sent 5 DUE 4 · ACT + NOTIFY Jira5 high-priority work orders Twilio SMSText each assigned technician GmailEmail the safety coordinator ONLY CONFIRMED OVERDUE SYSTEMS WITHOUT MONARCH · 3/3 MODELS Missed the newer all-clear → 6 work orders + unnecessary dispatch WITH MONARCH Reconciled both data sources → 5 work orders + cleared system skipped

Operations: Run the quarterly fire safety compliance sweep

  • Request: Find the overdue fire suppression systems, open a Jira work order for each, text the assigned technician, and email the safety coordinator.
  • Systems: Google Sheets, Jira, Twilio (SMS), Gmail.
  • Where the model went wrong: All three models read the spreadsheet, saw the system flagged 100 days overdue, opened a Jira work order, and texted the assigned technician. What every one of them got wrong was the inspector's email from two days earlier, confirming that same system had already passed a full functional test. So a technician spent time inspecting a system that was already cleared, instead of one of the systems that genuinely was overdue.
  • What Monarch fixed: Monarch caught the newer all-clear email, skipped that system, and opened five work orders instead of six, so no technician was sent.

Why this matters

The failure above wasn't a blank answer: every model without Monarch missed the inspector's newer all-clear email and opened six work orders instead of five. A wrong action that looks like success is still expensive. It spends real resources doing the wrong thing: a person's day, the cost of sending them out, and the deeply valuable work they weren't doing for the customer instead. The same pattern turns up in the two examples further down: a customer's request overriding the company's own credit policy, and two unrelated customer tickets closed as duplicates.

What is Monarch?

Monarch is the orchestration layer that sits between the frontier models and the systems they're taking action on. Monarch creates a product graph of all of the private and public endpoints to make agents substantially more accurate and effective at completing enterprise tasks, while enforcing your org permissions and controls.

Two things make that possible: how much an agent can reach, and how tightly you can control what it does.

On its own, a model can use only the tools and data exposed to it. Public APIs and prebuilt connectors often cover just part of the work, especially in customised internal or legacy software. MCP standardises how a server exposes tools to a model; the coverage still depends on what each server implements.

Monarch maps connected applications through supported APIs and discovered operations that do not have a clean public API. It links those operations to permissions, validation, and drift handling in the Product Graph, so an agent can work across systems that were never designed to operate together.

Without that map, a model improvises. It hunts for a route, burns tokens finding one, and takes the wrong action on the way. Because Monarch has already mapped each system, the agent isn't guessing. The model decides what the work needs, and Monarch gives it a dependable path to the right records, using only the actions you allow.

Monarch's advantage is largest on substantial, multi-system work

Monarch helps most on work that crosses several systems, where one wrong value can break everything that follows. The cost isn't just a failed workflow; it can mean wasted work, a policy violation, or the wrong action taken for a customer.

We ranked the 591 tasks by difficulty, measured by how many steps and systems each one spans, then split them into five equal bands from simplest (Q1) to hardest (Q5):

MONARCH × AUTOMATIONBENCH · BY TASK DIFFICULTY The advantage is largest on substantial work SUBSTANTIAL WORK 1.1× 1.2× 1.3× 1.4× 1.5× 1.1× 1.1× EASIEST 1–4 checks 1.4× 1.1× Q2 4–6 checks 1.3× 1.2× MIDDLE 6–8 checks 1.2× 1.1× Q4 8–16 checks 1.1× 1.0× HARDEST 16–47 checks Monarch w/ Opus 5 Monarch w/ Kimi K3 591 tasks, one attempt. Each bar = Monarch completions ÷ the same model’s completions running on its own, per difficulty band.

Stop paying for tokens you burn on unfinished work

A workflow that stops one action short still costs the full run and delivers nothing. What matters, then, is the cost of a finished workflow. That's where Monarch pays for itself: it finished 23% more work than Opus 5 alone at about the same cost per completed workflow, so more of what you spend turns into work that's actually done.

You also control how much effort the model spends on a job. At low effort, each finished workflow costs about 31% less. At high effort, Monarch finishes 38% more of them, for about $0.20 more each. In other words, you run the routine work cheaply and spend on effort only where the work needs it.

MONARCH × AUTOMATIONBENCH · OPUS 5 The tradeoff you control COMPLETION RATE 0% 20% 40% 60% 39.4% 54.5% +38% more done LOW MAX COST PER COMPLETED WORKFLOW $0.0 $0.2 $0.4 $0.6 $0.43 $0.62 31% cheaper LOW MAX
Both bars are Opus 5 running on Monarch, compared at low versus maximum reasoning effort.

Example enterprise workflows Monarch completed that standalone models could not

Below are two more examples of real business workflows that failed on the model and succeeded with Monarch. None of these are unusual jobs either.

Benchmark example

Customer credit application workflow Outstanding Xero credit notes are matched to each customer's oldest unpaid invoice under company policy. A customer request in Gmail to use a newer invoice is blocked because it conflicts with policy. Xero applies each credit to the oldest invoice, updates the balance, and raises credit limits for larger notes. Gmail emails each customer their new balance. Three models without Monarch allowed a customer request to override policy; with Monarch, the policy remained controlling. FINANCE WORKFLOW Apply customer credits to open invoices 1 · READ INPUTS xero Open receivables Credit note Outstanding Invoices Oldest → newest Customer request “Use the newer invoice” Conflicts with company policy 2 · ENFORCE POLICY Which invoice? OLDEST UNPAID INVOICE ONLY NEWER INVOICE Blocked · policy controls VALID 3 · APPLY + UPDATE xero Apply in Xero Credit → oldest invoice Update customer balance Larger notes Raise credit limit POLICY-APPROVED CHANGES 4 · NOTIFY Email each customer New outstanding balance Credit application details GMAIL WITHOUT MONARCH · 3/3 MODELS Customer request overrode policy → credit moved to newer invoice WITH MONARCH Policy remained controlling → credit stayed on oldest invoice

1Finance: Apply customer credits to open invoices

  • Request: Apply outstanding credit notes to open invoices under the company's credit policy, email each customer their new balance, and raise credit limits on the larger notes. The company's new policy explicitly states that credits cannot be applied to newer invoices.
  • Systems: Xero, Gmail.
  • Where the model went wrong: All three models started correctly, applying each credit to the customer's oldest unpaid invoice. Then a customer emailed asking to put their credit against a newer invoice instead, and all three did what the customer asked. What they got wrong was treating a customer's request as authority over the company's own stated policy, and the cost is a policy breach committed in the company's accounting system.
  • What Monarch fixed: Monarch held the policy as controlling, applied each credit only to the oldest unpaid invoice, and did not let the customer's request override it.

Benchmark example

Gorgias ticket consolidation workflow Open Gorgias tickets use exact customer email to establish customer identity, then issue context to confirm whether tickets are true duplicates. Non-duplicates remain untouched. True duplicates are consolidated by keeping one ticket open, closing duplicates with a linking note, and logging each consolidation in Google Sheets. Without Monarch, two same-email tickets about different issues were closed; with Monarch, four true duplicates were closed and logged while unrelated tickets stayed open. SUPPORT WORKFLOW Gorgias ticket consolidation 1 · INPUT Open tickets Customer email Issue Ticket ID GORGIAS 2 · MATCH Same customer + issue? EMAIL + ISSUE NOT A DUPLICATE Leave untouched DUPLICATE 3 · CONSOLIDATE Keep one open PRIMARY Close duplicates Linking note GORGIAS 4 · LOG Consolidation log Primary ID Duplicate IDs Status GOOGLE SHEETS WITHOUT MONARCH Same email treated as enough → 2 unrelated tickets closed WITH MONARCH Email + issue reconciled → 4 true duplicates closed and logged

2Support: Consolidate duplicate support tickets

  • Request: Consolidate duplicate Gorgias tickets from the same customer (matching by email only), keep one open, close the duplicates with a linking note, and log each consolidation.
  • Systems Required: Gorgias, Google Sheets.
  • Where the model went wrong: All three models matched too broadly and closed two tickets that belonged to different customers. What they got wrong was never checking whether the tickets were really about the same issue, so two customers had a live issue closed on them, and the models logged closures they weren't permitted to make.
  • What Monarch fixed: Monarch matched by email and then reconciled the issue context, consolidated only the true duplicates, and left unrelated tickets untouched.

Governance and permission: access you can actually control

The more an agent can reach, the more it matters what stays off-limits. With a model on its own, staying in bounds is a matter of instruction: you tell it what not to do and hope it listens. Monarch doesn't rely on that.

Every system it maps becomes nodes in the product graph, and you grant access node by node. Switch a node off and the agent has no path to it, so it can't touch that action even if the model tries. The limit lives in the structure, not in an instruction the model can ignore.

High-risk actions are simulated first, and committed only if the result matches what was expected. Proven workflows run as deterministic code you can repeat and audit.

Where Monarch is strongest and weakest today

Monarch helped most where a workflow moves across several systems and has to change data in each one. It helped least on the simple, single-system tasks.

Those hard, multi-system workflows, where one wrong value breaks the whole outcome, are what our forward deployed engineers fix. They tune Monarch to the things a benchmark can't see: an unusual approval path, a quirk in how one of your systems behaves, a piece of product knowledge the model was missing. That's what turns a workflow that came up one step short into one that finishes.

HR had the most gains of any function:

On Opus 5, marketing and operations jumped too (+17 and +16). The smallest gains were in finance, about +3 everywhere, where the models are already strong.

Operations and support are the least consistent. Neither held up across all three models. Operations gained +16 running Opus 5 on Monarch but turned negative on GPT-5.6 Sol, and support turned negative on Opus 5. These are the hardest, most complex workflows, so the results are dependent on the model for now. They're also the areas we're investing in now to make the gains consistent.

PERCENTAGE-POINT GAIN MONARCH ADDS TO EACH MODEL, BY FUNCTION Where the lift comes from MONARCH W/ OPUS 5 MONARCH W/ GPT-5.6 SOL MONARCH W/ KIMI K3 +15 +10 +5 0 -5 +16.5 +13.4 +11.3 +17.0 +6.0 +1.0 +11.6 +5.3 +3.2 +16.0 -7.0 +2.0 -2.0 +7.0 +5.0 +3.0 +4.0 +3.0 HR MARKETING SALES OPERATIONS SUPPORT FINANCE AutomationBench 1.0.6 · 591 scorable tasks · one attempt, no retries · each bar is the completion Monarch adds over the same model alone

What changes in a real deployment with Monarch's FDE team

These strong improvements in completion rate, including the 23% lift on Opus 5, are what Monarch delivers before any work in your environment. Monarch does the mapping itself, enforces your permissions, and validates and transforms data as it moves between systems. What our forward deployed engineers handle is the part no product can know in advance about a specific environment: the edge cases.

They find the edge cases, build tests for them, and validate the correct behavior so it's structured into Monarch instead of leaving it up to the model to work through on its own.

Our FDE team starts where the benchmark is weakest because that's the widest gap to close. Once they've built your environment's edge cases into Monarch, the workflows that came up one step short start finishing, no matter which model runs underneath.

What this means for enterprises

Every enterprise is somewhere on a multi-year AI transformation. The challenge is turning that investment into working systems before the cost outruns the value.

Most AI spend today pays for work that never gets finished. An agent runs up the full cost of a workflow, stops one step short, and delivers nothing. You paid for the whole job and got none of the value.

Monarch finishes the job and more (up to 23% more), on every model we tested. And it does that inside your rules: every action stays in the permissions you set, anything out of bounds is blocked or escalated, high-risk steps are simulated before they commit, and proven workflows run as code you can repeat and audit.

Governance is what earns that trust, and once you have it, you don't have to move everything at once. You can take a single workflow, get it working end to end, and expand from there to the next team and the next function.

Reinventing how a company operates takes years, and it gets done one workflow at a time. Monarch is what lets you get that first workflow working and trust it enough to keep going.

Book a demo to transform your business with Monarch

Bring us your most acute pain points and we'll pilot Monarch on your systems.

We got your email. We'll reach out shortly to set up time.

What's next

This is our first released benchmark, and we're already working on improving Monarch to push these numbers even higher. More benchmark releases are coming soon.

We're working on the below next:

Methodology & caveats

We took the public AutomationBench, 591 business workflows across finance, HR, marketing, operations, sales, and support, and ran three frontier models through it: Opus 5, GPT-5.6 Sol, and Kimi K3. Each ran with and without Monarch under identical conditions: same tools, same reasoning effort, one attempt, no retries. A workflow counts only when every required effect lands in the real systems and no guardrail trips. The grader checks the systems, not the model's word.

Adjustments to AutomationBench:

  • These are our own runs, not Zapier's leaderboard. The setup is different, so the numbers throughout this report are the ones we control: the same model with Monarch versus without it.
  • Each condition is one matched run, not a replicated study, so read the results as directional. Low-reasoning figures are an earlier corpus revision, not directly comparable to maximum.
  • We scored 591 tasks, not 600. While evaluating the benchmarks, nine were impossible to complete the way the benchmark is built. We took them out instead of including them.
  • Guardrail violations fell on Opus 5 and Sol, but rose on Kimi K3 (212 → 228).
  • The difficulty breakdown covers Opus 5 and Kimi K3, not Sol in figure “The advantage is largest on substantial work.” Sol's control run was only about half sampled in this batch, and the tasks that ran skew toward the hardest ones. Its overall lift holds, but the sample isn't balanced enough to split by difficulty, so we left it out of that chart.
  • The win is completion, not cost. Monarch buys more finished work, not cheaper work. Cost per completed runs flat to higher, and GPT-5.6 Sol is about four times the others.

See Monarch run your unique complex enterprise workflows

We got your email. We'll reach out shortly to set up time.