The time machine has more than one stop

Vitor on our team shared Gergely Orosz’s article on OpenAI’s agentic software factory. The chart showing Codex usage by department is what caught my attention.

Engineering moves first. Finance, recruiting, and legal follow with a much steeper climb. By June 2026, Codex accounts for roughly 90% of the average user’s AI output tokens in those functions. It has become their primary AI tool. The measure describes AI usage, rather than the percentage of their jobs automated.

Orosz describes non-developers building solutions to their own problems. They also improve documents or spreadsheets. They share what they create too. Teams are also sharing reusable instructions that let agents generate tools such as dashboards. The people asking for something can increasingly have a hand in building it.

Even from San Francisco, we look at the labs to see things we’re trying to bring into our own businesses. I think that gap is perhaps six to twelve months. Then there’s the longer transition into an enterprise with years of accumulated systems, responsibilities, and controls.

I’d bet that over roughly three years, we’ll see more enterprise teams working this way. I wouldn’t put every function on the same timeline. Some will be able to move quickly; others will need much more work on their systems and processes before they can trust the result.

To make that concrete, imagine how the working day could change in finance, recruiting, and legal. These are examples of what an enterprise could build toward.

Finance can investigate a problem without waiting for a report

Imagine a finance analyst trying to understand why a business unit’s margin has fallen. They request an extract. Then they reconcile revenue and staffing records while waiting for someone to change a report before they can test an explanation.

With an agent connected to the relevant systems, the analyst could ask it to investigate the movement. The agent could assemble the records, flag discrepancies, and build an analysis the analyst can explore. A follow-up question could become another calculation or view within the same piece of work.

The analyst would still need to challenge the explanation. A change might reflect a timing issue, a deliberate commercial decision, or an actual operating problem. Their knowledge of the business helps distinguish those possibilities.

To trust the result, finance needs agreed definitions, authoritative records, and calculations that can be inspected. If two systems disagree, the agent should surface that disagreement. A polished chart is little use if the analyst has to reconstruct the work to find out which numbers it used.

That gives the analyst more room to investigate what’s happening and work with the business on a response. When a question changes, they can change the analysis too. If other teams start relying on what they’ve built, they’ll need technical help with reliable upkeep.

Recruiting can improve the process candidates experience

Consider a recruiter trying to understand why candidates are dropping out between interview stages. The evidence may be spread across the hiring system, calendars, and conversations with hiring managers.

An agent could assemble a timeline, identify repeated scheduling delays, and help the recruiter test an explanation. With approved access, it could also help build a workflow that proposes interview slots, tracks missing feedback, and flags candidates who have been left waiting.

The recruiter could try a change and see whether candidates spend less time waiting. They could put more attention into conversations with candidates and hiring managers. That would replace time spent chasing calendars and missing feedback.

The recruiting team would need to decide which candidate records the agent can access and when it can contact someone. In this example, the recruiting team would retain responsibility for candidate decisions as well as sensitive conversations.

The recruiting manager would need to give someone time to test and improve the workflow. Otherwise, the person most able to fix a recurring delay could remain too busy chasing it every day.

Imagine a legal team handling a recurring category of supplier agreements. An agent could compare each incoming document against the current approved playbook, identify departures, and prepare a review linked to the relevant clauses.

The lawyer could focus on the exceptions, ask the agent to investigate related provisions, and decide which changes to propose. As the team encounters recurring issues, it could improve the review instructions, then test them against previously reviewed agreements.

The lawyer could get to the issue that needs their judgment sooner. Is this a risk the company should accept? Does the commercial relationship justify an exception? Those decisions and the advice given would remain their responsibility.

For that to work, the lawyer needs to know the reviewed agreement version. They also need confirmation that the agent used the current playbook. They need to catch important clauses the agent missed, as well as things it got wrong. Permission to prepare a review would not automatically give the agent permission to send terms or commit the company.

The team could then use what it learns from each review to improve the next one, with someone responsible for testing those changes.

Trust has to be built into the work

The article shows the direction of change at OpenAI. The team examples above are what I think that could look like inside an enterprise.

Orosz describes agents deeply connected to OpenAI’s internal systems, with unlimited internal token budgets. An enterprise has to establish which parts of that approach it can support economically and reliably.

People need to be able to follow what happened. An analyst should be able to trace a calculation back to its records. A recruiter needs to know whether an invitation was actually sent, and a lawyer needs to see which agreement was reviewed. If people have to repeat the whole job to trust it, much of the benefit disappears.

I’d give an agent more responsibility as the team sees evidence that it can handle it. That might mean reviewing every proposed action at first, then letting it carry out a defined set of routine actions once testing supports that decision. Someone still needs to be able to stop it, investigate a mistake, and get the work back on track.

This is part of rebuilding how the enterprise works. The existing systems may remain, while agents take on more of the work of moving between them.

That is part of the work we’re doing at Monarch. People in finance are going to build and improve workflows themselves. The same applies to recruiting and legal. Supporting that work requires agents to understand and act reliably across the systems the business already runs, within the permissions those people actually have. The model is only one part of the change.

People need a supported way to take on the new work

The finance analyst who improves an investigation workflow is doing useful business work. Their manager needs to make room for it and agree how that contribution counts.

Some people will want to build. Others will contribute by explaining the difficult cases, checking results, or helping colleagues learn. Give them ways to do that work with support. Making a team more capable shouldn’t mean leaving someone in finance responsible for software they don’t know how to maintain.

Managers need to explain what happens to the job as the work changes. What can the person stop doing? What are they now expected to take on? And who will help when they find a problem they can’t fix? People deserve answers beyond a promise that their work will become more interesting.

Training should include real examples from the team’s work, reviewed with experienced colleagues. Technical teams still have a substantial role in maintaining shared components, testing consequential changes, and supporting recovery. The ongoing workflow owner needs the capacity and authority to connect that support to business results.

I’d expect a business further along this path to solve problems that used to sit in a queue for months. The person who spots a problem would have a way to investigate it, build an improvement, and get the right people involved before others rely on it.

That’s what I’d want the enterprise to look like in three years: people across the business able to act on what they know, with the authority, systems, and support to do it well.