AI

Your AI Can Now Hire Other AIs. Nobody Set a Budget.

Two robotic hands passing a glowing orb of light
Two robotic hands passing a glowing orb of light

Here is a number that should bother you: in March 2023, processing one million tokens through GPT-4 cost roughly $30 on the input side and $60 on the output side. By mid-2025, the cheapest frontier-adjacent models were running at $0.10 to $0.40 per million tokens. Today in late 2026, small models sit below $0.05 per million tokens, and even reasoning models, the expensive tier, run around $2 to $7.

That is a price collapse of more than 99% for small models in three and a half years, and a steep drop at every tier above them.

If the price of electricity had dropped 99% in three years, every energy economist on earth would be writing books about it. The AI inference pricing collapse has been hiding in plain sight inside developer dashboards and API billing pages, noticed mostly by the engineers watching their monthly bills. The rest of us are only now starting to feel the downstream effect: AI systems that can economically justify hiring other AI systems to do their work.

This is not science fiction. It is accounting.

The Math That Changes Everything

A stack of glowing servers in a dark data center aisle

Walk through the logic with me. Suppose you build an AI agent to handle your company’s customer research workflow. The agent reads incoming market reports, cross-references them against your product database, and flags items that require human attention. Straightforward agentic task, widely deployed in 2025 and 2026.

Now suppose that research workflow gets complicated. A new market report references three separate regulatory filings, two competitor earnings calls, and a technical paper with contested methodology. A single AI agent, working linearly, will take several minutes to process all of that in sequence.

Or: your orchestrator agent spins up six specialized sub-agents simultaneously. One reads the regulatory filings. One parses the earnings calls. One evaluates the technical paper. One cross-references the product database. Two synthesize the results from the others. Total cost at today’s rates for that entire parallel run: something between $0.03 and $0.12, depending on model selection. The whole thing completes in under thirty seconds.

At those numbers, the decision to use a multi-agent architecture is not a technical preference. It is an obvious economic choice. The agents are cheaper than thinking about it differently.

Who Gave the AI a Checkbook?

Here is where it gets complicated, and where I think most of the discourse is missing the actual story.

When an orchestrator AI decides to spawn six sub-agents, it is making a resource allocation decision. It is, in a meaningful sense, hiring labor. The orchestrator did not consult a human before making that call. It evaluated the task, estimated complexity, and made a staffing decision. Then it billed the company’s API account for the result.

I have been running multi-agent workflows in my own work for about eight months now. Most of the time, the orchestrator makes sensible choices. Occasionally it over-spawns; it creates more sub-agents than the task actually requires, spending three times what it needed to. There is no budget cap by default. There is no approval step before the spawning decision. The limit is whatever the company set on their API account, and most companies set limits that would feel generous to a human team.

The question “who authorized this expenditure?” has no clean answer when the authorizing party is the AI itself, acting within permissions it was granted implicitly rather than explicitly.

This is the accountability gap that the agentic AI explosion has quietly opened. Not “can the AI do the task” but “who is responsible when the AI’s AI makes a bad call?”

The Governance Problem Nobody Is Funding

An empty boardroom with an open ledger and robot silhouettes in the window

The AI safety discourse in 2026 has two loud lanes. One lane worries about superintelligent systems with misaligned goals at civilizational scale. The other lane focuses on near-term harms: bias in hiring algorithms, misinformation, surveillance creep. Both lanes deserve the attention they get.

But there is a quieter governance problem between those two poles that is landing on corporate IT departments right now, this quarter, and most organizations are not ready for it.

When a company deploys an agentic AI workflow, they typically define the goal (“handle tier-one customer support inquiries”), give the orchestrator a set of tool permissions (read the CRM, write to the ticketing system, query the knowledge base), and set an API budget. What they rarely do is set a sub-agent governance policy: a rule specifying under what conditions the orchestrator may spawn new agents, what those sub-agents are permitted to do, and what paper trail captures the decision.

I have talked to people running these systems at mid-size tech companies. The honest answer I get is usually some version of “the orchestrator just does what makes sense.” Which is another way of saying: the AI has discretion, and we have not written down the limits of that discretion, because we did not realize we needed to.

The fair objection is that the market is not standing still. Major agent platforms already ship approval gates and spending controls, and heavy governance can strangle the iteration that makes agents useful in the first place. Both points are true. But a control that exists and a control that is switched on are different things, and the economics do not push teams toward the second. Every approval step adds friction and cost, while every unchecked sub-agent costs pennies. Nothing in the bill tells you to turn the gate on. That is why this is a preventive problem, not a reactive one.

The EU AI Act, which moved into enforcement mode in phases through 2025 and 2026, has provisions for high-risk AI systems. But multi-agent orchestration does not fit cleanly into the risk categories as written. An orchestrator that hires sub-agents to process customer data is touching the personal data provisions. An orchestrator that makes resource allocation decisions might touch the automated decision-making provisions. But the “who authorized the sub-agent” question falls into a gap between categories that the Act’s drafters did not anticipate.

America has been slower, and the gap shows up in the details. The executive orders from 2023 and 2025 created voluntary commitments and reporting requirements for frontier developers. Worth doing. But if you read those orders looking for guidance on what happens when an orchestrator decides to spawn sixty sub-agents, you are going to come up short. The question of who governs that hiring decision has no home in current US policy.

What Corporate IT Is Realizing Too Late

In the past six months, I have watched a pattern play out at several organizations. An engineering team deploys a multi-agent system. It works well. The team expands its scope. The orchestrator gains new tool permissions. Nobody notices when the system starts generating sub-agents that make decisions the original engineering team would have flagged for human review.

The discovery moment is usually a billing anomaly or a customer complaint. Either the API costs spiked because the orchestrator over-spawned for a complex task, or a sub-agent made a call that a human would not have made and a customer noticed.

The engineering team’s post-mortem almost always surfaces the same finding: “we authorized the AI to use its judgment, and it did, and it was wrong in a way we did not anticipate.”

This is not a failure of AI capability. The models are performing as designed. It is a failure of governance design. The organizations built a system that could hire other systems and then did not specify the hiring criteria.

The fix is not technically hard. You can add approval steps before sub-agent spawning. You can set hard caps on concurrent agents per workflow run. You can require human review for any sub-agent decision that touches a regulated data category. You can log every spawning decision with enough context for a retrospective audit. These are engineering choices, not research problems.

The problem is cultural. Organizations are deploying multi-agent systems at a pace driven by competitive pressure and the intoxicating cheapness of inference. The governance layer is treated as something you add after you know the system is working, not something you design into it from the start.

The Precedent That Matters

There is an analogy I keep coming back to. When companies started deploying automated trading systems in the early 2000s, the default assumption was that a human trader would always be in the loop for significant decisions. That assumption eroded gradually, then all at once. By 2010, algorithmic systems were executing trades faster than any human could review. The Flash Crash of May 2010 happened because those systems, each operating within their own rules, collectively created a feedback loop that no individual human authorized.

Nobody decided to remove humans from the loop. The economics of speed made it irrational to keep them there.

The economics of AI inference are doing something similar, but across a wider surface area than financial markets. When it costs less to spawn a sub-agent than to write a prompt asking a human for input, the economic incentive to keep humans in the loop disappears. Not because anyone made a policy decision, but because the math stopped supporting the old workflow.

The organizations that navigate this well will be the ones that design governance into their agentic systems before the cost curve makes it feel unnecessary. The ones that do not will rediscover, at some inconvenient moment, that “the AI used its judgment” is not an accountability framework.

Your AI can hire other AIs. The billing account is already open. The only question is whether anyone wrote down the rules before the first invoice arrived.

Chris Meredith writes about AI, technology, and the futures we are building faster than we can think about. Follow at @ChristopherMeredith.

Chris Meredith writes about AI, technology, and what it actually means for real people. Follow along on Substack: monkeyattack.substack.com

Leave a Reply

Your email address will not be published. Required fields are marked *