DARPA just flew a standard F-16 under AI control. The pilot was there, the AI did the flying, and the kit that made it happen wasn’t bolted to a one-off prototype.
There’s a switch in the cockpit of an F-16 sitting at Eglin Air Force Base in Florida. Flip it one way, a human flies the jet. Flip it the other way, software does.
That’s the part I keep coming back to. Not the AI. Not the dogfighting. The switch.
In most systems I’ve worked around, the interesting failure was never the automation. It was the handoff, the moment responsibility moves from one thing to another and nobody’s certain who owns the outcome mid-transfer. Aviation has spent fifty years learning this expensively, and the lesson keeps needing to be relearned: automation doesn’t remove the human problem, it relocates it.
So when DARPA announced an AI agent had flown a real F-16, my first question wasn’t whether the AI could fly. We’ve known that since 2023. My question was what happens at the switch.
What happened
On July 16, 2026, DARPA and the U.S. Air Force said an F-16 at Eglin had flown with an AI agent in control. The program is VENOM, for Viper Experimentation and Next-generation Operations Model, and it sits under a larger DARPA effort called Air Combat Evolution, or ACE.
The sequence matters more than the headline. Flight operations started in June with ordinary piloted missions, the boring kind, confirming the new hardware behaved safely in the air. Only then did the program move to flights where the AI took control while a safety pilot sat in the cockpit and watched.
A human pilot handled takeoff. The AI flew the majority of what came after. The pilot stayed in the seat throughout, with the authority to take the jet back.
Stevens, a test pilot with the 40th Flight Test Squadron at Eglin, described it as a beginning rather than an arrival. “Getting the aircraft into the air is always a monumental milestone for a complex test program. As we cross this starting line, we are excited to watch VENOM redefine the boundaries of autonomous flight.”
Starting line. Hold onto that.
Why a bolt-on kit changes the math
DARPA has flown AI fighters before. In September 2023, under the same ACE program, an AI agent flew the X-62A VISTA against a human-piloted F-16 in simulated dogfights over Edwards Air Force Base. No weapons were employed, but the flying was real, closing to roughly 2,000 feet at speeds near 1,200 miles per hour. Safety pilots could disengage the AI instantly. They never had to.
Genuine milestone. Also a science project. The X-62A is a one-of-a-kind aircraft that exists to test things. You can’t order forty more.
VENOM took F-16s from the operational fleet and added the VENOM Autonomy Kit, which DARPA describes as “a novel interface with the aircraft’s flight controls and mission systems.” The kit works without changing the jet’s core software.
To be straight: that clause isn’t a detail I dug out. It’s DARPA’s own pitch, sitting in their release, and several outlets led with it. What’s underappreciated is the second-order consequence.
Modifying a fighter’s flight control software means re-entering airworthiness certification, a process measured in years and tens of millions of dollars because the failure mode is a lost aircraft. It’s expensive on purpose. An interface that talks to those controls without altering them sidesteps that boundary entirely.
The risk doesn’t vanish, it moves, out of a certification regime built over decades for exactly this class of problem and into a kit that regime never contemplated.
Then there’s speed. If autonomy ships as a kit rather than a rebuild, how fast it spreads stops being set by aircraft production cycles and starts being set by how fast you can build kits. Those clocks differ by an order of magnitude.
The real shift is what the AI can’t see
Here’s the finding that got least attention, and it’s the most important thing in the program.
Brig. Gen. James “Fangs” Valpiani, the DARPA program manager, has been explicit that the earlier X-62A work gave the AI agents an advantage they won’t have again. In those tests, he said, the agents had “perfect” information and engaged with simulations rather than live sensors.
VENOM removes the crutch. These are fleet aircraft, with what Valpiani called “exactly the same sensors, exactly the same weapons flyout models, exactly the same dynamics and communication systems as every other aircraft that’s in the operational fleet.” The agents inherit the fleet’s blind spots. They “have to make decisions with the same knowledge limitations as human pilots.”
His own assessment: “This really is a sea change. All the previous research that DARPA has done and the Air Force has done up to this point with AI-based combat autonomy has involved some kind of white card, or some kind of simulation, standing in for what the real human experience is.”
That reframes the achievement. Flying a jet cleanly is a control problem, and control problems are where machine learning is strongest. Deciding correctly on incomplete, ambiguous, possibly wrong sensor data is judgment, a different animal. VENOM is the first time these agents face the second one.
What a toggle really means
The flip-of-a-switch framing is DARPA’s own language. It sounds like a convenience feature. It isn’t.
The government’s vocabulary here is mixed. DARPA’s materials use both “human-on-the-loop experimentation” and “human-in-the-loop test environment” for the same activity. That’s test-safety language, not settled doctrine, and I’d caution against reading a philosophy into a word choice that hasn’t stabilized.
The underlying tension is real regardless. In beyond-visual-range engagements, decisions compress into windows where human reaction time is a binding constraint. The whole argument for combat autonomy is that the machine operates inside a decision cycle a human can’t match. If that’s true, the human supervising it can’t fully evaluate its choices as they happen.
The switch isn’t the only safeguard, and I won’t strawman the engineering. The X-62A campaign ran automated constraints at machine speed: a 10,000-foot floor, 1,000 feet of minimum separation, knock-it-off criteria that fire without waiting for a human. The VENOM kit likewise respects the jet’s flight envelope.
But guardrails bound the failure. They don’t evaluate the decision. An agent can stay inside every envelope limit and still be wrong about what it’s looking at, and no altitude floor catches that.
Which leaves the switch carrying the judgment layer, and here’s the harder problem: a toggle assumes the fault will present as a fault. These systems fail confidently. Smooth, envelope-legal behavior built on a wrong read of the world looks exactly like correct behavior until the consequences arrive.
What happens to the pilot
The pilot doesn’t disappear. The job changes into something humans are measurably worse at.
Flying is an active skill, continuously exercised. Monitoring automation is vigilance work, and the human factors literature is unkind about it. Attention degrades with time-on-task, precisely when nothing’s going wrong, which is most of the time. Skills atrophy. Commercial aviation has fought this for decades and hasn’t won.
And there’s a sharper version of the handoff problem VENOM hasn’t reached. The transfers here are pilot-elected: a person decides to take the aircraft back. Aviation’s catastrophic handoffs run the other direction. Air France 447 killed 228 people after the autopilot quit on its own terms and dumped a degraded aircraft on a crew who didn’t have the picture. Nobody flipped that switch on purpose.
The question VENOM eventually has to answer isn’t what happens when a pilot reaches for the switch. It’s what happens when the kit disengages itself, at speed, and hands back an aircraft in a state the human didn’t watch develop.
Where this goes
VENOM’s F-16s are the foundation for DARPA’s next program, AIR, for Artificial Intelligence Reinforcements. Lt. Col. Patrick “Dice” Highland, the incoming AIR program manager, said the team now has “the opportunity to create dominant autonomy for beyond-visual-range, multi-ship combat.”
Multi-ship beyond visual range isn’t dogfighting scaled up. It’s a harder category. Within visual range, you can see what you’re maneuvering against. Beyond it, you’re deciding whether to trust a track on a display, built from sensor returns that can be degraded, spoofed, or simply wrong. The decision isn’t whether to engage an aircraft. It’s whether to believe a symbol.
Simulation work has run since 2024, building from one-versus-one to two-versus-two across both ranges, a single scenario repeatable a thousand times to study how small variations shift the AI’s decisions.
The eventual picture is a human supervising a formation of uncrewed aircraft: prove the autonomy on a crewed F-16 where somebody can take over, then move it onto aircraft with nobody aboard.
The part the excitement skips
VENOM aircraft are not going to war. They’re integration and verification aircraft, and they won’t fly without a human aboard. That constraint is official. Calling the program experimental rather than combat-deployable is a fair reading, not a formal designation.
So the claim that this scales across the fleet tomorrow isn’t something DARPA has said, and I won’t put it in their mouth. The FY24 budget request funded six aircraft from the start, at roughly $50 million. Six is the plan, not a waypoint on a curve. The first three reached Eglin in April 2024, the last in April 2025.
What’s true is narrower and still significant: the technical barrier to installing autonomy on a fleet-standard airframe just got lower, and the certification barrier got sidestepped rather than cleared. Whether that becomes doctrine involves budgets, airworthiness authorities, and decisions about lethal autonomy no engineering test resolves.
And verification remains unsolved. Run an agent a thousand times in simulation and you can study the distribution of its choices. Put it in an aircraft at 1,200 miles per hour and you get one sample. Weapons here exist as flyout models, not live ordnance, a real limit on what these tests establish. The gap between “performs well across simulated variation” and “is trustworthy in a situation nobody simulated” is precisely the gap machine learning has been worst at closing.
AIR’s stated ambition runs straight at it: uncertainty, deception, degradation. The conditions where these systems are least reliable are the conditions the next program is built to test. That’s the right call. It’s also the hard part, and it hasn’t been done yet.
The starting line
What strikes me about VENOM isn’t that the AI flew well. Of course it flew well. It’s the packaging. Somebody looked at a decade of autonomy research trapped inside one experimental airframe, asked how to get it out, and answered with a kit that doesn’t touch the software underneath and a switch that hands it all back. A modest-sounding answer to an enormous question.
Stevens called it a starting line, and he chose the right words. The interesting question was never whether the machine can fly the plane. It’s what we decide about the switch on the day the machine reads the situation better than we do, and reaching for it becomes the risk.
Chris Meredith writes about AI, technology, and what it actually means for real people. Follow along on Substack: monkeyattack.substack.com