Autonomous AI is moving from investor thesis to enterprise reality and the early results look less like demos and more like finished work.

Sanjot Malhi, who leads the global growth fund at Northzone, spent nearly two years building an investment thesis around it before the technology caught up.

“For the longest time it was a forward looking thesis, and the technology just didn’t exist,” he told me. “Then the tech showed up this year.”

Since Janaury his fund has backed Blitzy, an enterprise coding company that has raised $200 million at a $1.4B valuation, and XBOW, an offensive cybersecurity company that has raised $120 million.

Copilots fill in the blank while you type. Agents take an instruction, work for a few minutes, and hand back a result for a human to review. In fact KPMG says that nearly half of executives will pull back AI agents over cost.

Autonomous AI is different. You give the system a goal, it works for days or weeks, and it returns finished work with no humans in the loop. As one technologist put it to me, all autonomous AI systems are agents, but most agents are not fully autonomous.

Coding shows the progression clearly.

GitHub Copilot autocompleted code and then Cursor and Claude Code ran short supervised tasks.

Cognition’s Devin coded semi-autonomously and grew into a $26 billion compny that says Devin now writes 89 percent of its own code. The frontier now is systems that take on entire projects alone.

Autonomous AI Rebuilds Enterprise Codebases

Blitzy ingests hundreds of millions of lines of a company’s legacy code, absorbs its compliance policies, takes a goal and then builds it. Charles River Development, a State Street company, uses it to modernize decades old code that traditionally went to systems integrators on multiyear contracts, and interim CTO Jason Adams has spoken publicy about the results.

Builders FirstSouce, the largest supplier of structural building products in the United States and a Global 2000 company, tripled its software development velocity in its first three months on the platform. OpenAI cited Blitzy in its GPT-5.6 materials as a tester of the model’s coding capabilities.

Teams at Blitzy’s customers now kick off a project on Friday night and return Monday morning to complete work. Malhi calls that shift in human behavior the clearest sign that value is being delivered independently of people.

Autonomous AI Is Now The World’s Top Hacker

Cybersecurity tells the same story from the offensive side. Penetration testing, where ethical hackers attack a company’s own systems to find weaknesses, has always been periodic, partial, expensive, and dependent on scarce human talent.

XBOW, founded by GitHub Copilot creator Oege de Moor, built an automous hacker that tests everything, continuously, with no one in the loop, in a market Mordor Intelligence sizes at $2.72 billion this year and $5.54 billion by 2031.

In summer 2025 XBOW topped HackerOne’s US leaderboard, the first time the number one hacker was not a human being. Moderna’s deputy CISO Farzan Karimi has said publicly that XBOW caught a firewall bypass he had missed in his own review, the finding that convinced Moderna to sign on.

XBOW was amongst the only private companies globally given early access to Anthropic’s Mythos model during Project Glasswing. Anthropic themselves cited that XBOW’s early-access testing revealed that Mythos is a “significant step up over all existing models” and provides “absolutely unprecedented precision”.

And the category is bigger than one company. Horizon3, whose NodeZero platform also pentests autonomously, raised $250 million Series E this month at a valuation above $2 billion, triple where it stood a year ago.

The giants see the same future. Google's Big Sleep, built by DeepMind and Project Zero, autonomously found twenty new vulnerabilities in widely used open source software, and OpenAI folded Aardvark, its security researcher that caught 92 percent of known flaws in benchmark tests, directly into Codex. In chatting with the co-founders of Om Lab, Krish Chelikavada and Keon Kim, told me, "The next frontier is fully autonomous agents that run continuously toward a goal. Humans will focus on providing context and steering these systems. World models and preprocessing signals into easily accessible context layers will be helpful in making these agents truly reliable" Continuous autonomous testing also builds a harness that captures a company’s full security context, giving enterprises a middle layer they control rather than handing their data over to any single model provider.

Where Autonomous AI Goes Next

Blitzy’s stated ambition is to become the AI software factory for the enterprise, holding a company’s compliance rules and businesses so new products ship ready to go.

Malhi frames both investments as early versions of the enterprise brain , a system that holds an organization’s full context. Blitzy is the brain for the codebase, XBOW is the brain for security, and the models plug into that brain rather than owning it. Malhi sees these systems as the precursor to artificial general intelligence(AGI). AGI is AI that can navigate ambiguity, form hypotheses, hit dead ends, and iterate its way to a solution better than a person could.

Leaders should approach this with open eyes. Gartner predicts that by 2027, 40 percent of enterprises will demote or decommission autonomous AI Agents. They predict this will be done because of governance gaps discovered only after production incidents, so demand audit trails and clear accountability before handing over goals.

But waiting on the sidelines carries its own cost. Companies tripling their development speed in a single quarter are running product autonomous AI.