MLOps Is Dead. Long Live The New MLOps.
The past few weeks have seen the business and AI communities shaken up by companies announcing that their models have hacked other corporations. The first volley occurred between OpenAI and Hugging Face , followed by recent announcements from Meta that its Muse Spark 1.1 model breached another company’s systems during cybersecurity testing by exploiting a misconfiguration that provided the model with internet access. As discussed in my prior article, What Hugging Face Had That You Don’t , these developments are challenging organizations to develop new in-house core capabilities for AI management, which we have, for the past decade, called MLOps (Machine Learning Operations). When writing one of the first versions of the MLOps Wikipedia page, I articulated the elements of ML operations as we needed back then, comprising health, orchestration, governance, and others. Now, I argue, it is time to fundamentally rethink what MLOps is.
MLOps, in the past decade, has evolved substantially. Good definitions of where it started can be found here , for example. It covered many areas, some of which were:
- Monitoring model behavior and detecting anomalies. Are your models behaving well? Do you have practices in place to react if they do not, and diagnose what is going on?
- Orchestration of models. Can you introduce new models, retire old ones, and decide what to use when?
- Governance. Are you tracking how models were developed and approved? Are there legal obligations in your domain? Are you following them?
These are only a subset of what MLOps is today, and it is also worth noting that the arrival of Generative AI and Large Language Models spawned new requirements (sometimes called LLMOps ).
Even with all of these developments, MLOps is about to undergo another fundamental set of changes.
While MLOps today are sophisticated, it usually assumes that AI behaves in a predictable way that can be assessed, monitored, and responded to. The recent announcements show that new operational challenges are coming from the fact that AI models are now capable of new and adaptive behaviors, and can actively resist an organization’s efforts to counter them. This is a new domain for MLOps. It implies that your MLOps teams are now not just interacting with AIs that are executing patterns, but AIs that are able to adapt to countermeasures. The first casualty of this shift isn’t your response plan, it's your visibility. Before you can ask whether your rollback strategy still works, you have to ask whether you'd even know an incident happened. These changes challenge not just detection but response. MLOps rollback logic, as commonly used, assumes falling back to a known-good model is safe because "known-good" doesn't decay. However, against an adaptive adversary, "known-good" only means proven-safe against attacks that already existed.
What Is New And Already Here?
In the past, MLOps concepts like orchestration included the idea of what model to use where, but these AI models were far more limited in the scope of what they could do. The choice of model was often an engineering decision driven by factors like response time, compute costs, etc. While those are still valid, now organizations are not just deploying models; they are deploying AI agents capable of planning, decision-making, and increasingly independent execution. This means that orchestration decisions are also now a function of Corporate Taste, the instinct, judgement, domain expertise, and institutional knowledge that your organization possesses. How Corporate Taste gets translated into day to day operational decisions is a new element to MLOps entirely. That is the real work of orchestration now: not routing traffic, but encoding judgment.
I recently heard a corporate CTO say that every product team should now have only two people, a stellar builder and a stellar customer advocate. When AI agents build, test, and deploy the majority of the product, these two individuals provide complementary Taste that turns the army of AI Agent execution engines into product ROI. The exact number (2 people) is not, in my view, the key insight. It is that team structure can now be driven by what AI does, rather than the other way around. MLOps in the past was a layer added to an organization. You may have had an MLOps engineer and an MLOps team, etc. Now, your human t
teams may become structured to work with the new AI workflows, creating implications for everything from Human Resources to Hiring, Training, Promotions, etc.
As a business leader, there are several steps you can take to adapt your organization to excel at the new MLOps.
- Strategy: Identify which organizations or teams will develop the in-house capabilities for rapid response MLOps, now dealing with adaptive AIs (both within and outside of your organization) that can affect your business. Even if you are working with Forward Deployed Engineers from your vendor organizations, understand where their responsibility ends and your team’s work begins. The Hugging Face article outlines what capabilities these organizations need to possess.
- Starting Action: Identify one person in your organization who is ultimately responsible for detecting issues when they occur. All strategies will flow from there.
- Strategy: Establish an understanding of Corporate Taste as a valuable resource in your business. Ensure that employees learn how to recognise it, document and protect it, and pass it along to others.
- Starting Action: Identify and document what your model orchestration layer is optimizing for, and what its KPIs are. Sit down with your leadership team and decide if these goals are what will drive ROI for your organization.
- Strategy: Understand the appropriate AI workflows for your business, and continuously evaluate the match between your organizational structures and the effectiveness of your combined human and AI resources for business ROI.
- Starting Action: Pick one organization and put their AI workflow and their organization chart side by side. Which one drives the other? Is the order in the direction you need? What would happen if the order was reversed?
Takeaways: The Non-Negotiable ROI
MLOps will change in both evolutionary and revolutionary ways. This will in turn affect what your employees do and who you will hire to do what. The goal has not changed. Map each strategy directly to ROI, as directly as possible, and encourage all your teams to do the same.
As AI grows rapidly, what it means to manage AIs (yours and other people’s) in a way that protects ROI is the new MLOps. MLOps will continue to change, but setting up organizational structures that will grow with it can start now.
Loading article...