I’m talking to many companies these days that are trying to incorporate AI into their business practices. Most of them are still in what I call “the administrative phase,” when they just try modest automations that allow some workers to draft documents or presentations, answer emails, or do research on certain topics, but given my recent articles arguing for companies to go much farther, I’m starting to get more and more questions from enterprises looking to go way beyond that, and get AI to become more “strategic.”
For these companies, the main concern is expressed as a question: Will AI be smart enough? But I think a better and more useful question would be: How does the AI know whether what it did actually improved the business?
The examples are well known and not theoretical: from optimizing a salesperson for revenue, margin, or retention, to a factory in which we can improve throughput, but we could also improve quality, or reduce defects. In business, “optimization” rarely means just one thing: improving one metric can easily damage another. So, in the same way management already works through objectives and feedback, AI agents need to do exactly that. In fact, OpenAI now explicitly frames enterprise evaluation as “specify → measure → improve, and in doing so, it can turn what would be vague business goals into actionable and measurable expectations.
Outputs are easy to measure; outcomes are harder
Current AI metrics often measure variables such as answer quality, task completion, latency, cost, etc. However, what businesses actually care about are things like customer churn, margin, conversion, claims accuracy, delivery times, or customer lifetime value.
An agent could be completing its task perfectly as instructed, and still hurt the company by doing so. Businesses do not ultimately care about generated outputs: they care about outcomes. It is worth repeating, because it changes the entire conversation: businesses do not ultimately care about outputs; they care about outcomes. What we demand in corporate implementations of AI is completely different from what we need in our individual use.
A simple business example would be the well-known “successful” discount: If you give an agent the objective of getting a customer to renew, the agent could perhaps offer a 20% discount. When the customer renews, that would initially be seen as an apparent success, but it comes at a cost: the margin collapses, the customer learns to wait for discounts, and similar customers may start to demand the same. What counts as success depends on the scoreboard. As Goodhart’s law puts it, when a measure becomes a target, it ceases to be a good measure. Poorly chosen metrics become dangerous when they are optimized aggressively.
The industry is starting to realize that evaluation has to become continuous
OpenAI’s guidance for workflows puts a lot of emphasis on capturing full workflows, grading behavior, detecting regressions, improving prompts, routing, and guardrails. The important managerial translation of that is that AI cannot simply be approved once and then forgotten: evaluation has to necessarily become an active part of the operations.
The real world is our final exam: benchmarks happen before the systems are deployed, but companies operate in changing environments. According to NIST, controlled pre-deployment testing cannot capture every unexpected behavior or consequence that may emerge in real-world use. Like in “The Sorcerer’s Apprentice”, but in real life: that magic wand you use to optimize something can easily turn back against you.
The useful question is whether the system worked in the business, not merely in the test. Basically, we need to move from “Did the agent succeed?” to “Did the company improve?” According to recent work by McKinsey, many organizations are scaling agents faster than they can redesign the work beneath them, and boards and CFOs are starting to demand clearer value. And when we say value, take into account that token costs are only one of the pieces: agentic workflows must be judged economically at workflow level.
The managerial point is simple: task completion is not the same as business success.
AlphaZero as the intuitive precedent
Think about AlphaZero, the brilliant algorithm designed by DeepMind: AlphaZero did not become strong because it could describe chess: it improved because it acted, it saw the consequences of its actions, and it had an unambiguous objective: to win.
DeepMind says it learned through repeated trial and error, to favor moves that could increase its chances of winning. Companies, obviously, have goals that are much messier than chess, but the principle survives: learning requires feedback tied to an objective.
The hard part is choosing the scoreboard: “increase sales” is clearly not enough. What about margin, or churn, or returns, or compliance, or brand image? When establishing goals, multiple objectives can (and do) conflict. And those conflicts are not engineering problems alone: they are management problems. Therefore, boards and executives eventually have to decide what should improve, what must never be sacrificed under any circumstance, or what trade-offs could be allowed. Once systems can continually optimize, the objective itself becomes a governance decision.
As we saw in my previous article, something has to steer. But not only that: steering only makes sense if there is a clear destination and a way to measure progress toward it. So we are moving from intelligence to steering, to setting objectives, and finally, to learning. The missing control function must learn from repeated enterprise episodes toward an explicit objective.
The CEO questions
What I think CEOs should ask are questions like “What exactly is our AI optimizing?” or “How will we know whether its actions improved the business?” or “Does what it learns from those outcomes change what it does next time?” If executives cannot answer those questions in a proper and competent way, then I’m afraid that “autonomous” may simply mean “unsupervised activity”. Good luck.
“Smart enough” is not the same as “pointed in the right direction.” The next generation of enterprise AI will need more than intelligence and more than autonomy. It will need a scoreboard — and, most importantly, the ability to learn from it. Keep it in mind.