Are your AI agents measurable, visible, and controllable?

author photo
Author
,
September 24, 2026
•
15
min reading time

The adoption of AI agents has moved out of the isolated pilot phase and become part of companies' daily operations. This has raised a question most of them still cannot answer: How do you prove that an agent is actually delivering results based on what was invested?

‍Why doesn't a task completed by an AI agent necessarily mean success?

In the world of APIs and microservices, we know this scenario well. An HTTP status code 200 confirms that the call was processed. It does not necessarily mean the result was correct. An endpoint can return 200 with a semantically wrong payload, or even 200 with an error message in the payload, and no traditional monitoring catches this.

In the world of AI agents, the trap changes shape, but not nature. There is no equivalent HTTP status code. What exists is "task completed". The agent walks through the flow from start to finish, calls the right tools, and does the proper reasoning, but completion is not success. An agent can complete a task and still have recommended the wrong vendor or summarized a contract in a misleading way. The question that decides whether the investment was worth it does not live in "completed". It lives in the quality and assertiveness of the execution. That is exactly the layer most organizations still do not measure.

Some market data confirm the scale of the problem:

- Back in 2025, Gartner projected that over 40% of current agentic AI projects would be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls.

- In a recent survey released by Forbes, we see that half of the companies already running AI in production still cannot measure its return. The slice that has already put autonomous agents into production tends to feel this problem even more sharply, since there is no human in the loop reviewing every execution.

- In a Gartner survey of 227 sales leaders, conducted between August and September of 2025, 31% cited difficulty proving the ROI of AI tools as the main challenge for 2026.

‍Why is classic AIOps not enough for AI agents?

Gartner coined the term AIOps (Artificial Intelligence for IT Operations) in 2017 to describe the application of AI and machine learning to event correlation, anomaly detection, and response automation in IT environments. But when we talk about AI today, we are usually talking about generative AI, one among several other types of artificial intelligence.

An imprecise label becomes an undefined purchasing scope, which becomes an unmet expectation. This is exactly the kind of problem that shows up in the reasons why AI projects get canceled. Escalating cost and unclear business value rarely come from bad technology. They come from buying something without knowing exactly what is being bought.

This lack of consensus around the term is what opened the door to a new discipline, directly tied to agentic architecture: AgentOps.‍

What is AgentOps?

AgentOps is an operational discipline in a world that already has DevOps, MLOps, and LLMOps. While DevOps materializes as delivering and maintaining highly available systems, and MLOps keeps the model lifecycle under control, AgentOps keeps the behavior of AI agents measurable, visible, and controllable.

This goes beyond simply monitoring. It is a production operations layer for AI agents, responsible for directing how teams launch, observe, evaluate, debug, secure, govern, and evolve autonomous or semi-autonomous systems after they leave experimentation. In practice, this layer extends across four main fronts:

- Trajectory and execution: step-by-step tracing of the agent's reasoning and of the tools called throughout an agentic workflow.

- Memory and context: the retrieval and retention behavior of the information that feeds the AI agent's decisions.

- Runtime policies: guardrails, permission limits, and human approval for higher-impact actions.

- Evaluation and audit: the ability to replay, debug failures, and prove, after the fact, that an autonomous decision was made as expected.

Unlike classic AIOps, AgentOps is still a discipline under construction. There is no single market tool or practice. What exists is a heated ecosystem. That, however, is a risk. An ecosystem with no dominant standard means lock-in risk and rising integration costs with every new tool adopted.

This is where the most promising technical standardization effort of the moment comes in: the OpenTelemetry GenAI Semantic Conventions. They define a common vocabulary of spans for foundation model calls, tool calls, and agent execution, with native adoption in frameworks such as LangGraph, CrewAI, and AutoGen, and in APM tools such as Datadog, Dynatrace, and Grafana.

An important architectural detail: the instrumentation of tool calls is already addressed. The MCP Gateway layer, together with the LLM Gateway (responsible for governing calls to foundation models), makes up what we call the AI Gateway. It may seem like a detail, but it is extremely important, because most of an AI agent's actions today no longer rely on calling the model itself, but on calling tools to gather more context.

Related content: Why is the AI Gateway essential for companies that use AI agents?‍

Which AgentOps metrics really matter?

Each of the four AgentOps fronts has its own metrics. Each metric answers a specific business question. The ranges below are not a closed consensus. They are the maturity benchmark Sensedia uses as a reference, crossing market frameworks with what we have observed implementing agentic architectures in production.

Trajectory and execution

- Tool-call accuracy: measures whether the AI agent chose the right tool, with the right parameters, on the first attempt. The benchmark we recommend as the maturity floor is 95% on the first attempt. Below that, every extra call is rework (higher cost).

- p99 latency: the time the user, or the next step in a workflow, takes to get a response. Our recommended reference is up to 30 seconds for conversational interactions and up to 5 minutes for batch processing or more complex tasks.

Why does this matter? A turn (a full cycle of the agent reasoning, calling a tool, and observing the result) with 99% accuracy looks practically perfect in isolation. Repeated 100 times in a row, however, the probability that all steps succeed drops to roughly 37%. That is pure math, not an isolated failure, and that is why per-turn accuracy is as misleading as "task completed", as we discussed at the start of this content. In the integration world, this scenario is well known: the latency that matters is the real one (perceived from the start of the chain to the end), not that of a single service.

‍Memory and context

- Hallucination rate / groundedness: measures whether the AI agent's answer is grounded in the retrieved context, or whether the model filled the gap with something plausible but made up. This is the metric that determines how widely the agent can be used, whether by customers or for real, critical decisions, without constant supervision. The recommendation is 4.0 or higher out of 5 on the fact-grounding axis, following the same scale logic used by RAG evaluation frameworks and LLM-as-a-judge.

Runtime policies

- Instruction-following: measures whether the agent respected the explicit constraints of the prompt and of its Harness's policies. The benchmark we recommend as the maturity floor is 95% or higher. As a market practice, this percentage should be at least 95%.

- Refusal rate: how often the agent refuses to act when it should not (the opposite of hallucination, but just as costly). Up to 5% on legitimate queries is healthy.

Evaluation and audit

- Recovery rate: the agent's ability to recover from an error without human intervention. The recommendation is above 70% on transient tool failures. Also, 100% recovery is not a good sign, since it usually means the agent is masking a persistent failure with fabricated success.

- Cost-per-success: total tokens spent divided by successful completions, not by attempts. Up to 2x the cost of the task is the ceiling we recommend as a reference. It is the metric that puts technical execution and financial outcome on the same yardstick.

None of these metrics alone tells the whole story, nor do they answer the ROI question, whether the investment in the agent paid off.

Who is responsible when an AI agent fails?

None of the four fronts above answers the following question: If an agent fails, who is responsible?

The answer, today, in most companies, is "we don't know". And the problem usually is not a lack of tools. The security playbook most companies have was built for a different kind of threat.

That playbook's blind spot is exactly where most of today's AI risk lives: inside the company. Improper information sharing, use outside the expected scope, misdirected behavior of the agent itself. The risk does not necessarily come from outside. It comes from within. An agent is not an employee who makes a mistake once and can be questioned afterwards. It is a system that repeats the same decision, right or wrong, at scale and in seconds, with no pause to reconsider.

A quarterly audit, a manager's approval, or even an annual training all depend on the interval between one human decision and the next to work as control points. An agent gives you no such interval. The policy is still right on paper, but the problem is that it is not present at the exact moment an AI agent's decision is made.

In 2023, Gartner created AI TRiSM (Trust, Risk and Security Management). It is a framework built on a simple premise: policy declares intent, control enforces behavior. When the system acts on its own, only the second one sustains real AI governance. The methodology exists to turn what is today a document into a mechanism: rules that do not sit on a shelf waiting for the next audit, but that are enforced at the exact moment the agent decides to act.

In practice, this translates into three requirements that a written policy alone cannot fulfill:

- Visibility: knowing which agents exist, where they run, and what each one can access. An agent nobody has mapped (Shadow AI) is, by definition, ungovernable.

- Runtime enforcement: validating and blocking out-of-scope behavior at the moment of the action, not afterwards. It is the difference between discovering a problem and preventing it from happening.

- Traceability: recording every action with owner, scope, and authorization context, so that finding out "who authorized this", "when", and "why" is a query, not an investigation.

Notice that none of this is about locking the agent down. It is about ensuring it operates within the scope it received, and being able to prove it at any time. That is exactly what separates an agent that can take on a critical decision from one stuck in eternal experimentation.

‍How does AgentOps help prove the ROI of AI agent strategies?

As we saw at the start, many agentic AI projects get canceled for unclear business value or difficulty measuring. Leaders point to difficulty proving ROI as the main challenge. None of them say the technology does not work. They all say the same thing in different ways: nobody can prove it worked. And proving it has two sides that do not replace each other.

The first is execution. Is the agent doing what it is supposed to do? At what quality? At what cost? That is the territory of AgentOps and of the metrics we covered. Without it, claiming the agent is working is a perception, not a number that survives a budget conversation.

The second is legitimacy. Did the agent do that within what it was authorized to do? Is there a trail to prove it? That is the territory of AI TRiSM. Without it, no risk, compliance, or audit team signs off when it is time to give the agent real autonomy, and an agent without real autonomy is just eternal experimentation.

The point that ties the two together: an agent that can be neither measured nor held accountable never leaves experimentation. And experimentation is exactly the stage where cost is already full and return is still partial, because every execution still depends on someone watching. An agent's ROI does not appear when it starts working. It appears when it no longer needs supervision on every execution. That transition depends on evidence, not trust.

This is where architecture stops being a technical detail and becomes a business condition. Quality metrics per tool call, policy enforced at runtime, and per-agent identity share one requirement: they only exist if there is a central point through which every AI call flows. If each squad instruments its own agent its own way, the company may manage to answer isolated questions, but it cannot answer at the organizational level.

That is the role the AI Gateway plays in practice: it is the enabling middle layer for your strategy, centralizing and boosting AI security and governance both for traffic to foundation models and for tool and data source traffic to the agents. It is not just where traffic passes through. It is where policy leaves the paper and becomes real-time control. That is exactly where the data that proves the agent's value is born.

Related content: How to control AI costs through AI FinOps?‍

Conclusion

We will hardly be able to do everything at once. Trying to do everything at the same time usually does not work. Start with one agent, the highest-risk one or the highest-value one. Instrumenting one agent well teaches you more about what your operation really needs to measure than instrumenting ten halfway.

Before scaling that first agent, establish the baseline. Quantitative data is essential. With a baseline, talk turns into argument. It is the difference between "the team felt it got faster" and "cycle time dropped X%". Already with that first agent, give it its own identity and authorization scope. Per-agent credentials and authorization scope are cheap to do on the first agent and expensive to fix on the tenth. It is the same practice we follow with distributed systems, but with AI agents, growth is exponential. With one agent measured, with a baseline and with identity, you are not replicating an agent. You are replicating a method.

Cost issues are extremely important to an AI strategy. Control needs to happen at runtime, not in later reports.

We opened with the question of how to prove an agent is delivering value. The uncomfortable answer is that you do not prove it afterwards. You decide, before putting the agent into production, whether you will be able to prove it. An agent instrumented from day one answers that question in minutes. An agent that ran for six months without measurement, without a baseline, and without its own identity never answers it.

This is not a model problem. It is an architecture problem.

Want to know how to guarantee ROI in your AI agent-based strategies? Talk to our experts now!

‍

Begin your API journey with Sensedia

Hop on our kombi bus and let us guide you on an exciting journey to unleash the full power of APIs and modern integrations.

Blog

Related content

Check out the content produced by our team.

No items found.

Embrace an architecture that is agile, scalable, and integrated

Accelerate the delivery of your digital initiatives through less complex and more efficient APIs, microservices, and Integrations that drive your business forward.