An Agent Is Not Production-Ready Until Its Decisions Are Observable

A luminous agent path leaves observable traces through evaluation checkpoints. / Ein leuchtender Agentenpfad hinterlässt beobachtbare Spuren durch Evaluationspunkte.

A successful demo proves that an agent can complete one visible path. Production asks a harder question: can the organisation understand what happened when the path changes?

An agent interprets intent, selects tools, supplies parameters, uses results and decides whether to continue. The final answer may look correct even when the process underneath was wasteful, fragile or unsafe. Conversely, a disappointing answer may be caused by poor retrieval, a failed tool or an unclear instruction rather than the model itself.

If the team sees only the output, every incident becomes guesswork.

Observability is therefore not a technical feature added after deployment. It is the evidence layer that makes agent operation possible.

Logs are necessary, but they are not enough

Traditional logs tell you that a request happened, how long it took and whether a component returned an error. Agent systems need that information, plus the path through the work.

A useful trace should make several moments inspectable:

  • the user or system intent the agent received;
  • the plan or route it selected;
  • the tools it chose and the parameters it supplied;
  • the evidence returned by those tools;
  • the point where it acted, escalated or abstained;
  • the latency and cost accumulated along the path.

Tracing does not mean recording everything without restraint. Sensitive inputs, outputs and tool results need deliberate redaction, retention and access. Visibility without data discipline simply creates a new risk surface.

Evaluate outcome and process separately

An agent may arrive at a plausible answer through the wrong process.

Measure at least four evidence lanes:

  1. Outcome: Did the agent complete the intended task to a usable standard?
  2. Process: Did it select the right tools, use valid parameters and follow the expected boundary?
  3. Guardrail: Did it protect sensitive information, escalate the right cases and avoid prohibited actions?
  4. Economics: Did quality justify latency, token usage, evaluation cost and human review?

No single score represents production readiness. A high-quality answer with uncontrolled tool use is not ready. A perfectly compliant flow that users abandon is not ready either.

Turn production traces into learning

Pre-deployment test sets are essential. They are also incomplete because real users create combinations the design team did not anticipate.

Use production traces to identify:

  • recurring failures;
  • new intents;
  • unnecessary tool calls;
  • slow or expensive paths;
  • cases where humans repeatedly correct the same behaviour;
  • successful exceptions that should become a designed pattern.

Curate representative cases into a regression set. Run it when prompts, tools, models, permissions or knowledge sources change. The goal is not a perfect benchmark. It is a growing institutional memory of what the agent must continue to do well.

Build an operating response, not a dashboard

A dashboard that nobody uses is only decorated uncertainty.

Define what happens when a signal crosses a boundary:

  • Who receives the alert?
  • Can the agent or one capability be paused?
  • Who decides whether the incident is a data, tool, prompt, model or workflow problem?
  • Which evidence is preserved?
  • When must users or stakeholders be informed?
  • Which test prevents the same failure from returning?

Observability creates value only when it shortens the path from unexpected behaviour to a responsible decision.

The strongest objection: “Evaluation is imperfect too”

Correct. Automated evaluators can be inconsistent. LLM-based judges can miss context or reward superficial quality. Human review is expensive and variable.

That is not a reason to operate without evidence. It is a reason to triangulate.

Combine deterministic checks for permissions and parameters, domain-specific evaluation for quality, sampled human review for judgement and operational signals for real-world impact. Calibrate evaluators against examples your experts agree on. Treat evaluation as a managed measurement system, not an oracle.

The production-readiness test

Select one important agent action. Ask the team to reconstruct a real execution from intent to outcome, including every tool call, source, decision boundary, latency and cost signal.

Then ask: Which person can pause this behaviour today, and what evidence would they use to restart it?

If the path or authority is unclear, the agent may be deployed. It is not yet operational.

Production readiness is not the absence of failure. It is the organisation’s ability to see failure, understand it and respond before uncertainty becomes damage.

Deutsche Ausgabe: Ein Agent ist erst produktionsreif, wenn seine Entscheidungen beobachtbar sind

Sources and framing

Editorial note: Product features, preview status and pricing change. Verify current documentation before implementation.