AI & Machine Learning

LLM Observability: Architecture for Systems You Can Debug


Share

LLM observability architecture: a decision-level trace linking prompt, retrieval, model, tool call and output spans.

Instrumenting an LLM Application: Key Points

  • The debugging unit in an LLM application is a decision, not an HTTP request.
  • A useful span holds the prompt version, retrieved context, tool calls and output together.
  • OpenTelemetry's GenAI conventions keep message content capture opt-in, deliberately.
  • Sampling strategy is an architectural choice set by the latency budget, not a config toggle.
  • Evaluation and production telemetry belong in separate stores, or regressions stay invisible.
  • Reconstructing a past decision needs the retrieved evidence stored, not the model output alone.

What Is LLM Observability?

LLM observability is the instrumentation of a generative AI application so any past model decision can be reconstructed from stored telemetry. It links the prompt version, the retrieved context, the tool calls and the output inside a single trace. That trace, not the completion text, is what a debugger reads.

The Debugging Gap in Production LLM Systems

Application monitoring answers two questions: what is broken, and why. The four golden signals in Google's SRE Book, latency, traffic, errors and saturation, answer the first well and the second only when the cause sits in the system's own state. The enterprise AI architecture blueprint treats telemetry as a platform layer, and this is that layer.

An LLM call breaks the second question. The request succeeded, the status code was 200, latency sat inside budget, and the answer was still wrong. The cause lives in the prompt, the retrieved context or the model version, and a request-level span records none of them.

Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs and inadequate risk controls (Gartner, 2025). A team that cannot explain one production decision cannot argue either case.

Why Doesn't Request-Level Telemetry Explain an LLM Outcome?

Because the request is the wrong boundary. One user question fans out into retrieval, reranking, one or more inference calls, tool execution and post-processing. A span carrying model name, token counts and duration tells you the whole thing was slow, not which layer was wrong.

The OpenTelemetry GenAI semantic conventions already model this as distinct span types: inference, embeddings, retrieval, memory and tool execution (OpenTelemetry, Development status, 2026). A retrieval span is named {gen_ai.operation.name} {gen_ai.data_source.id}; a tool span is named execute_tool {gen_ai.tool.name}. The shape of the standard is the argument: a decision is a tree of operations, so the trace has to be a tree.

The attributes that carry the decision

  • gen_ai.operation.name and gen_ai.provider.name are required on an inference span
  • gen_ai.request.model, gen_ai.conversation.id and gen_ai.output.type carry the call context
  • gen_ai.prompt.name and gen_ai.prompt.version are conditionally required when a named prompt template is used

Prompt version as a first-class span attribute separates "we changed the prompt" from "the provider changed the model underneath us". Without it, both regressions look identical and the postmortem becomes an argument rather than a query.

Request-level monitoring versus decision-level tracing across prompt, retrieval, model, tools and output.

What Belongs Inside a Decision-Level Span?

Most designs miss in one of two directions: too little and the trace explains nothing, too much and a sensitive data set now lives in an observability platform.

The conventions are direct about it. Instrumentations SHOULD NOT capture model instructions, input messages or output messages by default, and SHOULD provide an opt-in (OpenTelemetry, 2026). The three usage patterns the spec names are three different architectures.

Capture strategyWhat the span carriesTelemetry volumeAccess modelWhere it fits
No content capture (spec default)Metadata, token counts, prompt version, timingsLowestSame as ops telemetryNothing needs reconstructing later
Content on span attributesgen_ai.input.messages and gen_ai.output.messages inlineHighest, often past envelope limitsTrace readers read user contentPre-production and evaluation
External store plus span referenceReference identifiers, payloads in object storageLow on the hot pathSeparate controls on the storeProduction with sensitive inputs

The third row is what production needs and what instrumentation defaults rarely give you. When the platform has to reconstruct decisions, the architectural pattern is an upload hook on the instrumentation path plus a content store with its own access boundary: the span holds pointers, the payloads sit under separate controls.

One line in the spec is worth copying into your design. The upload hook SHOULD run regardless of the sampling decision, because the sampled-out trace is exactly the one someone will ask you to explain.

LLM decision trace showing prompt version, retrieved context, tool calls and output as key observability data.

How Do You Keep LLM Observability Inside a Latency Budget?

Synchronous capture on the request path is the failure mode. Serialising a prompt plus retrieved context costs time proportional to context size, which grows as the application gets better.

Three levers, in the order they pay off:

  1. Asynchronous export and batching. The request path writes to a buffer, not to a backend.
  2. Payload upload off the request path. The hook handles content; the span carries the reference.
  3. Trace-level sampling with deterministic retention. Keep every error, every decision past a risk threshold, and a slice of ordinary traffic for a baseline.

When a screening path has a hard latency budget, the architectural response is head sampling at the trace root with explicit retention rules, not a lower log level. Head sampling decides once, cheaply, before the work happens; tail sampling decides after seeing the whole trace and pays for it in collector memory.

Teams designing this layer against a delivery date can see how LLM engineering services cover instrumentation through production rollout.

LLM requests return normally while telemetry is buffered and exported asynchronously to an evidence store.

Separating Evaluation Telemetry from Production Telemetry

NIST's AI Risk Management Framework 1.0 gives measurement its own core function, MEASURE, alongside Govern, Map and Manage (NIST, 2023). Measurement is a platform requirement with its own storage, not a dashboard opened after launch.

Eval runs and production traffic have incompatible shapes. An eval run is batch, repeatable and pinned to a dataset and prompt version. A production trace is streaming, sampled and pinned to a session. In one store, aggregate metrics move for reasons no query can separate.

One span schema, two destinations. Same attribute names, distinct resource attributes, separate stores, so the two stay diffable while a nightly sweep never moves the production baseline. That split belongs in the observability layer of the production AI architecture patterns, not the reporting tool above it.

What Changes When a Regulator Asks Why a Decision Was Made?

Debugging asks why the system produced this output now. Reconstruction asks why it produced that output months ago, from stored evidence, with the retrieval corpus already changed underneath.

When a platform has to answer for a past decision, the architectural response is to version everything that fed it and store the evidence at decision time: prompt template version, model identifier, retrieval index version, and the identifiers of the documents returned. A completion string proves what the model said, not what it was looking at.

In RegTech platform engineering the retrieval corpus is a watchlist that changes daily, so an index version on the span separates a reconstructable decision from a guess. The same discipline runs through designing auditable AI systems.

Debugging versus reconstructing an AI decision: prompt, model, index and retrieved-evidence versions compared.

Lessons from Instrumenting an LLM Alert-Triage Pipeline

We built the LLM alert-triage layer inside RapidAML's AI-powered AML screening platform: the model proposes a disposition on a screening alert, an analyst confirms it. The question afterwards was never "what did the model say", it was "why was this alert closed". That means the retrieved evidence, not the completion.

Three calls held up:

  • We logged the full prompt and retrieval context per decision, not input and output alone. The completion answers no reconstruction request by itself.
  • We chose trace-level sampling over full synchronous capture, because the screening path could not absorb serialising retrieved context inside a sub-200ms budget. Payload upload moved to an asynchronous hook.
  • We chose structured decision records in PostgreSQL with JSONB evidence columns over free-text traces in the log store, because the trail had to be queryable by alert identifier and list version. Instrumentation was the OpenTelemetry Python SDK; payloads went to object storage behind its own boundary.

Separating evaluation telemetry from production came last and should have come first. Until it did, every prompt change looked like a model regression.

Instrumenting Your LLM Stack with DigiWagon

DigiWagon builds the telemetry layer for production LLM systems, from span schema to the evidence store behind reconstruction.

  • Span schema aligned to OpenTelemetry's GenAI conventions
  • Prompt and retrieval capture behind a content-store boundary
  • Sampling and export design against a latency budget
  • Evaluation telemetry kept out of the production baseline

Designing the Trace Before the Dashboard

LLM observability comes down to one boundary: is the trace built around the request or the decision? Choose the request and you get healthy dashboards over a system returning wrong answers. Choose the decision and the prompt version, retrieved evidence, tool calls and output arrive together, which is the only shape that answers "why did it do that". Sampling strategy and the evaluation split follow from that call. Systems instrumented at the decision boundary stay debuggable by the team that built them and reconstructable by the team that inherits them.

Ask an AI about this article

Turn this article into your own next step

Pick a question, then the assistant you use. It opens in a new tab with this article as its source.

The question it opens withRead https://digiwagon.com/blogs/llm-observability-architecture and turn its key points into questions I should ask my own team, one per point. Stick to what the article says.

The question it opens withRead https://digiwagon.com/blogs/llm-observability-architecture and explain its argument in plain language for a CFO, with the one decision it asks a business to make. Stick to what the article says.

The question it opens withRead https://digiwagon.com/blogs/llm-observability-architecture and tell me what it means for a mid-size company, what to do first and what to avoid. Stick to what the article says and mark anything you are not sure about.

FAQ

Questions we get asked.

How much engineering work does LLM observability add to an existing application?
Wrapping the model client with an instrumentation library is the cheap part, often a single dependency. The real work is retrofitting prompt versioning, attaching retrieval identifiers to spans, and standing up a content store with its own access boundary. Applications already emitting OpenTelemetry traces absorb the change more cheaply than ones adopting distributed tracing for the first time.
When should you sample LLM traces instead of capturing every call?
Sample once serialising prompts and retrieved context competes with the request's latency budget, which arrives as context windows grow, not as traffic grows. Keep every error, every decision past a risk threshold, and a slice of ordinary traffic for baselines. Head sampling at the trace root keeps the decision cheap; tail sampling keeps more but buffers spans in collector memory.
What are the most common mistakes when instrumenting an LLM application?
Four recur. Logging only input and output, which cannot explain a retrieval failure. Putting prompt content on span attributes in production, which copies sensitive data into the telemetry store. Omitting prompt and index versions, which makes a prompt change indistinguishable from a model regression. Using free-text logs where the trail has to be queryable.
Can general-purpose APM tooling handle LLM traces?
Partly. Transport, storage and trace views are the same, and OpenTelemetry-based instrumentation exports into them unmodified. What general-purpose platforms usually lack is prompt-version-aware comparison, retrieval-evidence linkage and evaluation-run grouping. Teams either extend the APM schema with GenAI attributes or run a purpose-built LLM telemetry store beside it, holding one span schema across both so traces from either path stay comparable.
What is the difference between LLM observability and model monitoring?
Model monitoring watches aggregate behaviour: drift, accuracy against a labelled set, token and cost trends. LLM observability explains a single decision, which needs the prompt, the retrieved context and the tool calls behind it. Monitoring tells you the system changed. Observability tells you which decision changed and what fed it. Production systems need both, stored separately, sharing one schema.