Appearance
Observability and Debugging
Traditional observability answers "is the service slow, is it down"; agent observability has to answer a far more uncomfortable question: the system threw no errors, but the output is wrong — why? An agent run can complete with zero exceptions and all HTTP 200s, yet the task failed — the plan went sideways, a tool argument was filled wrong, the retrieved context was a terrible match for the question. Only trace-level observability can catch these "quiet failures."
This page covers three things: what observability data for agent systems looks like (the three pillars and the trace data model), which tools to plug in (LangSmith / Langfuse / Helicone / OpenTelemetry GenAI), and once you have traces, how to debug and how to turn them into production monitoring metrics.
1. The Three Pillars: trace, log, metric Redefined for Agents
The classic three pillars (trace / log / metric) all still exist in agent systems, but their weight has shifted completely:
| Pillar | Traditional backend | Agent system |
|---|---|---|
| Trace | Which microservices one HTTP request passed through | Every "think-act" step of an agent during one task: LLM calls, tool calls, retrievals, sub-agents |
| Log | Structured business logs, error stack traces | Full prompts, raw model outputs, tool inputs and outputs — essentially the complete recording of the "scene of the decision" |
| Metric | QPS, latency, error rate | Task success rate, token cost, tool failure rate, human takeover rate — many metrics can't be aggregated from logs and need scorers |
Two key differences are worth memorizing:
- Trace is the primary pillar; logs and metrics are derivatives. In the microservices world, metrics are the monitoring mainstay and traces are for troubleshooting; in the agent world it's the reverse — almost every metric that matters (success rate, cost attribution, tool failure rate) has to be computed from traces. When selecting tooling, first ask "can this tool's trace model express my agent's structure"; everything else is secondary.
- Non-determinism means "same input, different trace every time." In traditional tracing, diffs mainly happen after code changes; an agent's trace doesn't even have a fixed number of branches (how many times the agent loop iterates is decided at runtime). This directly shapes the debugging methodology later — you can't reproduce a problem by replaying a request; you can only rely on having fully recorded the inputs and outputs of the scene.
A litmus test
If your "monitoring" is just a wrapper around LLM API calls counting duration and status codes, that isn't agent observability — that's HTTP monitoring. The minimal unit of agent observability is the complete decision trajectory of one task execution (run), including every step's prompt, model output, tool calls, and intermediate state. Below that granularity, production problems can only be guessed at.
2. The Trace Data Model: runs, spans, and events
Terminology varies slightly by vendor (LangSmith says run, OTel says span, Langfuse says observation), but the data model has converged; it's essentially three layers:
Trace (one end-to-end task execution, e.g. "handle this support ticket")
└── Span (an operation with start/end times, nestable, forming a tree)
├── agent span one agent run (may contain multiple LLM calls)
├── generation span one LLM call (prompt, completion, token usage)
├── tool span one tool execution (inputs, outputs, errors)
└── retrieval span one retrieval (query, matched documents)
└── Event (point-in-time records attached to spans, no duration)
├── input/output messages (prompt messages, model replies)
├── evaluation results (score, the judge's explanation)
└── exceptions (stack traces, error types)A few details with real engineering consequences:
- The tree formed by parent_id matters more than time order. Laid out flat by time, a multi-step agent's spans are an undifferentiated soup; rendered as a parent-child tree, you can finally see "which plan produced this tool call." In multi-agent scenarios you also need trace context propagated across processes (see Multi-Agent Architecture), or sub-agent trajectories become orphaned traces.
- group_id / session_id is an independent dimension. The OpenAI Agents SDK's traces have a
group_idfield for chaining multiple traces from the same conversation thread (say, one thread ID). Langfuse's Session and Helicone's Session are the same idea: traces manage "one task," sessions manage "a continuous interaction." - Evaluation results are first-class citizens. The OTel GenAI semantic conventions added a
gen_ai.evaluation.resultevent in v1.38.0 (October 2025), carryingevaluation.name, a numericscore.valueor categoricalscore.label, and anexplanation, with a recommendedgen_ai.response.idfor correlation — so "scoring a particular generation" can attach directly back to the span being scored instead of living in some other system.
A minimal trace example from a real run
Below is the trace tree of a support agent handling "refund my order from last week" (indentation shows parent-child; the number after the colon is duration):
trace: "customer-service" trace_id=trace_9f3c... total 11.4s
├─ agent_span: support-agent 11.2s
│ ├─ generation_span: gpt-5 (turn 1) 2.1s
│ │ ├─ event: input.messages [system, user: "refund my order from last week"]
│ │ └─ event: output.messages tool_call: lookup_order(email=?)
│ │ ↑ note this: the model called the tool without asking for the email
│ ├─ function_span: lookup_order(email="") 0.3s status=ERROR
│ │ └─ event: exception ValidationError: email required
│ ├─ generation_span: gpt-5 (turn 2) 1.8s
│ │ └─ output: "Could you share the email on your account?" ← the agent chooses to ask back
│ │ ... (the trace continues after the user replies)
│ ├─ function_span: lookup_order(email="a@b.com") 0.4s
│ ├─ function_span: refund(order_id=88213) 1.9s status=OK
│ └─ generation_span: gpt-5 (turn 4) 1.6s
│ └─ output: "Your refund has been processed..."
└─ event: gen_ai.evaluation.result name=task_success score=1A trace like this contains nearly everything you need to debug: which step failed, what the model saw at the time, what the tools returned, how many tokens the whole task cost. Missing any one link and localization cost rises exponentially — that's the concrete meaning of "trace is the primary pillar."
3. The Tooling Landscape: Platforms, Gateways, and the Open Standard
The 2026 tooling landscape splits into three classes: all-in-one platforms (LangSmith, Langfuse), gateways/proxies (Helicone), and the open standard (OpenTelemetry GenAI). They aren't mutually exclusive; the common combination is "SDK instrumentation + OTel export + platform display."
LangSmith: the default option of the LangGraph ecosystem
LangChain's official closed-source platform. Zero-config integration with LangChain/LangGraph (two environment variables get you full traces), and the most mature eval-and-dataset tooling. The costs: it's closed source, only the Enterprise plan can be self-hosted, and billing is "seat fees + trace count" — as of mid-2026 the public pricing structure is roughly a free Developer tier (1 seat, 5k base traces per month), Plus at about $39/seat/month (10k included), overage billed per thousand traces, with the long-retention tier (400 days) priced at about twice the short-retention tier (14 days) (prices per the official site — verify before citing).
Trace billing hides a subtle trap: how you instrument directly determines the bill. An agent iterating 20 rounds costs several times more if each round is reported as an independent run than if it's merged into one nested trace. Work out your trace granularity before volume arrives.
Langfuse: the first choice for open source and self-hosting
MIT-licensed for the core features, self-hosting first, framework-agnostic — that's the deepest divergence from LangSmith. The data model is three layers: trace / observation (spans, generations, events) / score, and beyond the SDKs it ingests OTLP (the OpenTelemetry protocol) directly. The hosted version bills by "units" (any ingested event — traces, observations, and scores all count); as of mid-2026 there's a free Hobby tier (50k units per month) and a Core tier starting at $29/month; self-hosting has no such limits.
The typical way to attach a score to a trace (the entry point for online scoring; Section 6 builds on this):
python
from langfuse import Langfuse
langfuse = Langfuse()
# attach a score to an already-uploaded trace: from rules, human labeling, or LLM-as-judge
langfuse.score(
trace_id="trace-abc123",
name="task_success", # metric name; shows up on dashboards
value=1, # continuous 0.0-1.0 values and categorical values are also supported
comment="the refund flow completed end to end",
)Helicone: the one-line-of-code proxy model
A completely different approach: no code instrumentation — point the base URL of your OpenAI/Anthropic client at Helicone's proxy (one line plus a Helicone-Auth header) and requests/responses, latency, and cost land in the dashboard automatically. Open source, self-hostable, with gateway features bolted on (caching, rate limiting, retries). The official docs draw the two integration modes clearly: proxy mode is fast to set up and unlocks gateway features, but traffic passes through their servers; async logging mode moves reporting off the critical path, so network failures don't affect the main chain.
The proxy mode's ceiling is equally clear: it can only see "the traffic in and out of the LLM," not your agent's internal structure — the planning steps, the relationships between tool calls, why that retrieval came back empty. For multi-step agents, gateway metrics (cost, latency, error rate) are enough for monitoring, but debugging still needs SDK-level traces. It fits better as cost monitoring and a logging backstop than as the primary debugging tool.
OpenTelemetry GenAI semantic conventions: a standard still under construction
This is the most important structural change of 2025-2026: LLM/agent telemetry is converging on OTel, but as of August 2026 the standard isn't stable yet. The concrete state (based on a hands-on survey from July 2026):
- The GenAI semantic conventions have moved out of the main repo into the standalone
open-telemetry/semantic-conventions-genairepository; main repo v1.42.0 (June 12, 2026) marked allgen_ai.*content deprecated and migrated it out, and from v1.43.0 the main repo no longer carries GenAI content. - All GenAI-specific spans, events, metrics, and attributes were still in Development status as of July 2026, with none marked Stable; the standalone repo hasn't even cut a version tag. The generic attributes it references (
error.type,server.address, etc.) are Stable; thegen_ai.*namespace is not. - Naming has churned repeatedly in recent years, so production inevitably runs several generations side by side:
gen_ai.system→gen_ai.provider.name(v1.37.0, August 2025);prompt_tokens/completion_tokens→input_tokens/output_tokens(the earlier v1.27.0); one event per message → structured attributes likegen_ai.input.messages(v1.37.0). The core span names have converged oninvoke_agent,chat, andexecute_tool.
Don't treat a Development-status schema as a database contract
Four practical recommendations: (1) pin versions across the framework, instrumentation packages, and the OTel SDK — one minor upgrade can rename attributes; (2) when querying, COALESCE the old and new generations of fields (and note never sum them — during the compatibility window frameworks double-write the same data); (3) maintain your own versioned internal data model and map to the current conventions only at export; (4) keep custom attributes out of the reserved gen_ai.* namespace. Some instrumentation supports switching to the latest experimental conventions with OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental, but behavior differs by framework — trust the spans actually exported.
The tracing that ships inside vendor SDKs
Mainstream agent frameworks all have built-in tracing, but their relationships to OTel differ:
- OpenAI Agents SDK (see OpenAI Agents SDK): tracing is on by default, automatically wrapping
Runner.run()in a trace and generatingagent_span/generation_span/function_span/guardrail_span/handoff_spaninternally (plustranscription_spanandspeech_spanfor voice). By default it batch-uploads viaBatchTraceProcessorto OpenAI's backend; you can append third-party backends withadd_trace_processor()(Langfuse, Arize Phoenix, Braintrust, Datadog, and twenty-plus other ecosystem integrations on the official list) or replace them wholesale withset_trace_processors(). Note it does not natively emit the OTel GenAI semantic conventions — OTel requires community instrumentation or a custom processor. Tracing is unavailable to organizations under ZDR (zero data retention). - Claude Agent SDK / Vercel AI SDK and others: generally take the "direct OTel support" route. The Vercel AI SDK's telemetry has been split into a standalone
@ai-sdk/otelpackage that emits newer-generation attributes likegen_ai.provider.name, with the oldai.*namespace kept as a compatibility path. - LangGraph: LangSmith is the first choice, but it can also go through OTel via instrumentation like OpenLLMetry.
A side-by-side comparison:
| Tool | Positioning | Open source / self-host | Billing basis (mid-2026, per official sites) | Fits |
|---|---|---|---|---|
| LangSmith | All-in-one platform | Closed source; Enterprise only for self-hosting | Seats + trace count, tiered by retention | All-in LangChain/LangGraph |
| Langfuse | All-in-one platform | MIT, self-hosting first | Hosted bills by units (event count); self-hosting free | Needing data control, framework-agnostic |
| Helicone | Gateway/proxy | Open source | Request volume | Quick cost/latency monitoring, or as a gateway |
| Arize Phoenix | Platform | Fully open source, OTel-native | Free self-hosting | Wanting OTel-native at zero self-host cost |
| OTel GenAI | Standard/protocol | — | — | The underlying convention for avoiding single-platform lock-in |
4. The Debugging Methodology: Working Backward from Failed Traces to Root Cause
Once you have the full trace, the core debugging move is walking the tree backward from the end: look at the final failure point first, then ask at each level, "what was this step's input, and why is the output what it is?" The vast majority of agent failures fall into four root-cause classes, each with a characteristic "fingerprint" in the trace:
1. Planning failure
Fingerprint: the tools all executed successfully, but the tool sequence itself was wrong — a missing critical step, reversed order, terminating the loop at the wrong moment, or spinning in a pattern of repeatedly calling the same tool (several consecutive function_spans with nearly identical arguments). The test is comparing against "what the task actually required." The fix is usually to revise instructions, add planning constraints, or introduce an explicit planning module — not to touch the tools.
2. Tool failure
Fingerprint: an exception hangs on the function_span, or the tool returned an error/empty result. Split further into two kinds: the model filled the arguments wrong (like calling lookup_order without the email in the example above — the external sign of a generation problem) and the tool/environment itself is broken (API timeouts, insufficient permissions, a changed response schema). Fix the prompt or tool description for the former; fix the tool for the latter. The trace must store inputs, outputs, and the raw error together, or the two can't be told apart.
3. Context failure
The most hidden class. The model's behavior is "reasonable" but based on wrong information: the retrieved documents are irrelevant, key information got truncated in a long conversation, the memory system injected a stale fact. The fingerprint is in the generation_span's input.messages — which is why you must persist the complete prompt sent to the model; looking only at final output will never surface this class. The fixes live on the context engineering and RAG side, not the model side.
4. Model failure
The catch-all: the input context is complete and correct, tools are fine, planning is sound, but the model just reasoned wrong, hallucinated, or ignored the format. The fingerprint is whatever remains after eliminating the first three. Fixes: switch models, adjust temperature, split the task, or add structured-output constraints and retries for that step.
In practice, run the investigation in reverse order: check tool failures first (easiest to localize), then context failures, then planning failures, and only last blame the model. Empirically, "the model isn't good enough" is the conclusion that becomes valid last — most of the time it's a context or tool problem being scapegoated onto the model. For more anti-patterns, see Common Pitfalls.
Trace replay and diffing
For verifying fixes, two techniques are most useful:
- Trace replay: pull the
input.messagesout of a failed trace verbatim and re-run them on the modified prompt/model/tools, comparing the new trajectory. Because models are non-deterministic, one replay proves nothing — accumulate failure cases into a dataset and replay in bulk — which slides into the territory of evaluation, covered in Section 6. - Trace diffing: compare the structural differences between the "failed trace set" and the "successful trace set" — tool-call count distributions, retrieval hit rates, loop rounds. Structural differences usually point at systemic problems faster than the details of any single trace. If failed traces average 3× the tool calls of successful ones, the odds are the planning is spinning, not that one tool has a bug.
5. Production Monitoring: Which Metrics an Agent System Should Watch
The traditional RED (Rate / Errors / Duration) covers only half of an agent's health. A passing production dashboard needs at least these five metric groups, and all of them should be sliceable by prompt version, model version, and user cohort — aggregate numbers have almost no diagnostic value in agent systems:
| Metric group | Concrete metrics | Alert signals | Data source |
|---|---|---|---|
| Task success rate | End-to-end completion rate (requires defining "complete"), abandonment/timeout rate | Version-over-version drop > 5% | traces + evaluation scoring |
| Cost | Tokens and dollar cost per task, attributed by step | Sudden jump in P95 per-task cost | usage attributes on generation spans |
| Latency | End-to-end P50/P95/P99, per-iteration duration, tool-time share | P99 over SLA; one tool's time share anomalous | span durations |
| Tool health | Per-tool failure rate, argument-validation failure rate, retry rate | Single-tool failure rate over threshold | function span statuses |
| Human takeover rate | HITL trigger rate, post-takeover edit rate, escalation-to-human rate | Sustained rise = agent capability regression | business events |
Three easily overlooked points:
- Look at the cost distribution, not the mean. Agent cost is heavy-tailed — most tasks cost a few thousand tokens while a few runaway loops burn dozens of times more. Watch P95/P99 and the "top N most expensive traces," and set a per-run token budget cap (see Cost Optimization). In the OTel GenAI conventions, token usage lives on
gen_ai.usage.input_tokens/output_tokens(the older generation usedprompt_tokens/completion_tokens); remember to merge both generations in aggregate queries. - "Task success rate" cannot be computed from logs. An HTTP 200 doesn't mean the task succeeded. It must come from evaluators: rule checks (like "was the refund actually created"), LLM-as-judge, or human spot checks. This is the biggest difference between agent monitoring and traditional APM, and the topic of the next section.
- Human takeover rate is the most honest capability metric. A user clicking "transfer to human," editing the agent's draft, or abandoning the session is all free failure labeling. These signals should be written back as scores on the trace, not left in the business database.
TIP
Treat every release (prompt changes, model switches, tool changes) as an experiment: stamp version labels into trace metadata and slice the dashboard by version. Agent system regressions are often not "got worse" but "the distribution changed" — one version makes P50 better and P99 worse, and you'll never see it without slicing.
6. Closing the Loop Between Online and Offline Evaluation
Observability and evaluation aren't two systems; they're one loop. The OTel GenAI conventions making gen_ai.evaluation.result a standard event is exactly this loop expressed at the specification level. The complete cycle:
Offline eval (before release) Online eval (while running)
─────────────────────── ─────────────────────
Failure-case dataset ─────replay────▶ batch-score the new prompt/model
▲ │
│ ▼
│ score sampled traces (judge/rules)
│ │
└──── low-scoring traces auto-join the dataset ◀────┘
│
▼
Debug (Section 4) → fix → back to offline evalA few operational lessons:
- Online scoring must sample; full coverage is impossible. LLM-as-judge has its own cost and latency. The common approach is stratified sampling: all failure signals (human takeover, tool exceptions) get scored, plus a random 1-5% of successes.
- The judge's explanation is worth more than the score. Scores go to dashboards; explanations go to the debugging workflow. Low-scoring traces arriving with the judge's rationale in an auto-generated to-do queue make the team's most efficient morning-meeting material.
- Offline datasets are raised on online data. A cold-start dataset is invented and won't cover the real distribution; only by continuously enrolling low-scoring traces and edge cases does the dataset evolve with traffic. This is the core loop emphasized in Evals in Practice.
- Put the regression gate before release. Every prompt/model change replays against the accumulated failure-case set first; below the pass threshold, it doesn't ship. Without this step, "fixed one case, broke ten" is the norm.
7. Privacy and Data Residency
Agent traces are among the most sensitive of all data: full prompts may contain user PII and business secrets, and tool outputs may contain database rows. Several decisions you must make:
- Use the sensitive-data switches. The OpenAI Agents SDK defaults to
trace_include_sensitive_data=True(recording complete inputs and outputs for generations and functions); turn it off viaRunConfig.trace_include_sensitive_dataor theOPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATAenvironment variable — at the cost of a large drop in debugging power. A finer-grained option is PII masking before upload (a custom processor or the platform's masking features), keeping structure while erasing sensitive fields. Note that the SDK's tracing is simply unavailable to organizations under ZDR (zero data retention). - Data residency decides the tooling. In finance, healthcare, and government scenarios, traces can't leave the domain — which rules out closed-source SaaS outright and points to self-hosting Langfuse / Phoenix, or building your own OTel collector + backend (Jaeger, Grafana, ClickHouse). This is why you think through compliance constraints before choosing a platform: migrating trace storage later is expensive.
- Retention period is a compliance question, not just a cost question. LangSmith prices by retention tier (14 days vs 400 days at twice the unit price), but the prior question is "how long are we allowed to store users' raw conversations?" Many jurisdictions impose explicit maximum-retention requirements on telemetry containing PII, and the long-retention tier may simply not be an option.
- Gateway mode means traffic passes through a third party. Managed proxies like Helicone handle all LLM traffic; for sensitive scenarios choose a self-hosted gateway or async logging mode, demoting "passing through" to "carbon-copied."
For the broader security surface (prompt-injection detection and auditing, tool permission boundaries), see Security & Protection — traces are also the post-hoc forensic record for security incidents, which requires traces themselves to be tamper-proof and access to be audited.
References
- The state of the OpenTelemetry GenAI semantic conventions (July 2026) — a hands-on July 2026 survey of the OTel GenAI migration, stability status, and what each framework actually emits; the primary fact source for this page's OTel section.
- open-telemetry/semantic-conventions-genai — the standalone repo the GenAI semantic conventions moved to in 2026 (no releases yet).
- Tracing - OpenAI Agents SDK — the official docs on the SDK's built-in tracing: span types, processor architecture, and the sensitive-data switch.
- Trace the OpenAI Agents SDK with Langfuse — the official integration guide for ingesting OpenAI Agents SDK traces into Langfuse via OTLP/Logfire.
- Proxy vs Async Integration - Helicone — the official trade-off explanation of Helicone's two integration modes.
- Langfuse vs LangSmith vs OpenObserve: LLM Observability Compared (2026) — a July 2026 comparison of the three vendors' billing bases (units / trace count / traffic GB).
- The Complete Guide to LLM Observability Platforms - Helicone — a vendor-perspective platform comparison; watch for its biases.