AI Engineering · 6 MIN
What should you log for LLM observability?
Log the prompt version, model, inputs, retrieved context, tool calls, tokens, latency, cost, and user feedback per request. Here is the schema and the traps.
For LLM observability, log one trace per request that captures the prompt version, model and parameters, the inputs, any retrieved context, every tool call, the output, token counts, latency, cost, and a link to user feedback. With that record you can answer the three questions production always raises: what happened on this request, how often does this kind of failure occur, and did the last change make things better or worse. This post lists the fields, the sampling and privacy decisions, and the dashboards worth building.
- An LLM request is a pipeline, so log it as a trace with a span per step, not a single text line.
- Always record the prompt version and model version, because most regressions trace back to one of them.
- Token counts, latency, and cost per request are the minimum for managing spend and speed.
- Logs contain user data, so decide redaction, retention, and access before you turn them on.
- Every failed or thumbs-down request should be one click from becoming a test case in your eval set.
- OpenTelemetry has GenAI semantic conventions, so you can avoid inventing your own field names.
- Nactore wires tracing and eval feedback into each AI feature before launch, so quality problems are visible on day one.
Why is normal application logging not enough?
Traditional logs tell you a request succeeded or threw an error. An LLM request can return a 200 status and still be wrong, slow, expensive, or unsafe. You need the content of the interaction and the intermediate steps to understand why.
LLM systems are also multi-step. A single user question might trigger query rewriting, retrieval, a model call, a tool call, a second model call, and validation. Without a trace that ties those steps together, you cannot tell which one failed.
What fields should every trace contain?
Use this as a starting schema.
| Group | Fields | Why you need it |
|---|---|---|
| Identity | trace ID, user or tenant ID (hashed if needed), session ID, timestamp | Group requests and find a specific complaint |
| Versioning | prompt version, model name and version, parameters (temperature, max tokens), schema version, code release | Locate regressions to a change |
| Input | user input, system prompt reference, rendered prompt (or its hash) | Reproduce the request |
| Retrieval | query used, passages returned with IDs and scores, filters applied | Separate retrieval failures from generation failures |
| Tools | tool name, arguments, result, duration, errors | Debug agents |
| Output | raw response, parsed result, validation outcome, stop reason | See truncation, refusals, and schema failures |
| Usage | input tokens, output tokens, cached tokens, estimated cost | Control spend |
| Performance | time to first token, total latency, retries, fallbacks used | Find slow steps |
| Feedback | thumbs, edits, escalations, downstream outcome | Connect quality to real results |
Provider usage fields are the source for token and cache numbers. For instance, Anthropic returns cache read and creation token counts in the response usage object, and OpenAI reports cached tokens in its usage details, as described in their prompt caching documentation.
Is there a standard for naming these fields?
Yes. OpenTelemetry maintains semantic conventions for generative AI covering spans, metrics, events, and provider-specific attributes. They live in the OpenTelemetry GenAI semantic conventions repository, and the specification has been evolving, so pin the version you use. Adopting the standard names makes it easier to move between tracing backends and avoids a pile of one-off attribute names.
What should you do about privacy?
Prompts and outputs often contain personal or confidential data. Decide before launch.
- Classify the data. Know which inputs may contain personal data, secrets, or regulated content.
- Redact at the edge. Mask emails, phone numbers, and account numbers before they hit the log store when the task does not need them.
- Set retention. Keep full payloads for a short window and aggregates longer.
- Restrict access. Not every engineer needs to read customer conversations.
- Check provider terms. Confirm how the vendor handles your data and whether logging to a third-party tool is permitted under your contracts.
For US, UK, and EU customers, involve whoever owns privacy compliance early, because retention and data-subject rules apply to logs too.
How much should you sample?
For a low-volume internal tool, log everything. At higher volume, keep metadata (versions, tokens, latency, cost, validation outcome) for every request and sample full payloads. Always keep full payloads for failures: validation errors, guardrail trips, refusals, timeouts, negative feedback, and escalations.
Add a "save to eval set" button to your trace viewer. Failures that become test cases in one click are what make an eval set grow with the product. See why AI features need evals before production.
Which dashboards and alerts are worth building?
Start with a short list and add only when a question demands it.
- Quality. Eval pass rate on a scheduled run and online grader scores on a sample.
- Reliability. Schema-valid rate, refusal rate, fallback rate, tool error rate.
- Cost. Cost per request and per successful task, cache hit rate, top spenders by tenant.
- Speed. Median and 95th percentile latency per step.
- Safety. Guardrail trips by type. See LLM guardrails in production.
- Drift. Input length and topic distribution over time, so you notice when users change.
Alert on rate changes against a baseline (for example, a jump in fallback rate after a release) instead of on single events.
What mistakes do teams make?
- Logging only the final answer. You lose the retrieval and tool steps where failures hide.
- Not versioning prompts. If prompts live as strings in code with no version, you cannot link a score change to an edit.
- Storing everything forever. It is expensive and a liability.
- Collecting logs nobody reviews. Schedule a weekly review of the worst traces.
Frequently asked questions
Do we need a dedicated LLM observability product?
Not necessarily. Many teams start with OpenTelemetry traces into the backend they already use. A dedicated tool helps when you want prompt management, trace review, and eval workflows in one place.
Should we log the full prompt?
Log it or a hash plus the template version and variables, so you can reproduce it. Apply your redaction and retention rules either way.
How do we measure quality online without labels?
Run a calibrated LLM grader on a sample of production traces, track user feedback signals, and review the lowest-scoring traces by hand. See how to evaluate LLM output quality.
What is the single most useful field?
The prompt and model version. It turns "something got worse" into "it got worse after this change."
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.