Home  /  Blog  /  AI · Infrastructure

LLM Observability: Why Your Logs Aren't Enough

Standard application logs miss critical data for debugging AI features. Learn what to capture, when to use specialized tools, and how to start.

Topic  AI · Infrastructure Author  Andrew Dainty Published  10 October 2026 Read time  5 min

Your LLM application just failed in production on October 10, 2026. Was it the prompt? The model? A token limit? Did the API call even complete? Standard application logs often provide just a timestamp and a “failed” message. That’s useless for debugging AI-powered features. We need to see the full context of every LLM interaction to understand what went wrong, iterate on prompts, and track costs. Without specific observability for your AI calls, you’re flying blind. You can’t improve what you can’t measure.

The Common Wrong Approach

Most teams start by treating LLM API calls like any other external service call. They log the request and response, maybe an error message, and a timestamp. Sometimes, they store the full prompt and response in a database field for later inspection. This seems reasonable initially. It’s simple to implement with existing logging frameworks. The problem is, this approach quickly breaks down. You don’t get token counts, which are crucial for cost management and debugging unexpected truncations. You miss the latency breakdown for external API calls versus internal processing. Intermediate steps in complex agents are invisible. You can’t easily compare prompt versions or track how changes impact performance over time. When an issue arises, you’re left sifting through unstructured log data, guessing at the root cause, and wasting engineering cycles.

The Better Approach

True observability for LLM applications requires capturing specific, structured data for every interaction. This includes the full prompt, the model’s response, the model ID, latency, token counts (input and output), and any metadata like user ID, session ID, or prompt version. This data needs to be easily queryable and visualizable. We see three main paths to achieve this, depending on your team’s size and existing infrastructure:

1. Specialized LLM Observability Platforms: Tools like Langfuse are purpose-built for AI observability. They offer SDKs that wrap your LLM calls, automatically capturing prompts, responses, token usage, and latency. These platforms excel at visualizing traces for complex agentic workflows, tracking prompt versions, and providing dashboards for cost and performance. They often integrate with evaluation frameworks, letting you connect prompt changes directly to outcome quality. The main benefit is the speed of implementation and the rich, AI-specific features out-of-the-box. We recommend these for teams that prioritize rapid iteration and deep insight into their LLM interactions.

2. OpenTelemetry for AI: If your team already uses OpenTelemetry (OTel) for distributed tracing, extending it to LLM calls offers a vendor-agnostic approach. You instrument your LLM calls as spans, adding custom attributes for the prompt, response, model, and token counts. This integrates AI interactions seamlessly into your existing observability stack, allowing you to trace a user request from the UI, through your backend, and into the LLM API. While more setup is involved than with a specialized platform, OTel provides maximum flexibility and avoids vendor lock-in for your observability data. It’s an excellent choice for organizations with mature observability practices.

3. Enhanced Custom Structured Logging: For smaller teams or those with specific data residency needs, you can build a robust system using your existing structured logging infrastructure. For every LLM call, log a rich, structured JSON object containing all the critical data points mentioned above. Send these logs to an aggregator like Elasticsearch, Splunk, or even just CloudWatch Logs. Then, use a visualization tool like Grafana to build dashboards for latency, token usage, cost, and error rates. While we recommend external tools for comprehensive LLM observability, our own Email Triage product, for example, maintains a public prompt transparency page, which requires careful internal tracking of prompt versions and their associated fingerprints. This level of detail is only possible with robust data capture.

Regardless of the path, the core principle is to make every LLM interaction a first-class citizen in your observability stack. This allows you to quickly identify issues, compare prompt versions, and understand the real-world performance and cost implications of your AI features. For more on managing the underlying infrastructure, read our guide on How to Manage LLM API Keys and Secrets in Production.

Where This Breaks

Implementing a comprehensive LLM observability stack isn’t without its challenges. The primary concern is cost. Storing full prompts and responses, especially for high-volume applications, can quickly become expensive. You might need to implement data retention policies or truncate large prompts/responses for logging purposes, potentially losing some debugging context. Privacy and security are also critical; sensitive user data in prompts or responses must be handled carefully, potentially requiring redaction before logging, especially if using external observability services. For instance, our CV Matcher truncates CV and job text to a fixed budget before sending to the model, a practice that reduces data logged and processed.

For very small projects with minimal LLM usage, the overhead of a dedicated observability platform might be overkill, adding unnecessary complexity. A simple, structured log of key metrics might suffice. The “better approach” also introduces another dependency, another system to maintain. Over-instrumentation can lead to alert fatigue or obscure critical issues amidst noise. It’s a balance between comprehensive data capture and operational simplicity. Sometimes, the most common issues are not LLM-related at all, as we discuss in What Breaks First When You Put an LLM in Production.

Practical Next Step

Start small, but start now. Instrument your next LLM API call to capture a structured JSON log that includes: timestamp, request_id (a unique ID for this specific call), model_id, prompt_hash (a SHA256 hash of the full prompt to avoid storing sensitive data but still track prompt changes), response_hash, input_tokens, output_tokens, latency_ms, and status (success/failure). Log this to your existing log aggregation system. This minimal set provides immediate visibility into performance and cost without a significant infrastructure investment. As your AI features grow, you can then consider integrating a specialized platform or expanding your OpenTelemetry implementation. If you need help designing an observability stack that scales with your AI ambitions, we’re here to help you start a project.