---
name: Agent Observability
slug: agent-observability-2
category: AI Engineering
description: Agent Observability instruments AI agents with logs, traces, metrics, token usage, latency, and cost telemetry. Use it to debug reliability issues, set SLOs, and monitor agent or LLM workloads.
github: "https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/agent-observability"
language: Shell
stars: 709
forks: 89
install: "npx degit https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/agent-observability ~/.claude/skills/agent-observability"
installs_to: ~/.claude/skills/agent-observability
source_path: devops/ai/agent-observability/SKILL.md
collection_size: 25
category_size: 2451
collection_url: "https://dirskills.com/collections/BagelHole/DevOps-Security-Agent-Skills"
added: 2026-08-24T05:15:51.357Z
last_synced: 2026-08-24T05:15:51.357Z
canonical_url: "https://dirskills.com/skills/agent-observability-2"
---

# Agent Observability

Agent Observability instruments AI agents with logs, traces, metrics, token usage, latency, and cost telemetry. Use it to debug reliability issues, set SLOs, and monitor agent or LLM workloads.

**Install:**

```bash
npx degit https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/agent-observability ~/.claude/skills/agent-observability
```

## README

# Agent Observability

Monitor AI agent behavior with logs, traces, metrics, and cost telemetry. This skill covers the full observability stack for LLM-powered applications: from raw Prometheus counters to Grafana dashboards, OpenTelemetry tracing, structured logging, cost tracking, SLO definition, and PII redaction.

---

## When to Use

Apply this skill whenever you operate:

- **Autonomous AI agents** that make multi-step tool calls (e.g., coding agents, support agents, data-pipeline agents).
- **LLM-backed APIs** serving chat completions, summarisation, or classification behind a REST or gRPC gateway.
- **RAG pipelines** where a retriever fetches context from a vector store before prompting a model.
- **Multi-agent orchestrations** (crew-style or graph-based) where several agents collaborate on a single task.
- **Batch inference jobs** that process thousands of prompts against a model endpoint.

Key signals that you need this skill:

1. You cannot answer "what is p95 latency for agent responses this week?"
2. You have no per-request cost attribution.
3. Debugging a bad agent response requires grepping raw application logs.
4. You have no alerting on token-usage spikes or elevated error rates.

---

## Core Metrics

Define these metrics at the application layer. All examples use the Prometheus client library naming conventions.

### Latency

```python
from prometheus_client import Histogram

# Total end-to-end latency for a full agent turn (user prompt -> final response)
AGENT_LATENCY = Histogram(
    "agent_request_duration_seconds",
    "End-to-end latency of an agent request",
    labelnames=["agent_name", "model", "status"],
    buckets=(0.25, 0.5, 1, 2, 5, 10, 30, 60, 120),
)

# Latency of a single LLM API call (one completion request)
LLM_CALL_LATENCY = Histogram(
    "llm_call_duration_seconds",
    "Latency of an individual LLM API call",
    labelnames=["model", "provider", "stream"],
    buckets=(0.1, 0.25, 0.5, 1, 2, 5, 10, 30),
)

# Latency of tool/function calls executed by the agent
TOOL_CALL_LATENCY = Histogram(
    "agent_tool_call_duration_seconds",
    "Latency of a tool call executed by the agent",
    labelnames=["tool_name", "agent_name", "status"],
    buckets=(0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10),
)
```

### Token Usage

```python
from prometheus_client import Counter, Histogram

PROMPT_TOKENS = Counter(
    "llm_prompt_tokens_total",
    "Total prompt tokens sent to the model",
    labelnames=["model", "agent_name"],
)

COMPLETION_TOKENS = Counter(
    "llm_completion_tokens_total",
    "Total completion tokens received from the model",
    labelnames=["model", "agent_name"],
)

CACHED_TOKENS = Counter(
    "llm_cached_tokens_total",
    "Prompt tokens served from KV-cache (provider-reported)",
    labelnames=["model", "agent_name"],
)

TOKENS_PER_REQUEST = Histogram(
    "llm_tokens_per_request",
    "Total tokens (prompt + completion) per request",
    labelnames=["model", "agent_name"],
    buckets=(100, 500, 1000, 2000, 4000, 8000, 16000, 32000, 64000, 128000),
)
```

### Cost

```python
from prometheus_client import Counter

LLM_COST = Counter(
    "llm_cost_dollars_total",
    "Estimated cost in USD for LLM usage",
    labelnames=["model", "agent_name", "cost_type"],  # cost_type: prompt | completion
)
```

### Tool Calls

```python
from prometheus_client import Counter

TOOL_CALLS_TOTAL = Counter(
    "agent_tool_calls_total",
    "Total tool calls made by agents",
    labelnames=["tool_name", "agent_name", "status"],  # status: success | error | timeout
)
```

### Errors and Retries

```python
from prometheus_client import Counter, Gauge

LLM_ERRORS = Counter(
    "llm_errors_total",
    "Errors returned by the LLM provider",
    labelnames=["model", "provider", "error_type"],  # error_type: rate_limit | timeout | 5xx | auth
)

LLM_RETRIES = Counter(
    "llm_retries_total",
    "Retried LLM API calls",
    labelnames=["model", "provider", "retry_reason"],
)

AGENT_ACTIVE_REQUESTS = Gauge(
    "agent_active_requests",
    "Number of agent requests currently in flight",
    labelnames=["agent_name"],
)
```

---

## OpenTelemetry Integration

Use the OpenTelemetry Python SDK to create traces that capture every step of an agent turn: the top-level request, each LLM call, each tool execution, and retrieval operations.

### Setup

```python
# otel_setup.py
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource

def init_tracing(service_name: str, otlp_endpoint: str = "http://localhost:4317"):
    resource = Resource.create({
        "service.name": service_name,
        "service.version": "1.0.0",
        "deployment.environment": "production",
    })
    provider = TracerProvider(resource=resource)
    exporter = OTLPSpanExporter(endpoint=otlp_endpoint, insecure=True)
    provider.add_span_processor(BatchSpanProcessor(exporter))
    trace.set_tracer_provider(provider)
    return trace.get_tracer(service_name)
```

### Tracing LLM Calls

```python
# llm_tracing.py
import time
from opentelemetry import trace
from opentelemetry.trace import StatusCode

tracer = trace.get_tracer("agent.llm")

def traced_llm_call(client, messages, model="gpt-4o", **kwargs):
    """Wrap an LLM completion call with a full OpenTelemetry span."""
    with tracer.start_as_current_span("llm.chat_completion") as span:
        span.set_attribute("llm.model", model)
        span.set_attribute("llm.provider", "openai")
        span.set_attribute("llm.message_count", len(messages))
        span.set_attribute("llm.temperature", kwargs.get("temperature", 1.0))
        span.set_attribute("llm.max_tokens", kwargs.get("max_tokens", 0))

        start = time.perf_counter()
        try:
            response = client.chat.completions.create(
                model=model, messages=messages, **kwargs
            )
            elapsed = time.perf_counter() - start

            usage = response.usage
            span.set_attribute("llm.prompt_tokens", usage.prompt_tokens)
            span.set_attribute("llm.completion_tokens", usage.completion_tokens)
            span.set_attribute("llm.total_tokens", usage.total_tokens)
            span.set_attribute("llm.duration_seconds", elapsed)
            span.set_attribute("llm.finish_reason", response.choices[0].finish_reason)
            span.set_status(StatusCode.OK)

            # Update Prometheus counters
            PROMPT_TOKENS.labels(model=model, agent_name="default").inc(usage.prompt_tokens)
            COMPLETION_TOKENS.labels(model=model, agent_name="default").inc(usage.completion_tokens)
            LLM_CALL_LATENCY.labels(model=model, provider="openai", stream="false").observe(elapsed)

            return response

        except Exception as exc:
            elapsed = time.perf_counter() - start
            span.set_status(StatusCode.ERROR, str(exc))
            span.record_exception(exc)
            LLM_ERRORS.labels(model=model, provider="openai", error_type=type(exc).__name__).inc()
            raise
```

### Tracing Tool Execution

```python
# tool_tracing.py
import functools
from opentelemetry import trace
from opentelemetry.trace import StatusCode

tracer = trace.get_tracer("agent.tools")

def traced_tool(tool_name: str):
    """Decorator that wraps a tool function with an OTel span and Prometheus metrics."""
    def decorator(func):
        @functools.wraps(func)
        def wrapper(*args, **kwargs):
            with tracer.start_as_current_span(f"tool.{tool_name}") as span:
                span.set_attribute("tool.name", tool_name)
                span.set_attribute("tool.args_count", len(args) + len(kwargs))

                import time
                start = time.perf_counter()
                try:
                    result = func(*args, **kwargs)
                    elapsed = time.perf_counter() - start
                    span.set_attribute("tool.duration_seconds", elapsed)
                    span.set_status(StatusCode.OK)
                    TOOL_CALLS_TOTAL.labels(
                        tool_name=tool_name, agent_name="default", status="success"
                    ).inc()
                    TOOL_CALL_LATENCY.labels(
                        tool_name=tool_name, agent_name="default", status="success"
                    ).observe(elapsed)
                    return result
                except Exception as exc:
                    elapsed = time.perf_counter() - start
                    span.set_status(StatusCode.ERROR, str(exc))
                    span.record_exception(exc)
                    TOOL_CALLS_TOTAL.labels(
                        tool_name=tool_name, agent_name="default", status="error"
                    ).inc()
                    TOOL_CALL_LATENCY.labels(
                        tool_name=tool_name, agent_name="default", status="error"
                    ).observe(elapsed)
                    raise
        return wrapper
    return decorator

# Usage
@traced_tool("web_search")
def web_search(query: str) -> str:
    # ... tool implementation ...
    pass

@traced_tool("sql_query")
def sql_query(statement: str) -> list:
    # ... tool implementation ...
    pass
```

### Propagating Trace Context Across Services

```python
# context_propagation.py
from opentelemetry import context
from opentelemetry.propagate import inject, extract
import httpx

def call_downstream_service(url: str, payload: dict) -> dict:
    """Propagate the current trace context to a downstream HTTP service."""
    headers = {}
    inject(headers)  # injects traceparent + tracestate headers
    response = httpx.post(url, json=payload, headers=headers)
    response.raise_for_status()
    return response.json()

def extract_context_from_request(request_headers: dict):
    """Extract trace context from incoming request headers (for the receiving service)."""
    ctx = extract(request_headers)
    token = context.attach(ctx)
    return token  # call context.detach(token) when done
```

---

## Structured Logging

Emit JSON logs for every agent action so they can be ingested by Loki, Elasticsearch, or Datadog.

### Python Logging Configuration

```python
# logging_config.py
import logging
import json
import sys
from datetime import datetime, timezone

class AgentJSONFormatter(logging.Formatter):
    """Structured JSON formatter for agent logs."""

    def format(self, record: logging.LogRecord) -> str:
        log_entry = {
            "timestamp": datetime.now(timezone.utc).isoformat(),
            "level": record.levelname,
            "logger": record.name,
            "message": record.getMessage(),
            "module": record.module,
            "function": record.funcName,
            "line": record.lineno,
        }
        # Merge any extra fields attached to the record
        for key in ("trace_id", "span_id", "agent_name", "model",
                     "tool_name", "request_id", "user_id",
                     "prompt_tokens", "completion_tokens", "cost_usd",
                     "duration_seconds", "status", "error_type"):
            value = getattr(record, key, None)
            if value is not None:
                log_entry[key] = value

        if record.exc_info and record.exc_info[0] is not None:
            log_entry["exception"] = self.formatException(record.exc_info)

        return json.dumps(log_entry, default=str)


def configure_logging(level: str = "INFO"):
    handler = logging.StreamHandler(sys.stdout)
    handler.setFormatter(AgentJSONFormatter())

    root = logging.getLogger()
    root.setLevel(getattr(logging, level))
    root.handlers = [handler]

    # Suppress noisy libraries
    logging.getLogger("httpx").setLevel(logging.WARNING)
    logging.getLogger("opentelemetry").setLevel(logging.WARNING)
```

### Logging Agent Actions

```python
# agent_logging.py
import logging
from opentelemetry import trace

logger = logging.getLogger("agent")

def log_llm_call(model: str, prompt_tokens: int, completion_tokens: int,
                 duration: float, cost: float, status: str = "ok"):
    span = trace.get_current_span()
    ctx = span.get_span_context() if span else None
    logger.info(
        "LLM call completed",
        extra={
            "trace_id": format(ctx.trace_id, "032x") if ctx else None,
            "span_id": format(ctx.span_id, "016x") if ctx else None,
            "model": model,
            "prompt_tokens": prompt_tokens,
            "completion_tokens": completion_tokens,
            "duration_seconds": round(duration, 3),
            "cost_usd": round(cost, 6),
            "status": status,
            "agent_name": "default",
        },
    )

def log_tool_call(tool_name: str, duration: float, status: str, error: str = None):
    span = trace.get_current_span()
    ctx = span.get_span_context() if span else None
    extra = {
        "trace_id": format(ctx.trace_id, "032x") if ctx else None,
        "span_id": format(ctx.span_id, "016x") if ctx else None,
        "tool_name": tool_name,
        "duration_seconds": round(duration, 3),
        "status": status,
        "agent_name": "default",
    }
    if error:
        extra["error_type"] = error
    logger.info("Tool call completed", extra=extra)
```

Example log output:

```json
{
  "timestamp": "2026-03-24T14:22:01.337Z",
  "level": "INFO",
  "logger": "agent",
  "message": "LLM call completed",
  "module": "agent_logging",
  "function": "log_llm_call",
  "line": 12,
  "trace_id": "0af7651916cd43dd8448eb211c80319c",
  "span_id": "b7ad6b7169203331",
  "model": "gpt-4o",
  "prompt_tokens": 1842,
  "completion_tokens": 356,
  "duration_seconds": 2.417,
  "cost_usd": 0.013770,
  "status": "ok",
  "agent_name": "support-agent"
}
```

---

## Grafana Dashboards

### Agent Overview Dashboard

Save this JSON as `agent-overview.json` and import it into Grafana.

```json
{
  "dashboard": {
    "title": "AI Agent Overview",
    "uid": "agent-overview-v1",
    "tags": ["ai", "agent", "llm"],
    "timezone": "browser",
    "refresh": "30s",
    "panels": [
      {
        "title": "Request Latency (p50 / p95 / p99)",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 12, "x": 0, "y": 0 },
        "targets": [
          {
            "expr": "histogram_quantile(0.50, sum(rate(agent_request_duration_seconds_bucket[5m])) by (le))",
            "legendFormat": "p50"
          },
          {
            "expr": "histogram_quantile(0.95, sum(rate(agent_request_duration_seconds_bucket[5m])) by (le))",
            "legendFormat": "p95"
          },
          {
            "expr": "histogram_quantile(0.99, sum(rate(agent_request_duration_seconds_bucket[5m])) by (le))",
            "legendFormat": "p99"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "s",
            "thresholds": {
              "steps": [
                { "color": "green", "value": null },
                { "color": "yellow", "value": 5 },
                { "color": "red", "value": 15 }
              ]
            }
          }
        }
      },
      {
        "title": "Token Usage (prompt vs completion)",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 12, "x": 12, "y": 0 },
        "targets": [
          {
            "expr": "sum(rate(llm_prompt_tokens_total[5m])) by (model)",
            "legendFormat": "prompt - {{ model }}"
          },
          {
            "expr": "sum(rate(llm_completion_tokens_total[5m])) by (model)",
            "legendFormat": "completion - {{ model }}"
          }
        ],
        "fieldConfig": {
          "defaults": { "unit": "short" }
        }
      },
      {
        "title": "Cost per Hour (USD)",
        "type": "stat",
        "gridPos": { "h": 4, "w": 6, "x": 0, "y": 8 },
        "targets": [
          {
            "expr": "sum(rate(llm_cost_dollars_total[1h])) * 3600",
            "legendFormat": "$/hr"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "currencyUSD",
            "thresholds": {
              "steps": [
                { "color": "green", "value": null },
                { "color": "yellow", "value": 10 },
                { "color": "red", "value": 50 }
              ]
            }
          }
        }
      },
      {
        "title": "Error Rate (%)",
        "type": "gauge",
        "gridPos": { "h": 4, "w": 6, "x": 6, "y": 8 },
        "targets": [
          {
            "expr": "sum(rate(llm_errors_total[5m])) / (sum(rate(llm_call_duration_seconds_count[5m])) + 1e-10) * 100",
            "legendFormat": "error %"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "percent",
            "min": 0,
            "max": 100,
            "thresholds": {
              "steps": [
                { "color": "green", "value": null },
                { "color": "yellow", "value": 1 },
                { "color": "red", "value": 5 }
              ]
            }
          }
        }
      },
      {
        "title": "Tool Call Success vs Failure",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 12, "x": 0, "y": 12 },
        "targets": [
          {
            "expr": "sum(rate(agent_tool_calls_total{status='success'}[5m])) by (tool_name)",
            "legendFormat": "ok - {{ tool_name }}"
          },
          {
            "expr": "sum(rate(agent_tool_calls_total{status='error'}[5m])) by (tool_name)",
            "legendFormat": "err - {{ tool_name }}"
          }
        ]
      },
      {
        "title": "Active Requests",
        "type": "timeseries",
        "gridPos": { "h": 8, "w": 12, "x": 12, "y": 12 },
        "targets": [
          {
            "expr": "sum(agent_active_requests) by (agent_name)",
            "legendFormat": "{{ agent_name }}"
          }
        ]
      }
    ]
  }
}
```

---

## Cost Tracking

### Per-Model Cost Calculation

```python
# cost_tracker.py
from dataclasses import dataclass

@dataclass
class ModelPricing:
    prompt_cost_per_1k: float    # USD per 1,000 prompt tokens
    completion_cost_per_1k: float  # USD per 1,000 completion tokens

# Updated pricing as of early 2026 -- adjust to your negotiated rates
MODEL_PRICING: dict[str, ModelPricing] = {
    "gpt-4o":           ModelPricing(0.0025, 0.0100),
    "gpt-4o-mini":      ModelPricing(0.00015, 0.0006),
    "gpt-4.1":          ModelPricing(0.002, 0.008),
    "gpt-4.1-mini":     ModelPricing(0.0004, 0.0016),
    "gpt-4.1-nano":     ModelPricing(0.0001, 0.0004),
    "claude-sonnet-4":  ModelPricing(0.003, 0.015),
    "claude-haiku-3.5": ModelPricing(0.0008, 0.004),
    "claude-opus-4":    ModelPricing(0.015, 0.075),
}

def calculate_cost(model: str, prompt_tokens: int, completion_tokens: int) -> float:
    """Return estimated cost in USD. Falls back to zero if model is unknown."""
    pricing = MODEL_PRICING.get(model)
    if pricing is None:
        return 0.0
    prompt_cost = (prompt_tokens / 1000) * pricing.prompt_cost_per_1k
    completion_cost = (completion_tokens / 1000) * pricing.completion_cost_per_1k
    return prompt_cost + completion_cost

def record_cost(model: str, prompt_tokens: int, completion_tokens: int, agent_name: str = "default"):
    """Calculate cost and record it in the Prometheus counter."""
    pricing = MODEL_PRICING.get(model)
    if pricing is None:
        return
    prompt_cost = (prompt_tokens / 1000) * pricing.prompt_cost_per_1k
    completion_cost = (completion_tokens / 1000) * pricing.completion_cost_per_1k
    LLM_COST.labels(model=model, agent_name=agent_name, cost_type="prompt").inc(prompt_cost)
    LLM_COST.labels(model=model, agent_name=agent_name, cost_type="completion").inc(completion_cost)
```

### Budget Alerting -- Prometheus Rules

Save as `agent-cost-alerts.yaml` and load it into Prometheus or Cortex ruler.

```yaml
# agent-cost-alerts
