Skip to main content
Prerequisites: Telemetry introduces the environment variables and telemetry architecture. This page covers metrics collection in detail. Mellea automatically records LLM metrics across all backends using OpenTelemetry. Metrics follow the Gen-AI Semantic Conventions for standardized observability. The metrics API also lets you create your own counters, histograms, and up-down counters for application-level instrumentation.
Note: Metrics are an optional feature. All instrument calls are no-ops when metrics are disabled or the [telemetry] extra is not installed.

Enable metrics

You also need at least one exporter configured — see Metrics export configuration below.

Token usage metrics

Mellea records token consumption automatically after each LLM call completes. No code changes are required.

Token instruments

Token attributes

All token metrics include these attributes following Gen-AI semantic conventions:

Backend support

Note: Token usage metrics are only tracked for generate_from_context requests. generate_from_raw calls do not record token metrics.

Token recording timing

Token metrics are recorded after the full response is received, not incrementally during streaming:
  • Non-streaming: Metrics recorded immediately after await mot.avalue() completes.
  • Streaming: Metrics recorded after the stream is fully consumed (all chunks received).
This ensures accurate token counts from the backend’s usage metadata, which is only available after the complete response.

Latency histograms

Mellea tracks request duration and time-to-first-token (TTFB) automatically after each LLM call. No code changes are required.

Latency instruments

Latency attributes

Histogram buckets

Custom bucket boundaries are configured for LLM-sized latencies:
  • mellea.llm.request.duration: 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30, 60, 120 seconds
  • mellea.llm.ttfb: 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10 seconds

Latency recording timing

  • mellea.llm.request.duration: Recorded for every generate_from_context call, both streaming and non-streaming.
  • mellea.llm.ttfb: Recorded only for streaming requests, measuring elapsed time from the generate_from_context call until the first chunk arrives.
Access latency data directly from a ModelOutputThunk:

Error metrics

Mellea records LLM errors automatically after each failed backend call. No code changes are required. Errors are classified into semantic categories for consistent filtering across providers.

Error counter

Error attributes

All error metrics include these attributes:

Error type categories

The error_type attribute maps exceptions to human-friendly semantic labels:

When errors are recorded

Error metrics are recorded when a backend raises an exception during generation, after the request has been dispatched to the provider. Construction-time errors (e.g. missing API key) are not captured by the error counter.

Cost metrics

Mellea estimates request cost automatically after each LLM call when pricing data is available. No code changes are required.

Cost instrument

Cost attributes

Pricing data

Cost metrics use litellm (mellea[litellm]) as a pricing library. This is independent of the LiteLLM backend — pricing works with any Mellea backend, but cost is only recorded for models that litellm has pricing data for. Local and private model IDs (Ollama, HuggingFace, custom deployments) will log a one-time warning per model and produce no cost metric. Pricing is auto-enabled when litellm is installed. Use MELLEA_PRICING_ENABLED to override:

Custom pricing

Override or add pricing for any model using a JSON file with litellm’s native per-token schema:
Minimal entries with only cost fields are accepted. Errors loading the file are logged as warnings and litellm’s built-in pricing is used as a fallback.

Operational metrics

Mellea records metrics for its internal sampling, validation, and tool execution loops. These counters give visibility into retry behavior, validation failure rates, and tool call health — independent of the underlying LLM provider.

Sampling counters

All sampling metrics include:

Requirement counters

Guardian Intrinsics and metrics: guardian_check(), policy_guardrails(), factuality_detection(), and factuality_correction() are not Requirement subclasses and do not emit mellea.requirement.checks or mellea.requirement.failures metrics. If you migrate from GuardianCheck to Guardian Intrinsics, Guardian-related requirement counters will stop appearing in your metrics. Wrap Guardian Intrinsic calls in a custom Requirement subclass if you need to preserve this telemetry.

Tool counter

Metrics export configuration

Mellea supports multiple metrics exporters that can be used independently or simultaneously.
Warning: If MELLEA_METRICS_ENABLED=true but no exporter is configured, Mellea logs a warning. Metrics are collected but not exported.

Console exporter (debugging)

Print metrics to console for local debugging without setting up an observability backend:
Metrics are printed as JSON at the configured export interval (default: 60 seconds).

OTLP exporter (production)

Export metrics to an OTLP collector for production observability platforms (Jaeger, Grafana, Datadog, etc.):
OTLP collector setup example:

Prometheus exporter

Register metrics with the prometheus_client default registry for Prometheus scraping:
When enabled, Mellea registers its OpenTelemetry metrics with the prometheus_client default registry via PrometheusMetricReader. Your application is responsible for exposing the registry. Common approaches: Standalone HTTP server (simplest):
FastAPI middleware:
Flask route:
Verify with:
Prometheus server configuration:
Access Prometheus UI at http://localhost:9090 and query metrics like mellea_llm_tokens_input.

Multiple exporters simultaneously

You can enable multiple exporters at once:
This configuration prints metrics to console for immediate feedback, exports to an OTLP collector for centralized observability, and registers with the prometheus_client registry for Prometheus scraping. Typical combinations:
  • Development: Console + Prometheus for local testing
  • Production: OTLP + Prometheus for comprehensive monitoring
  • Debugging: Console only for quick verification

Custom metrics

The metrics API exposes create_counter, create_histogram, and create_up_down_counter for instrumenting your own application code. These return no-ops when metrics are disabled, so you can call them unconditionally.

Programmatic access

Check if metrics are enabled:
Access token usage and latency data from a ModelOutputThunk:
The generation attribute is a GenerationMetadata dataclass. Its usage field is a dictionary with three keys: prompt_tokens, completion_tokens, and total_tokens. All backends populate this consistently. streaming and ttfb_ms are set automatically based on whether streaming mode was used.

Performance

  • Zero overhead when disabled: When MELLEA_METRICS_ENABLED=false (default), no auto-registered metrics plugins are active and all instrument calls are no-ops.
  • Minimal overhead when enabled: Counter increments and histogram recordings are extremely fast (~nanoseconds per operation).
  • Async export: Metrics are batched and exported asynchronously (default: every 60 seconds).
  • Non-blocking: Metric recording never blocks LLM calls.
  • Automatic collection: Metrics are recorded via hooks after generation completes — no manual instrumentation needed.

Troubleshooting

Metrics not appearing:
  1. Verify MELLEA_METRICS_ENABLED=true is set.
  2. Check that at least one exporter is configured (Console, OTLP, or Prometheus).
  3. For OTLP: Verify MELLEA_METRICS_OTLP=true and the endpoint is reachable.
  4. For Prometheus: Verify MELLEA_METRICS_PROMETHEUS=true and your application exposes the registry (curl http://localhost:PORT/metrics).
  5. Enable console output (MELLEA_METRICS_CONSOLE=true) to verify metrics are being collected.
Missing OpenTelemetry dependency:
Install telemetry dependencies:
OTLP connection refused:
  1. Verify the OTLP collector is running: docker ps | grep otel
  2. Check the endpoint URL is correct (default: http://localhost:4317).
  3. Verify network connectivity: curl http://localhost:4317
  4. Check collector logs for errors.
Metrics not updating:
  1. Metrics are exported at intervals (default: 60 seconds). Wait for the export cycle.
  2. Reduce the export interval for testing: export OTEL_METRIC_EXPORT_INTERVAL=10000 (10 seconds).
  3. For Prometheus: Metrics update on scrape, not continuously.
  4. Verify LLM calls are actually being made and completing successfully.
No exporter configured warning:
Enable at least one exporter:
  • Console: export MELLEA_METRICS_CONSOLE=true
  • OTLP: export MELLEA_METRICS_OTLP=true + endpoint
  • Prometheus: export MELLEA_METRICS_PROMETHEUS=true
Full example: docs/examples/telemetry/metrics_example.py

See also:
  • Telemetry — overview of all telemetry features and configuration.
  • Tracing — distributed traces with Gen-AI semantic conventions.
  • Logging — console logging and OTLP log export.