These pages describe the intended behaviour of the gateway. The open-source release and the enterprise edition are being split out; details will be updated when the first release ships.
Metrics
What is counted, at what granularity, and how to scrape it.
Metrics are produced for every request and exposed for scraping. They are off until an exporter is configured.
Counters
| Metric | Dimensions |
|---|---|
| Requests | provider, model, key, route, status class |
| Tokens | direction (input, output, cached), provider, model |
| Attempts | target, outcome |
| Held targets | target, failure class |
| Retries | reason, target |
Latency
Latency is recorded as a histogram so percentiles can be computed at query time.
- Total request duration
- Time to first byte, for streaming responses
- Per-attempt duration
The last one is what distinguishes a slow gateway from a slow provider.
Cost
Cost is derived from token counts and a price table. Where a price is unknown for a model, the request is counted and the cost series is skipped rather than reported as zero.
Exposition
The metrics listener is an infrastructure endpoint, so it lives in the bootstrap file rather than in the console:
[telemetry.metrics]
prometheus = true
listen = "0.0.0.0:9090"
Metrics are also exportable over OTLP. See Logs for the trace side of the same pipeline.
Cardinality
Every dimension above can multiply the series count. Model identifiers and key names are therefore constrained by what is configured in the console rather than accepted from the request, so a caller cannot create new time series by inventing a model name.