Observability
Analytics and alerts
Every request produces one record. Records are queryable in the dashboard and exportable over OpenTelemetry.
recorded fields
- Requests
- by provider, model, key, workspace
- Tokens
- input, output, cached
- Latency
- p50, p95, p99, time to first token
- Cost
- per request, per model, per team
- Errors
- by failure class and upstream status
- Health
- per target, rolling window
dashboard (sample)
last 24h- requests
- 1,284,930
- error rate
- 0.31%
- p99 latency
- 1.42s
- cost
- $2,418
+4.1%
-0.12%
+0.04s
+2.8%
latency distribution
| target | share | errors | state |
|---|---|---|---|
| openai / gpt-6-sol | 52% | 0.18% | serving |
| anthropic / claude-sonnet-5-5 | 31% | 0.24% | serving |
| bedrock / claude-opus-5-5 | 14% | 1.10% | degraded |
| azure / gpt-6-astra | 3% | — | held |
sample data, shown for layout
Alerting that fires before the limit does
Rules run against the same counters the dashboard reads, so what you chart is what you are paged on.
Export targets
OpenTelemetry (OTLP), Prometheus scraping, structured JSON logs, and webhooks.
Quota runway
Projects when a key or account will reach its limit, based on the current rate rather than the current total.
Error rate
Fires on a sustained rate per target, per failure class, over a configurable window.
Latency
Fires on p99 regression against a target or against that target’s own baseline.
Health transitions
Fires when a target is held out of rotation or released back into it.
Next step
Talk to us about your architecture
Tell us how your traffic is shaped today. We will walk through where a forward path fits, what the failure policy should look like, and how the gateway would be deployed in your environment.
No pricing on this page. Every deployment is scoped individually.
- A review of your current provider and key layout
- A failure policy matched to your error classes
- A deployment shape for your network