Features
What runs alongside the forward path
Forwarding is the fast path. Everything here is evaluated per request, outside the bytes you send.
Features
Failover
Ordered fallback chains across providers, models and keys. Each attempt carries its own timeout budget, so a slow primary cannot consume the whole request deadline.
failover
Load balancing
Weighted distribution across keys and endpoints, with health-aware scoring that adjusts to observed latency and error rates instead of a fixed table.
load-balancing
Circuit breaking
A target that reports a credential, quota, access or rate-limit failure is held out of rotation. It returns to service after a single probe confirms it can serve again.
circuit-breaking
Self-hosted
One static binary. Docker, Kubernetes or bare metal, inside your own network. No external control plane is contacted at runtime.
self-hosted
Analytics
Requests, tokens, cost and latency broken down by provider, model, key and workspace, so spend can be attributed to the team that produced it.
analytics
Capacity alerts
Threshold and trend alerts for quota usage, error rate and p99 latency, delivered to email, Slack or a webhook before a limit is reached.
capacity-alerts
Next step
Talk to us about your architecture
Tell us how your traffic is shaped today. We will walk through where a forward path fits, what the failure policy should look like, and how the gateway would be deployed in your environment.
No pricing on this page. Every deployment is scoped individually.
- A review of your current provider and key layout
- A failure policy matched to your error classes
- A deployment shape for your network