Skip to content
SolidwayLLM gateway

Rust LLM gateway

Never rewrite your traffic.

Solidway sits between your application and your model providers. Requests and responses pass through byte for byte, while failover, load balancing, circuit breaking and telemetry happen around them.

Self-hosted · Runs inside your own network · Single static binary

solidway · forward pathok
01

Client

Your application, any SDK

02

Solidway

Routing, accounting, health

03

Provider

Unchanged request

POST
/v1/messages
model
claude-sonnet-5-5
body
unchanged
stream
passthrough
0
bytes rewritten in the forward path
1
static binary to deploy
3
resilience mechanisms working together

Solidway

Three things Solidway does not do

Most gateways earn their keep by changing your traffic. Solidway is built around the opposite constraint.

  • 01does not

    It does not touch the payload

    Request bodies, response bodies, header order and streaming chunk boundaries are forwarded as they arrived. No schema translation, no dropped fields, no rewritten prompts.

  • 02does not

    It does not surface retries to the caller

    When a provider is slow or unavailable, the request moves to the next target inside the gateway. Your client receives one response.

  • 03does not

    It does not hide what happened

    Latency, token usage, cost and the error class of every attempt are recorded per provider, model and key, and can be exported over OpenTelemetry.

Features

What runs alongside the forward path

Forwarding is the fast path. Everything here is evaluated per request, outside the bytes you send.

  • Failover

    Ordered fallback chains across providers, models and keys. Each attempt carries its own timeout budget, so a slow primary cannot consume the whole request deadline.

    failover

  • Load balancing

    Weighted distribution across keys and endpoints, with health-aware scoring that adjusts to observed latency and error rates instead of a fixed table.

    load-balancing

  • Circuit breaking

    A target that reports a credential, quota, access or rate-limit failure is held out of rotation. It returns to service after a single probe confirms it can serve again.

    circuit-breaking

  • Self-hosted

    One static binary. Docker, Kubernetes or bare metal, inside your own network. No external control plane is contacted at runtime.

    self-hosted

  • Analytics

    Requests, tokens, cost and latency broken down by provider, model, key and workspace, so spend can be attributed to the team that produced it.

    analytics

  • Capacity alerts

    Threshold and trend alerts for quota usage, error rate and p99 latency, delivered to email, Slack or a webhook before a limit is reached.

    capacity-alerts

Architecture

Transparent by design

The request path has three stages, and only the middle one is ours.

  1. client

    Client

    Your application, any SDK

  2. gateway

    rust

    Solidway

    Routing, accounting, health

    • select target
    • record usage
    • judge health
    • preserve bytes
    • primary

      upstream 5xx

      provider-a/gpt-6-astra

    • fallback

      returns response

      provider-b/gpt-6-astra

A request that fails on the primary target is served by the next target in the chain.

Rust

Written in Rust

One statically linked binary. No garbage collector, no interpreter, and no runtime package to install on the host.

  • No garbage collector

    Memory is released at deterministic points, so the tail of the latency distribution is not affected by collection pauses.

  • One file to ship

    The build produces a single binary for x86-64 or arm64. There is no runtime package, no sidecar and no version matrix to maintain.

  • Flat memory under load

    Throughput scales with available cores. Steady-state memory is bounded by the configured connection pools rather than by a heap.

  • Memory safety at compile time

    In safe Rust, the borrow checker rejects the class of defects that produce buffer overflows and use-after-free before the binary is built.

Observability

Analytics and alerts

Every request produces one record. Records are queryable in the dashboard and exportable over OpenTelemetry.

dashboard (sample)

last 24h
requests
1,284,930

+4.1%

error rate
0.31%

-0.12%

p99 latency
1.42s

+0.04s

cost
$2,418

+2.8%

latency distribution

p50p95p99
targetshareerrorsstate
openai / gpt-6-sol52%0.18%serving
anthropic / claude-sonnet-5-531%0.24%serving
bedrock / claude-opus-5-514%1.10%degraded
azure / gpt-6-astra3%—held

sample data, shown for layout

Next step

Talk to us about your architecture

Tell us how your traffic is shaped today. We will walk through where a forward path fits, what the failure policy should look like, and how the gateway would be deployed in your environment.

No pricing on this page. Every deployment is scoped individually.

  • A review of your current provider and key layout
  • A failure policy matched to your error classes
  • A deployment shape for your network