inferoute
OpenAI-Compatible Inference Gateway
Single Go binary that fronts multiple LLM backends, including Ollama, vLLM, or hosted providers like Groq, behind one OpenAI-compatible API. Round-robins across healthy backends, fails over on error, streams responses unbuffered, and semantically caches repeat prompts via NuclaDB so identical requests never hit a backend twice.
// why this exists
vLLM, Ollama, and the rest already solved 'run a model.' Nobody had solved 'route across five of them, fail over when one dies, and stop paying for the same prompt twice.' inferoute is the thin, boring glue layer that does exactly that, shipped as a single static Go binary instead of another framework to install.
// numbers, not adjectives
Routing overhead, ab -n 2000 -c 20
| Requests/sec | Mean latency | p50 | p99 | |
|---|---|---|---|---|
| Direct to backend | ~11,400 | 1.8ms | 0ms | 1ms |
| Through inferoute | ~3,450 | 5.8ms | 1ms | 31ms |
Semantic cache, 700ms simulated backend, 20 trials
| Mean latency | |
|---|---|
| Cache miss | 705.7ms |
| Cache hit | 0.80ms |
// how it's built
Routing & Failover
- Round-robins across backends serving the same model
- Health-checked failover with no dropped requests
- Config hot-reload on SIGHUP, no restart
Semantic Cache
- NuclaDB-backed prompt embedding cache
- Streaming SSE responses replayed byte-for-byte on a hit
- Distance-threshold matching, not exact string match
Rate Limiting & Observability
- Per-API-key token bucket, Redis-shared across instances
- Prometheus metrics for latency and cache hit/miss
- Single static binary, one JSON config file