Distributed SystemsGolangInfrastructureLLMs

inferoute

OpenAI-Compatible Inference Gateway

Single Go binary that fronts multiple LLM backends, including Ollama, vLLM, or hosted providers like Groq, behind one OpenAI-compatible API. Round-robins across healthy backends, fails over on error, streams responses unbuffered, and semantically caches repeat prompts via NuclaDB so identical requests never hit a backend twice.

GoRedisNuclaDBPrometheusDocker

// why this exists

vLLM, Ollama, and the rest already solved 'run a model.' Nobody had solved 'route across five of them, fail over when one dies, and stop paying for the same prompt twice.' inferoute is the thin, boring glue layer that does exactly that, shipped as a single static Go binary instead of another framework to install.

~880x faster on cache hits

// numbers, not adjectives

Routing overhead, ab -n 2000 -c 20

Requests/secMean latencyp50p99
Direct to backend~11,4001.8ms0ms1ms
Through inferoute~3,4505.8ms1ms31ms

Semantic cache, 700ms simulated backend, 20 trials

Mean latency
Cache miss705.7ms
Cache hit0.80ms

// how it's built

Routing & Failover

  • Round-robins across backends serving the same model
  • Health-checked failover with no dropped requests
  • Config hot-reload on SIGHUP, no restart

Semantic Cache

  • NuclaDB-backed prompt embedding cache
  • Streaming SSE responses replayed byte-for-byte on a hit
  • Distance-threshold matching, not exact string match

Rate Limiting & Observability

  • Per-API-key token bucket, Redis-shared across instances
  • Prometheus metrics for latency and cache hit/miss
  • Single static binary, one JSON config file