Experiential
9.6k
Blog

Writing

Notes from Experiential Labs on the gateway, reliability, and the intelligence layer.

Performance · August 2026

How we make the API fast

An inference gateway is one hop in front of the model, and a naive one charges tens of milliseconds for it. Ours charges about one. Sitting next to Postgres, one database call before dispatch, a Rust data plane, an async ledger that took a worker from tens to hundreds of requests per second, and the $0 load rig that keeps the numbers honest.

10 min
Reliability · August 2026

Failing over LLM providers without changing the model

How another authorized provider can serve the same model when a route fails. Which failures allow fallback, how streaming commits a route, how observed and estimated usage affect accounting, and why prices, capacity, and deadlines still matter. From the open-source engine.

12 min
Reliability · August 2026

Hardened to carry all your traffic

How we decide when to retry a provider, switch routes, keep useful cache, or stop. Explore the routing policy, see how token counts and nested fallbacks work, and read the production findings and test results.

55 min
Paper · June 2026

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders

Inner product scoring conflates directional alignment with input magnitude. We propose cosine scoring for dictionary learning on normalized representations. ICML Spotlight.

2 min
Paper · June 2026

CLaaS: Continual Learning as a Service for Sample Efficient Online Learning

Online continual learning for deployed LLM agents behind a chat API: experience replay for gradient reuse, superior forward transfer, and less forgetting than in-context learning.

2 min