Writing
Notes from Experiential Labs on the gateway, reliability, and the intelligence layer.
How we make the API fast
An inference gateway is one hop in front of the model, and a naive one charges tens of milliseconds for it. Ours charges about one. Sitting next to Postgres, one database call before dispatch, a Rust data plane, an async ledger that took a worker from tens to hundreds of requests per second, and the $0 load rig that keeps the numbers honest.
10 minReliability · August 2026Failing over LLM providers without changing the model
How another authorized provider can serve the same model when a route fails. Which failures allow fallback, how streaming commits a route, how observed and estimated usage affect accounting, and why prices, capacity, and deadlines still matter. From the open-source engine.
12 minReliability · August 2026Hardened to carry all your traffic
How we decide when to retry a provider, switch routes, keep useful cache, or stop. Explore the routing policy, see how token counts and nested fallbacks work, and read the production findings and test results.
55 minPaper · June 2026Size Doesn't Matter: Cosine-Scored Sparse Autoencoders
Inner product scoring conflates directional alignment with input magnitude. We propose cosine scoring for dictionary learning on normalized representations. ICML Spotlight.
2 minPaper · June 2026CLaaS: Continual Learning as a Service for Sample Efficient Online Learning
Online continual learning for deployed LLM agents behind a chat API: experience replay for gradient reuse, superior forward transfer, and less forgetting than in-context learning.
2 min