The same model is often sold by several providers. fable-5 is available from Anthropic directly, on Amazon Bedrock, and through OpenRouter. Each route has its own rate limits, credits, availability, latency, and error formats. Another route can keep a provider outage from becoming a failed request, but only if an authorized, compatible, funded alternative remains and the request still has time to use it. This post explains that bounded fallback, with the code.
The engine behavior below was checked against published release 0.7.117, commit b7ec2801. That identifies the source we reviewed, not the version running on every hosted worker. Hosted accounting and catalog details were checked separately against platform source f7309626c on September 25, 2026; that source check is not a new production acceptance run. Historical implementation examples are labeled separately. The longer reliability article covers cache policies, limits, production observations, and candidate linked-model and recovery work without presenting those candidates as live.
Retries are not enough
When a provider call fails, there are three possible responses. You can retry the same provider, which works for a blip and does nothing for a persistent outage. You can fail over to the same model on a different provider, which is what this post is about. Or you can silently switch to a different model, which is the dangerous one: the caller asked for a specific model, tested against it, and tuned prompts for it. Swapping it under them changes behavior without telling anyone. The ordinary provider waterfall keeps the exact model fixed. Explicit cross-model links are a separate candidate contract, not an accidental consequence of exhausting that waterfall.
Which failures are allowed to fail over
The first question is who owns the failure. A platform-managed provider account can have bad credentials, exhausted funding, a missing deployment, a transient conflict, or an outage. Another authorized route may serve the same valid request. Invalid caller credentials, malformed input, and request-wide security or spending restrictions must instead stop the request. A failed customer-owned provider key does not authorize the gateway to buy that request on a platform account.
HTTP status is only the starting point. Provider 402 means an account funding problem; 408 is a timeout; 409 and 425 can represent transient provider conflicts. A recognized quota error inside a 429 is not an ordinary throttle. A 404 for a caller-supplied response or conversation reference is not the same as a missing provider deployment. The native transport inspects bounded error bodies before deciding, rather than treating every remaining 4xx as invalid input. The native status classification and body-specific corrections show the distinction. These branches from the matching Python contract illustrate why 408, 409, and 425 are not ordinary caller errors; they are excerpts, not the complete classifier:
if status_code == 408: return GatewayFailure( failure_class=GatewayFailureClass.TIMEOUT, safe_message="provider request timed out; retry the request", retryable_same_deployment=True, failover_eligible=True, safe_details=details, ) if status_code in {409, 425}: return GatewayFailure( failure_class=GatewayFailureClass.PROVIDER_INTERNAL, safe_message="provider reported a transient conflict; retry the request", retryable_same_deployment=True, failover_eligible=True, safe_details=details, )
Eligibility is not a promise of another attempt. Authorization, capabilities, funding ownership, the remaining deadline, and the total attempt allowance still apply. Throttles also follow the configured cache and retry policy. Refusals do not fail over by default: a model declining a request is model behavior rather than a provider fault. An explicit refusal-failover option operates inside the streaming boundary described below.
The waterfall
Each direct rung is a deployment: one way to reach one exact model. A configured chain might include a direct provider on your key, a cloud route on a platform account, and another authorized reseller. That is an example, not a claim about the current routes enabled for every model. An ordered set of deployments is an ExactModelPool. Each deployment must carry its pool's exact_model_id, and a multi-deployment pool requires operator certification of equivalence. These checks are in the pool validator:
if len(set(self.deployment_ids)) != len(self.deployment_ids): raise ValueError("exact-model pool deployments must not repeat")if len(self.deployment_ids) > 1 and self.equivalence is None: raise ValueError("multi-deployment pools require operator equivalence certification")if len(self.deployment_ids) == 1 and self.equivalence is not None: raise ValueError("singleton pools must not assert equivalence certification")
The catalog separately rejects a pool containing a deployment for another exact model. Ordinary provider failover therefore chooses a safe retry or another deployment within that model's authorized pool. Cache affinity and health checks can change which usable route is selected first; the configured list alone is not a trace of the attempts a request actually made.
Failover during streaming
The important boundary is when the engine commits model output, not when the client confirms it has read a byte. Text, tool calls, images, exposed reasoning, and output structure can commit a route. The engine makes that decision before the committed prefix is delivered, so an empty client buffer does not prove that retrying remains safe. Stream framing and bookkeeping alone do not commit the route. Private reasoning that is deliberately withheld follows a separate rule and can still permit fallback before outward output is committed.
Before commitment, an eligible failure may retry the same rung or advance within the remaining policy and request budget. Sequence numbers are rewritten into one public sequence. After commitment, the route is fixed: a provider failure can leave a truncated stream, with no resume or silent restart that could replay or reorder output. The commitment predicate lists the output events covered by that rule.
With refusal failover enabled, the engine can withhold a refusal and try another eligible rung before committing it. That buffer is bounded at 64 KiB or 256 events. If adding a refusal event would exceed the bound, the engine commits and flushes the buffered output instead of promising another fallback. A refusal can also be returned when no alternative succeeds. The option is not a guarantee that a refusal will never reach the caller.
Billing through failures
A platform-funded physical dispatch needs a bounded monetary reservation before calling the provider. Each actual attempt has its own usage accounting, and finalization must not charge that attempt twice. Exactly-once accounting does not mean every failed attempt costs zero. Input processing, generated output, or private reasoning can consume provider tokens even when the client receives no complete answer.
On a disconnect, provider-reported counters are retained where available; missing portions can be estimated from the input and output already observed by the gateway. The result is labeled estimated rather than presented as a final provider measurement. Hidden reasoning can be included in those token counts without being sent to the caller or stored in logs. The disconnect tests exercise complete, partial, and missing provider meters. Neither a broken stream nor a missing final usage event proves zero provider cost. A complete engine estimate can settle at the accepted prices. If the meter remains incomplete and is not a complete estimate, the hosted ledger keeps a pending reservation instead of inventing a zero charge. That hold can outlive the stream and carry into later UTC periods until it is resolved. Pending money is not the same as finalized spend, and neither observed nor estimated usage is a verified provider invoice.
Caller retries have their own contract. /v1/chat/completions and /v1/responses accept an Idempotency-Key header. Reusing the key with different content is a typed conflict. Within the bounded in-process replay store, the same key and body can replay a completed response; concurrent callers join the owner. An abandoned owner fails closed rather than silently rerunning the operation. Those guarantees belong to the matching replay identity and that worker's retained entry, not an unlimited cross-worker response archive.
Replaying a stored result does not dispatch another provider attempt or create another usage charge. Each rung also derives an upstream idempotency key, reused on same-rung retries and changed for a new rung. Whether the provider honors that upstream key is up to the provider.
Tool calls across providers
Failover assumes the next provider can preserve the request's required behavior. Tool calls and structured outputs differ across OpenAI, Anthropic, Bedrock, and Gemini wire formats. The engine normalizes supported provider events and re-encodes them into the caller's dialect, such as Chat Completions, Responses, or Messages. API-shaped similarity alone does not establish support for every tool, image, reasoning, or structured-output feature.
Capability declarations and their test evidence therefore matter as much as the route list. A dated certification result belongs to the exact provider, model, surface, and revision tested. An untested cell, including one blocked by missing credentials, is not a passing result. Another route must not silently drop a required feature to obtain an answer.
Knowing a provider is down
Health checks consult in-process records rather than making a database or network call during selection. They still do local counter and lock work. Typical operational failures need two failures before a thirty-second open period and a limited half-open trial. Credential and missing-deployment failures can suppress a route immediately. These records are scoped to the relevant deployment and credentialed connection, not proof that an entire provider is down.
Throttling uses a separate cooldown. Thirty seconds is the fallback when usable provider timing is absent, not a fixed duration for every 429; usable Retry-After is bounded by the health policy. Active throttle windows are not bypassed merely because every route is throttled. Other suppression states have bounded last-resort behavior, but that does not guarantee an available model or justify repeated calls into an active provider throttle window.
When credits run out
Running out of credits looks like a failure but routes differently, and the scope decides. A restriction specific to one provider account or deployment may leave another authorized route usable. An organization, key, or requested-model funding refusal must not be bypassed by choosing a different destination. A customer-owned provider credential or funding error does not silently become a platform-funded request. Where a recorded quota reason is available, the hosted response can preserve that explanation alongside insufficient_quota; a generic message is the fallback, not the only possible response.
BYOK does not spend the platform's provider balance, but it still needs authorization, compatible capabilities, admission capacity, and the customer's own provider allowance. Skipping the platform credit reservation does not mean bypassing all gateway policy or rate checks.
What the ledger keeps
The accounting ledger is content-free: identifiers, route and failure information, token counts, prices, and costs explain an attempt without a prompt or response body in those accounting rows. A request hash is not the request text. That boundary is narrower than saying no part of the system ever retains response content.
Exact replay needs the completed response. The bounded in-memory replay entry contains its status, headers, media type, and body bytes, as the cached-response type makes explicit. Separately, stored Responses conversations follow the organization's content-retention setting. Hosted prompt and response capture uses separate content tables and the organization's privacy setting; zero-data-retention routing has its own capture fence. A test proving that secrets are absent from ledger rows or logs does not prove that replay memory or permitted content storage contains no content.
The gaps
Several limits remain important. Provider equivalence is operator-attested configuration, not a continuous experiment comparing outputs across providers. A final provider label and an attempt count are not a complete explanation of every selection, skip, and retry. Diagnostics and usage records must be read with their actual scopes and timestamps, rather than treating every later attempt as a rescued request.
The non-streaming Chat path requests a provider stream and assembles a completed response. That is an implementation detail, not by itself a measured latency penalty across all API surfaces. Controlled failure tests also do not establish real-provider cache savings, invoice accuracy, or multi-provider capacity under load. Those claims need their own dated experiments on the exact deployed combination.
Hardening the rungs
Failover shipped, and running it in production taught us one lesson twice: a waterfall is only as strong as the metadata on its rungs. A route can be reachable and still be unusable because the gateway cannot bound its cost. Health and pricing are separate checks. A successful health trial does not supply a missing price.
The original pricing defect treated a missing reservation ceiling as zero. That could let a paid request pass checks against balance and spending limits without reserving its exposure. The database now refuses an unknown ceiling before balance, budget, and promotion checks. This excerpt uses its integer nano-USD argument; comments are omitted:
if p_maximum_cost_nano_usd is null then raise exception using errcode = 'P1013', message = 'deployment_price_unknown: this route has no known price, so ' || 'its spend cannot be bounded against the credit balance, a daily ' || 'cap, or a scope budget; it is ineligible and another route may ' || 'serve the request';end if;
An unpriced route is ineligible rather than free. Another priced route may still serve the request, subject to the other authorization, funding, and deadline checks. If no usable route remains, the request fails. Losing one route is preferable to guessing its cost and overspending the caller's allowance.
The released engine's token-priced monetary ceiling includes both estimated input and bounded output. It selects conservative applicable input and output rates, including cache-write and reasoning rates and a reachable long-context pricing tier, then rounds up in nano-USD. Missing a required price still makes the ceiling unknown. It is not simply maximum output multiplied by the output price; the reservation calculation accounts for both sides.
A missing output ceiling also has an implemented rule, not an unfinished rollout. A declared context window can provide a conservative output reservation bound when the explicit output metadata is absent. If neither a declared bound nor a caller-supplied maximum exists, the engine requires an explicit max_tokens rather than guessing unlimited generation. The missing-metadata tests cover both cases. Reserving a context-window bound does not claim that the provider can actually generate that much output for every prompt.
Catalog updates can reach running workers without a restart, but a database edit is not a guarantee of activation within fifteen seconds. The original refresher example below illustrates the watermark check and snapshot swap. It is a historical sketch, not a description of every publication fence or the current host's polling schedule:
class GatewayCatalogRefresher: def __init__(self, *, poll_interval_seconds: float = 15.0): self._poll_interval = poll_interval_seconds def refresh_now(self) -> bool: watermark = read_catalog_watermark(self._connection) if self._state is not None and watermark == self._watermark: return False # nothing changed; no work, no swap ... # reload rows, rebuild the snapshot, swap it in return True
Serving depends on a successfully published compatible snapshot and workers accepting that version. The host's default poll interval is fifteen seconds, while full rebuilds normally have at least thirty seconds between them. Hydrating a published revision can happen without waiting for that full rebuild. Publication fences, retries, and failed refreshes can still delay a change. A metadata correction is not proven live until the serving catalog reflects it; requests already admitted against a snapshot keep their accepted configuration.
Silent providers also have implemented bounds. Connection setup, waiting for the first response bytes, and waiting for model progress are separate phases. The first-token allowance is clamped to leave part of the request budget for an eligible fallback. The actual wait depends on host and route configuration, not a universal thirty-five-second deadline.
Provider pings alone are not model progress. A test with hidden reasoning followed only by pings reaches the bound and uses the next provider without exposing that reasoning. Genuine private reasoning progress has separate handling. None of these timeouts permits switching after output commitment or promises that the next route will succeed.
The failover engine is open source at experientiallabs/experiential. The old 0.5.0, 0.5.1, and 0.5.2 examples describe its early releases, not the current deployment. Native serving now requires the Rust extension: a missing or failed extension is an error, not a switch to a Python data-plane fallback. Released source, hosted activation, and measured production outcomes remain separate evidence.