Experiential
9.6k
BlogReliability

Hardened to carry all your traffic

August 2026 · 55 min

When a model provider fails, the gateway may send your request to another provider. Whether it should do that depends on what failed. A connection error may justify another attempt. An exhausted customer budget does not. Once you have received part of an answer, starting a different answer on another provider can make the result unusable.

Cache adds another consideration. Returning to the same provider can let it reuse the beginning of a long conversation instead of processing that text again. Switching providers may produce an answer sooner, but require that work to be repeated. The gateway needs to balance that cost against the chance that waiting will help.

This article explains those decisions, the separate calculations behind our limits, and the failures that led us to change the implementation. It also covers our candidate changes for linked fallback models, more careful recovery, and provider status feeds. Those sections describe the proposed behavior and its constraints, not a claim that it is all running in production. Methodology and test status lists the versions and tests behind the article. Our latency work is covered separately in How we make the API fast.

Provider selection and fallback

A request is the operation your client sends to the gateway. An attempt is one call the gateway sends to a provider. If the first provider fails and the second succeeds, you made one request and the gateway made two attempts. Calling the same provider a second time also counts as another attempt.

A route identifies a provider deployment and the account used to access it. A model can have several routes: a direct provider, another cloud serving the same model, or a connection using your own provider key. We call their configured order a waterfall and each entry a rung. The ordinary provider waterfall changes the provider while keeping the exact model fixed. Changing models requires the explicit links described later.

A gateway worker is one running process that handles requests. Before making an attempt, it checks which routes you may use, which support the request, and which have recent failures or capacity limits. Cache affinity can also change the order. The first provider in the configured list is therefore not always the first one called.

Skipping a route is not an attempt. A request can even end without any provider call, for example because the gateway rejected its credentials, capabilities, or budget. This distinction matters when interpreting request counts, provider errors, and bills.

What can trigger another attempt

A connection failure, provider 5xx error, missing provider deployment, or excessive wait for the first response can justify trying another route. A problem with a platform-managed account can also make that route unusable. In these cases, another authorized provider may be able to serve the same valid request.

Invalid caller credentials, malformed input, and organization-wide security or spending restrictions are different. Another provider must not bypass them. A capability missing on one route may leave another suitable route available; a capability missing everywhere does not. A model refusing the content is not a provider outage either. Trying another provider after a model refusal requires an explicit opt-in and is allowed only before the client receives meaningful output.

Bring your own key, or BYOK, also changes who owns an error. If your provider key is rejected or its account has run out of funds, the gateway returns that customer-owned error rather than silently buying the request on our account. Later requests can skip a connection known to be unhealthy. That only removes a route: it does not add a platform-funded replacement you never configured.

Every attempt remains part of the original request. It keeps the same deadline, authorization, and operation identity. The gateway records the actual destination, account, prices, usage, and result for each attempt. Moving to another provider does not restart the deadline or reset how much you are allowed to spend.

Three ways to handle provider throttling

A provider throttle means the provider is not accepting that attempt at its current rate. It usually returns HTTP 429, but the status alone does not tell us why: a provider can also use 429 for a spending cap. Retrying cannot fix an exhausted monthly allowance. The rules below apply to errors classified as throttles, not every 429.

The gateway can return a throttle or try another route. Returning it may preserve useful cache for a later retry. Switching may get an answer now but lose that cache.

The three modes below control this choice when neither a cache threshold nor a scheduled throttle retry is configured. Those two controls are explained immediately afterward.

Maximizing Cache

maximize_cache returns the provider throttle instead of switching. Your client can retry later, with a chance of reusing the same provider's cache. The gateway does not hold this request in a queue, and the mode does not guarantee a cache hit. It still allows fallback for other eligible failures, such as a connection error or 5xx.

Maximizing Availability

maximize_availability allows a provider throttle to trigger another attempt on a different usable route. The next provider may need to process the full prompt. All authorization, spending, deadline, and output-safety rules still apply.

Maximizing Cache Affinity

maximize_cache_affinity also allows throttle fallback. Its additional purpose is to keep successive turns on a consistent provider. It hashes an identifier for the conversation to produce a stable, weighted route order. The algorithm is weighted rendezvous hashing.

If the worker already remembers a valid route for that conversation, that remembered choice takes priority over the hash order. This is a sticky binding. When configured, it can remember a fallback provider for later turns instead of immediately returning to the original provider. Remembering where a request went does not prove that the provider cached it.

Set a cache threshold

A cache threshold lets the decision depend on recent cache use. The worker estimates what fraction of an organization's input tokens were read from cache on a particular route. A value of 0.8 means 80%. This estimate comes from previous completed usage reports, not a cache measurement for the request that just received a throttle.

With no scheduled throttle retries, the rule is simple: at or above the threshold, return the provider throttle; below it, try another eligible route. For a threshold of 0.5:

Estimated cached inputResult on a provider throttle
20%Try another eligible route
50%Return the provider throttle
80%Return the provider throttle

An explicit threshold overrides the throttle decision in all three modes. You can therefore keep affinity-based placement while using this table to decide whether to switch after a throttle. A threshold of zero always returns the throttle under this no-schedule rule, including when the worker has no cache evidence and its estimate is zero. This threshold does not control local concurrency or rate-limit checks.

Administrators can change these settings under Edit policy at /admin/telemetry/models/[slug], or through PUT /api/admin/dispatch-policies/{slug}. For example:

JSON
{    "failover_mode": "maximize_cache",    "throttle_cache_threshold": 0.5}

If neither field has been set, the platform code supplies maximize_cache and a 0.5 threshold. Explicitly setting a mode without a threshold removes that automatic threshold. The mode then makes the throttle decision itself, unless a retry schedule is set.

In an API update, omitting a field leaves it unchanged; null clears its stored value. To select affinity without a threshold, set failover_mode to maximize_cache_affinity and throttle_cache_threshold to null. Supported settings and code defaults do not tell us which settings an operator has selected for every production model.

Wait and retry the throttled provider

A throttle_redial schedule tells the gateway to wait briefly and retry the same provider before moving on. A redial is just another provider attempt after that wait. This gives a useful cached route a limited opportunity to recover within the current request.

The schedule replaces the no-schedule rule above. When its retry allowance is used up, or a wait would exceed the allowed time, the gateway moves to the next eligible route regardless of the cache estimate. It does not check the threshold again and decide to stay. The request can still end with a throttle if no next attempt is possible or the request has used its total attempt allowance. The two paths are separate in the released execution code.

Cache use determines how many of the configured retries an ordinary route receives. At admission, the worker takes the configured maximum N, its cache estimate F, and a positive threshold T:

Python
retry_allowance = min(N, floor(N * F / T))

With three configured retries, a threshold of 0.5, and a cache estimate of 0.2, the allowance is one retry. An estimate of 0.5 gets all three. With no threshold, the route gets the full configured allowance. With no schedule, it gets no scheduled retries. This calculation allocates retries; it does not estimate the provider's remaining quota.

Three cases always get the full configured allowance: the last usable route, the route already remembered for this conversation, and the provider that issued reasoning data the conversation still needs. A worker may have no recent cache samples even when the conversation has useful cache there. The last route also has no alternative left. Those cases motivated the fixes in engine #938 and #942.

The first failed attempt has already used a local rate reservation. If that reservation fills the worker's rate window, rejecting its scheduled retry on the same local limit would defeat the schedule. Engine #970 allows the retry past that rate_limit decision. Concurrency and fair-share checks still apply, as do provider quota, authorization, spending limits, cancellation, and the request deadline.

Retry delays increase exponentially up to the configured maximum. Each actual delay is chosen randomly between half and all of that capped delay, so many throttled requests do not all retry together. This is called equal jitter. A provider's Retry-After can increase the delay, but only within the schedule's maximum. If the provider asks for a longer wait, the gateway moves on instead.

The delay must also fit within the time allowed for a first response, and leave enough of the request deadline for that response. These are bounded waits for an already accepted request, not a general queue. If the request ultimately returns a throttle, its explanation retains the largest observed Retry-After. Clients should respect that value when retrying.

Interactive example · not live traffic

What happens after a provider throttle?

Change the policy to see the next decision. Assume this is the first throttle on this route, no output has been sent, all other checks pass, and every retry in this example also throttles.

Next decision

Try the next route

No scheduled retries. The threshold, when set, overrides the mode.

Effective mode
maximize cache
Effective threshold
50%
Scheduled allowance
0 retries before the total cap
Request budget left
7 of 8 attempts

The estimate belongs to this organization and route on one worker. It is not this request's cache hit. Scheduled retries get the full allowance when the threshold is zero or unset, or the route is last, sticky, or reasoning-pinned. Waits must still fit the deadline. Affinity also changes route ordering; this example shows only the throttle decision.

Stop before retries can change the answer

All retries share one request deadline and one total attempt allowance. The released defaults allow eight attempts overall. An ordinary retry on the same deployment allows two attempts including the original. A throttle schedule can allow up to six retries after the original, but those attempts still count toward the same total of eight. A long waterfall may therefore stop before reaching its last provider.

A retryable provider 408 can retry the same route in all three modes, within the applicable limits. A provider that never sends its first headers or bytes can instead trigger fallback. Neither case permits silently replacing an answer that the client has already started receiving.

The important boundary is the engine committing model output, which can happen before the client reads it. Text, tool calls, images, exposed reasoning, and output structure can commit a route; stream framing alone does not. Deliberately withheld private reasoning has separate handling and need not commit it. Once committed, the gateway cannot start a different answer on another provider even if the client has not yet received the prefix. If the stream then breaks, the answer can remain incomplete. Preventing duplicate charges does not reconstruct it.

Concurrency limits and fair sharing

Concurrency is the number of attempts running at the same time. Before sending another attempt, each worker checks its local counters for that route. It either admits the attempt or tries a later eligible route. This per-route check does not queue the request. The diagnostic reason queue_bound means the route reached its concurrency bound, despite the name. Separately, each worker shares admission capacity across all keys in an organization. A request waiting for that capacity can receive a retryable 429 after 20 seconds. This limits the wait, not the number of queued requests or the lifetime of a running stream.

That admission slot covers the work up to the first reservation, including any native execution-queue wait. It is released while the stream continues. Later fallback reservations and settlement calls are outside that boundary. It is not a fleet-wide concurrency cap or a guarantee that a slow provider cannot occupy other shared resources.

Administrators configure per-route overrides through rungs[].facts in the model policy. dispatch_concurrency_bound sets an explicit bound, and dispatch_fair_share enables weighted sharing. dispatch_requests_per_minute and dispatch_tokens_per_minute set separate local rate reservations. Weighted sharing and these rate reservations are optional; omitting an explicit concurrency bound does not necessarily leave a route unbounded.

The host supplies a default route bound derived from the worker's maximum active requests and a configurable share. The engine rounds that product up, with a minimum of one. With the stock maximum of 64 and share of 0.5, the default is 32 in-flight attempts per route on that worker. Those are configuration defaults, not a measured setting for every deployed worker, an organization queue length, or a fleet-wide cap.

Leave room for existing conversations

A fresh conversation can be sent elsewhere before the concurrency bound is reached. This leaves some capacity for conversations already assigned to the route. The setting is dispatch_fresh_session_spill_fraction.

For example, in affinity mode, a concurrency bound of ten and a spill fraction of 0.8 send new conversations elsewhere once eight attempts are running on the route. Conversations already assigned to it can use the remaining capacity, subject to rate and fairness checks. The decision depends on total current occupancy, not a separate count of fresh conversations. A remembered route, not a particular cache percentage, gives a conversation access to that remaining capacity.

A local decision such as fair_share_shed or rate_limit means the worker chose not to admit an attempt under its policy. It is not a provider 429 and does not say the provider is down.

Share busy routes without wasting idle capacity

Fair sharing gives active organizations a share of a route's concurrency according to their weights. The weight comes first from an explicit manual tier assignment, then the system tier derived from purchases, then the default tier. A Pro badge alone does not determine the weight. Operators can change the weights; seed values are not permanent guarantees of a traffic ratio.

An organization counts as active for ten seconds after an admission or a local policy rejection. Organizations with work still running count too. Recording rejected demand matters: an organization must not lose its share just because the scheduler kept turning its requests away.

Unused capacity remains available to others. One active organization can use the whole route. When another becomes active, the scheduler adjusts future admissions; it does not stop already-running requests to make the ratio exact immediately.

Give useful cache more weight when a route is busy

An optional setting, dispatch_cache_priority_alpha, gives more weight to organizations that have recently used cache on the route. The boost increases as the route gets busier:

Python
effective_weight = base_weight * (    1 + alpha * congestion * cached_fraction)

Here, congestion is the number of running attempts divided by the concurrency bound. The cached fraction is measured separately for each organization, deployment, and credentialed connection on that worker. One organization's cache use cannot raise another's weight. A report from one provider account doesn't tell us whether a different account has used cache.

Each valid usage sample divides cached input by total input. Total input already includes the cached tokens. Cached counts are restricted to the range from zero to total input. The first valid sample sets the estimate directly; later samples update a moving average:

Python
sample = min(max(cached_tokens, 0), total_input) / total_inputelapsed = max(seconds_since_previous_sample, 1)retained = 0.5 ** (elapsed / 600)estimate = previous * retained + sample * (1 - retained)

The 600-second half-life means that, when a sample arrives ten minutes after the previous one, the old estimate and new sample have equal weight. Samples arriving closer together change the estimate less. The calculation uses at least one second of elapsed time.

Reading the estimate does not gradually reduce it between samples. It remains unchanged until another sample arrives, and reads as zero once the last sample is more than an hour old. Sampling can run even when fair sharing or the cache boost is off. This hour of cache history does not reserve capacity for an idle organization: active demand still uses the separate ten-second rule.

With no cache evidence, the fraction is zero. With alpha unset or zero, there is no boost. When nothing is running, congestion is zero and the boost also disappears. None of these weights increase the provider's quota.

Suppose two organizations have base weight one, alpha is one, and the route is 75% busy. One organization has an 80% cache estimate and the other has none. Their effective weights are 1.6 and 1.0. On an eight-slot route, that corresponds to shares of about 4.92 and 3.08 slots. Those are targets for admission, not promises of fractional running requests.

Rounding those shares too early wastes capacity. Consider weights 3:1:1 on an eight-slot route: the shares are 4.8, 1.6, and 1.6. If the organizations currently run four, one, and one attempts, another attempt from the first would bring the total to seven. The other two are each short 0.6 slots. Their combined shortfall is 1.2, which the scheduler rounds down to one reserved slot. Seven plus one fits. Rounding each organization's share first would give a different, less accurate admission decision.

When local limits refuse instead of overflowing

If the waterfall ends with only local policy rejections, the engine may consider a last-resort admission called saturated_overflow. It examines the first bypassed route. A rejection caused by the worker's default route bound is not force-admitted. An explicitly configured saturation="refuse" policy also stops that overflow. Exhausting every route at those bounds can therefore return a retryable refusal rather than send more work into the full worker.

Other eligible local policy rejections can retain the older overflow behavior, so not every authored bound is an absolute cap. The distinction depends on the bound and saturation policy; it is not a universal exception whenever all routes are busy. Overflow still cannot bypass caller authorization, platform spending limits, the request deadline, or an active provider throttle window.

Use provider-reported remaining capacity carefully

Some providers return headers containing a token limit and the remaining allowance. The platform can use them to move suitable fresh text requests before throttling begins. The optional setting is dispatch_headroom_spill_below. Headroom here means the remaining capacity reported by that provider, not the worker's cache estimate or learned RPM limit.

A reading belongs to the actual provider account and route. Readings are ordered by when their settlement first reached the worker, the available proxy for response receipt, and become unknown after sixty seconds. Retrying settlement does not give an old reading a new time. The platform requires two consecutive low readings from the same account before moving traffic. One low reading is marked low_unconfirmed. The newest reading must be fresh; the earlier one need not be. This prevents one bad header pair from redirecting a large amount of traffic.

Once low headroom is confirmed, the platform puts a selected alternative ahead of the primary route and keeps the remaining order. It leaves requests that must return provider-issued reasoning to its issuer alone. Affinity mode makes its own hash-based order, so this platform reorder has no effect there. A headroom_spill record explains a selection decision, not a failed provider call. Platform changes #1874 and #1906 added this policy and the two-reading safeguard.

Cache location and sticky sessions

Provider cache is not stored in the gateway. It can be specific to a model, endpoint, region, account, or provider-defined sharing scope. Two routes offering the same model do not necessarily share cache. Moving between them may require the full prompt to be processed again. The gateway cannot copy one provider's internal cache to another.

A prompt_cache_key helps identify related requests for placement. The gateway combines caller-provided material with the organization and caller identity before hashing it. Two customers choosing the same string do not thereby share a routing identity. The resulting key is a hint, not proof that the actual message prefix matches or that a provider still holds it.

A sticky binding remembers the route a conversation used. Its idle lifetime is configured by dispatch_sticky_spill_seconds. Activity can refresh that idle lifetime, but total age cannot exceed four times the configured lifetime. This prevents a continuously active binding from renewing forever. A failed or throttled route can lose its binding, and a worker restart or eviction can lose the local record.

Four cache questions need separate answers:

Only the third is a measurement of cache use. Even that does not prove the prefix remains cached now. A sticky binding records placement, not a successful cache write.

Continue conversations without misusing reasoning data

Plain conversation text can often be sent to another compatible route. Encrypted or otherwise opaque reasoning returned by a provider is different: only its issuer may be able to use it. Sending that data to another provider just because both support the same API can fail.

The gateway therefore keeps the issuing route first for such a continuation. Its first reservation can bypass local load sheds to preserve that reasoning. A real provider failure can still justify another route. The change in engine #944 allows the remaining compatible routes for the exact same model to be tried, with reasoning that the new issuer cannot use removed. We can't assume other opaque state works across different models, clouds, or API formats.

Stored Responses conversations have a separate rule. The lookup uses the organization, caller identity, and response ID, not the current catalog revision. Changing one model's price must not make an unrelated conversation disappear. Storage follows the organization's content-retention setting. If a response ID is unknown, expired, or was never stored, the client must resend the conversation; the gateway cannot recover content it intentionally did not retain.

Three rate-limit systems

RPM means requests per minute. TPM means tokens per minute. Those names are not enough to interpret a limit. You also need to know who counts it, whether retries count, and whether the count uses estimated or actual tokens. Our workers, the platform's organization limits, and the provider each answer those questions differently.

A cached request can cost less without consuming less provider TPM. A request can be below the organization's allowance but above a worker's local limit. There is no single formula that combines rate, cost, cache, and concurrency into one number.

1. What one worker reserves before sending

Each worker keeps sixty seconds of local reservations for each deployment and credentialed connection. A reservation records that the worker admitted an attempt before sending it. Local RPM counts those reservations, not successful model responses and not an independent count of confirmed network dispatches. A local policy rejection does not add a reservation.

If both a configured RPM limit and a learned limit exist, the lower one applies. If only one exists, it applies. Other workers have their own counters, so this is not a complete view of the provider account.

The local token reservation is estimated whole input plus maximum output. Text uses the packaged o200k_base tokenizer, with additions for message framing, tools and media, plus a 15% margin. Output uses the applicable bounded output ceiling. This is a planning estimate, not a guarantee that every provider tokenizer will count fewer tokens. Cached input receives no discount in this reservation.

Finishing an attempt frees its concurrency slot. It does not refund its local RPM or token reservation; those entries leave the window after sixty seconds. An attempt that reserved a large answer but generated only a few tokens can therefore continue occupying a large token reservation until it ages out.

There is one special case for a large request: if the token window is empty, a reservation larger than its cap can enter. Otherwise that prompt could be rejected forever. The large reservation then occupies the window and can prevent further normal admissions until it expires. This is separate from the last-resort overflow used when every route rejected the request only because of local policy.

How the worker lowers and recovers its RPM limit

When a provider throttles, the worker reduces its learned RPM limit using the recent local reservation count:

Python
learned_RPM = max(0.1, observed_reservations_in_60s * 0.9) # For each elapsed minute without another throttle:learned_RPM += max(1, ceil(learned_RPM * 0.05))# Do not exceed a configured RPM ceiling, if there is one.

For example, 100 reservations followed by a throttle produce a learned limit of 90. A minute without another throttle raises that by five, to 95, unless the configured ceiling is lower. After six hours without a throttle, the learned limit expires. The minimum of 0.1 prevents a permanent zero limit; it is not a promise that the provider accepts fractional requests.

The worker updates the limit when it next checks it, using the elapsed time since the last update. Newly admitted traffic supplies the next result. No timer sends synthetic paid probes. This controller learns RPM, not TPM. It neither raises the provider's token allocation nor discovers its exact account-wide token limit. A small successful request also does not prove that a much larger prompt will fit the provider's remaining token allowance.

2. What the platform counts for an organization

Postgres enforces the organization's platform-funded limits. A key override or gateway tier can determine the applicable limit, but host-funded attempts across all keys in the organization contribute to the count. Creating more keys does not create separate organization capacity. BYOK attempts do not consume these host-funded windows; they still face worker policy and their provider's limits.

Organization RPM counts platform-funded dispatches in the trailing sixty seconds. For a limit of L, the exact check looks up the Lth most recent dispatch and asks whether its timestamp is still inside that window. For an organization with prior recent traffic, its first receipt starts a 61-second warm period using the older conservative boundary while that traffic ages out. New or idle organizations can use the exact check immediately. Each actual fallback attempt counts; inspecting or skipping a route does not.

TPM and tokens per hour combine finished weighted usage over minute and hour windows with full input reservations that remain outstanding. Those reservations include work still running and pending meters that have not been resolved after an attempt stopped. Finished daily usage follows the current UTC day, not a rolling twenty-four hours; unresolved holds continue to count in current-period gates even across a UTC boundary.

Finished usage can receive a cache discount in these token windows. Let I be total input, C the cached input already included in I, and O total output. Reasoning tokens already included in output must not be added again. The calculation is:

Python
# If input_price > 0 and a cache-read price exists:discount = 1 - min(cache_price, input_price) / input_pricewindow_tokens = I - min(C, I) * discount + O # Otherwise discount = 0.

Use the prices that apply to that attempt. If its input crosses the long-context pricing threshold, use that schedule rather than the base prices. A missing cache price or zero input price gives no discount. A cache price above the input price cannot make one cached token count as more than one fresh token. The database functions gateway_attempt_cached_input_discount and gateway_attempt_window_tokens define the same calculation used to explain a refusal.

Suppose an attempt reports 100,000 input tokens, including 90,000 cache reads, and 2,000 output tokens. If cached input costs one tenth of fresh input, the window counts 21,000 tokens: 10,000 fresh, 9,000 discounted cached, and 2,000 output. If its long-context schedule instead prices cached and fresh input equally, the count is 102,000. A discount on a different pricing tier does not apply.

While input remains reserved, the organization token counter uses that full input without assuming a cache hit. Unlike the worker's token reservation, this token-window hold does not include maximum output. A complete engine estimate can settle at the accepted prices; an incomplete meter without such an estimate keeps a pending hold. Ending a stream is not enough to clear it or report zero spend. The token gate sums weighted usage and outstanding input, rounds the total once to a whole token, and refuses the next dispatch when that existing total is at or above the limit.

It does not require the next attempt's entire worst case to fit under the remaining token allowance. One request can cross the limit, and concurrently finishing streams can add more output. Monetary reservations are a separate check; this token-window rule alone is not a hard bound on the eventual total.

Promotion allowances have additional rules

Promotions can have separate free-input and free-output allowances, per model and across the promotion. Their hour and day boundaries use UTC. Cached input receives the applicable price discount unless the promotion explicitly makes it free. Output has its own allowance and reservation checks. These rules should not be inferred from the organization token-window formula above.

Customer promotion limits are shown in tokens. When a limit was configured in money and converted to tokens, the display marks the token number as approximate. A request explicitly asking for the free-only option must stop when that allowance is exhausted, not quietly continue using credit.

3. What the provider counts

The provider enforces its own account, model, region, and deployment limits. Direct OpenAI says cached tokens still count toward TPM. Most direct Claude models exclude cache reads from their input-token rate limit, but include fresh input and cache writes. The documented Haiku 3.5 cache-read exception is historical; that model is retired on the direct service. Claude's output-token rate limit counts generated output, not the requested max_tokens ceiling.

These provider rules do not change the gateway's conservative reservation. They also should not be assumed for the same model hosted on Bedrock, Vertex, or another reseller. A cache hit can lower the bill while leaving OpenAI TPM unchanged. On another provider, cache reads may also leave more input-rate capacity. Provider headers describe that provider's counters, not our local reservations.

Illustrated example · not live data

Cache changes usage, not the upfront reservation

Worker reservation108k
100,000 input + 8,000 reserved output. Cache does not reduce this reservation.
Organization · no cache reads102k
100,000 input + 2,000 actual output = 102,000 tokens.
Organization · 90% cache reads21k
10,000 fresh + 90,000 cached × 10% + 2,000 actual output = 21,000 tokens.
Numbers and assumptions
CounterTokensCalculation
Worker reservation108,000100,000 input + 8,000 reserved output. Cache does not reduce this reservation.
Organization · no cache reads102,000100,000 input + 2,000 actual output = 102,000 tokens.
Organization · 90% cache reads21,00010,000 fresh + 90,000 cached × 10% + 2,000 actual output = 21,000 tokens.

These are worker reservations and the platform-funded organization token window, not provider quota counters or dollars. We use the same input estimate in both systems to show the accounting difference. Actual tokenizer counts can differ. While a request is running, the organization counter holds full reserved input.

One 100k-token prompt. Reserve 8k output; generate 2k. Cached input is priced at 10% of fresh input in this example.

Explicit model fallbacks

If all providers for a model are unavailable, you may prefer another model to a failed request. That must be an explicit choice. Two models accepting the same API format are not necessarily interchangeable. Our candidate linked-model design adds that choice to the waterfall. The behavior below is the design we are testing, not a report that cross-model fallback has been enabled in production.

Activation needs more than a new engine package. Every writer and serving worker must understand the chain, and an older or rolled-back worker must not serve an obsolete plan that drops its model links. We keep activation separate from the code merge and verify it before describing the feature as live.

Each model may contain one reference to another model, after at least one direct provider or BYOK rung. The reference enters the other model's own waterfall; it does not copy that model's provider rows. The child can itself reference another model. Each model keeps its own failure mode, cache threshold, and throttle retry policy. A reference never overrides a refusal that must stop the request.

Allow reciprocal links without repeating models

Model A may reference B even if B references A. The gateway prevents a loop while executing the request, rather than banning that configuration. It remembers each model's canonical ID, so using another alias for the same model does not evade the check.

Illustrated example · not live data

Try B, skip the loop, then finish A

MODEL AMODEL Benter Bcontinue Bresume A① a1② b1↩ link to ASKIPPED③ b2④ a2a2 waitsuntil B ends
Example assumptions

Candidate behavior, not a production trace. a1, b1, and b2 fail before output; a2 succeeds. Each failure allows advancement, with no same-provider retries. Both models are authorized and root funding has passed. All calls share the request's deadline and attempt budget. No fallback is allowed after meaningful output.

A's link enters B. B's link back to A is skipped—not restarted. After B's providers fail, continue at a2.

A later request that starts on a remembered child must likewise not use a back-reference to restart models already behind its starting position.

Expansion is limited to sixteen distinct canonical models and 256 examined or expanded rungs, in addition to the request-wide deadline and attempt budget. Exceeding a configuration bound is an error, not permission to silently omit the remaining configuration. A provider retry inside a stage is separate from entering the same model again.

Cache affinity and fair sharing can reorder providers only within the group before or after a model reference. They cannot move a provider across that reference. Otherwise a scheduling preference would change when the caller explicitly chose to switch models. Recovery never sends an in-progress request backward through the chain.

Check the actual request at every destination

Each destination must be authorized and support the request's tools, images, structured output, reasoning, and context and output sizes. Provider-bound continuation data also needs compatible handling. The gateway must not silently remove required behavior just to make another model accept the request. The linked-model contract covers conversational Chat, Responses, and Messages requests, not batches, image generation, or embeddings.

Save the configuration as one change

The API and database enforce the same rules as the editor. Every rung has exactly one target: a deployment, a BYOK connection, or a model. They check access to the target, allow at most one model reference, and require a direct first stored rung. A positive position number is not enough: deleting earlier rows can leave a reference first even though its stored position is still greater than zero.

Replacing a chain happens in one database transaction. Deleting the old rows and inserting the new ones in separate transactions would expose an incomplete chain between the two operations. Saves also carry a revision number. If two editors both load revision seven, the second save must not overwrite the first without detecting the change.

Submitting an empty list explicitly resets an organization's override to the default. The database keeps the revision even after that reset. Otherwise an old editor holding revision zero could become valid again after someone else created and then cleared an override. Keeping this small revision record prevents that stale save.

No override, an empty authored chain, an available chain, and an unavailable chain are different states. Deleting a credential must remain possible even if a chain depends on it. That deletion must not make a model reference the first rung or erase an unusable override into an inherited paid default. The editor also refuses to switch off its last enabled direct route. The user must repair the chain or deliberately reset it, rather than have the system choose a new funding source for them.

Keep identity and order correct in the editor

A private model can have the same slug as a public model. Looking up a referenced child by slug alone could therefore show or serve the wrong model. The candidate carries the canonical ID through navigation, page loading, and the API request, then checks the returned identity. Display names and promotion information still use the resolved model's appropriate catalog fields.

Some provider rows may be hidden in the editor but still enabled. Combining a visible edit with those rows must preserve the groups separated by model references. Toggling a visible provider must not accidentally move a hidden provider ahead of the child model. Disconnecting a connection must also refresh the chain's revision, not just its visible detail, so the next save uses the right version.

The existing public-model system does not support an organization saving only a different order for platform-owned routes: it has no separate serving plan for that preference. A supported BYOK chain does produce one. New BYOK connections require Pro unless the organization already has a provider connection. Managing or using an existing connection does not create a new Pro requirement.

Keep the accepted routing configuration stable

The catalog is a published snapshot of models, provider routes, capabilities, and policies. An organization overlay adds that organization's connections and chain choices. The candidate fixes one view of each when resolving a request. A policy edit in the middle of the request must not change what its later child will do.

Every stage carries its own policy, even if it has only one provider. A one-provider child must not lose its cache threshold simply because it needs no provider-selection algorithm. The snapshot also identifies the selected pool of routes. The model ID alone may not be enough when several authorized pools serve that same model.

Older accepted requests keep access to their original indexed catalog generation. They do not read today's mutable policy under yesterday's catalog hash. Organization-specific variants are derived from the fixed overlay instead of publishing a copy for every organization and model. Resolving a request must not rebuild the catalog or traverse an unlimited graph; that work would occupy the limited callback threads needed to serve other requests.

Enforce both the requested and destination budgets

The original model is the root; a referenced model is a child. A request for A that eventually runs on B still needs to obey A's request-level budget and funding rules. The actual attempt must also obey B's applicable model, pool, and deployment budgets. If the same budget applies through both paths, it must not be counted twice.

Each provider call has its own reservation and records its token usage at the destination's prices. Passing through three parent chains must not create three charges for that call. A failed attempt that consumed tokens still matters; a skipped reference consumed no provider tokens. The logical request is finalized once after its attempts finish.

B's prices and promotions apply to B's attempt. A discount on A does not become a discount on B. Switching models must not bypass a free-only request, a root payment refusal, or another request-wide restriction. Skipping an unhealthy root route doesn't prove that the caller could have paid for it.

This is an important limitation of the current candidate. Before its first platform-funded child attempt, it requires evidence of a valid platform-funded admission for the root. An attempt number, a BYOK attempt, a failed reservation, or another organization's attempt does not qualify. If every root route was skipped, starting directly on a platform-funded child is refused.

Supporting that case needs a trustworthy root-funding check that makes no reservation and changes no state. Creating and rolling back a fake root attempt is not the same thing: it can distort accounting. The candidate does not claim that child-first case works without this check.

Destination daily spend is added through the existing once-only usage finalization, using bounded counters. It must not restore large history scans under the organization lock. Tests cover root and destination being the same model, repeated settlement, billed failures, UTC day changes, renamed aliases, and overlapping budgets. A retained hash used to match alias history still needs access restrictions and deletion review. Hashing makes it pseudonymous, not anonymous.

Report the model and the cost that were actually used

The response's model field retains the requested alias, while x-gateway-canonical-model identifies the actual model that produced the committed response. These serve different purposes: one identifies what the caller requested, the other what served it. That header is not a complete stored trace of every child, skip, and recovery decision. The current request detail does not provide that full traversal history.

Cost also has several meanings. List-priced usage is not necessarily the amount charged to the customer: a promotion may pay some of it. A BYOK attempt can have zero platform charge but a nonzero provider bill. Failed and cancelled attempts may consume tokens too. Prices must come from the actual attempt, including its cache and long-context rates. A list-price calculation is not a verified provider invoice.

Returning to a recovered provider

Suppose a conversation moved from provider A to B after A failed. B now works and may hold useful cache. A later request should not return to A merely because some time passed. The candidate recovery design asks two questions: has the reason for leaving A cleared, and is this conversation's cache on A still likely to be useful? In the candidate, recovery has its own runtime wiring and verification. Switching off model-reference authoring does not, by itself, switch off recovery.

The decision applies between requests. It never restarts a request already working through B's waterfall. A successful return can restore the preferred route for later turns. If the trial fails, the healthy fallback remains the remembered choice.

Require evidence that matches the failure

A transport error, timeout, or 5xx can affect shared infrastructure. Later successful traffic can help show recovery, but only for the matching provider, model, endpoint, and region. A success from another region is not evidence that the failed region works. Newer failures, old observations, an active cooldown, or an already claimed recovery trial can prevent a return.

A throttle depends on the relevant account or model limit. That window and Retry-After must permit another attempt, and local admission must allow it. Another account succeeding does not replenish yours. Authentication, account funding, and model-access failures similarly need evidence for the affected credentials and model, or a credential rotation. A provider-wide green status is not enough.

If the gateway left because its own worker was full or its fair-share policy rejected the attempt, the relevant evidence is local capacity. A remote provider success says nothing about a worker's free slots. A request-wide security, payment, or budget refusal is not temporary provider failure and must not be bypassed at all.

Require this conversation's cache evidence

Other customers can help demonstrate that shared infrastructure has recovered. They cannot prove that your credentials work or that your prompt is cached. A return also needs this conversation's own successful cache-use evidence, a matching actual prefix, the same model and credential context, and a conservative cache lifetime. The route must still be authorized and compatible with the request.

If A works again but its cache is cold or unknown, the design keeps using a healthy, likely warm B until the bounded sticky lifetime expires. After expiry, normal health-based selection resumes. Expiry is not proof that A recovered. Unknown cache support, missing history, a worker restart, or an evicted record must remain unknown.

Illustrated example · not live data

A recovers. This request stays on B.

A · preferredB · fallback
A failsA recoversNext request
Current requestLater request

A's cache likely warm

BBAtrial eligible*

A's cache cold or unknown

BBBkeep fallback
*What makes a trial eligible?

Candidate behavior. Recovery must match the original failure, with no newer contrary or incomplete evidence. This conversation needs matching prefix, account, and cache-use history. Authorization, capacity, cooldown, and a one-use recovery lease must allow the trial. It is not a guaranteed cache hit. The fallback is healthy and warm here; after sticky expiry, ordinary health-checked selection resumes.

A return can happen on a later request only when recovery evidence and this conversation's cache history both support it.

Check recovery separately from the circuit breaker

A circuit breaker records recent route failures and temporarily avoids more calls to that route. An open circuit means ordinary selection avoids it; after a cooldown, a limited half-open trial can test it. A closed circuit permits normal calls again.

Native health records are separate for the catalog identity, deployment, and credentialed connection. Typical operational failures need two failures before a thirty-second open period. Authentication, missing-model, and account-quota failures can open immediately. Throttles use their own bounded cooldown and usable Retry-After information. Existing last-resort and forced paths mean an open circuit is not an absolute promise that the route will never be called.

Permission to call A again is not enough reason to leave a healthy, cached B. The extra recovery check can also reject a return sooner than the breaker does. One new local failure must outweigh an older shared success even when two failures are normally required to open the circuit. Otherwise the system would ignore the failure it just saw.

Keep recovery records small and separate from serving

A recovery record needs to explain why the route was left, when it failed, what later evidence cleared that failure, and which cache observation justified returning. Records must distinguish a repeated model being skipped, a warm fallback being retained, a failed recovery trial, cache expiry, and an exhausted chain.

Workers publish small, time-limited outcome summaries in the background, not during the inference call. Shared summaries can describe infrastructure health; they do not share tenant cache history or provider secrets. Recovery trials need limited, expiring permission so workers do not all return at once. No paid synthetic probes run by default.

Recovery bugs we found and corrected

We found the following defects while implementing and testing the candidate recovery logic. They are not a list of measured production incidents. Each could cause a later request to return too soon or rely on evidence that was no longer valid.

Activity could keep old history alive forever

Refreshing a recovery record also refreshed its maximum age. A busy conversation could therefore keep an old record indefinitely, even though the cache lifetime was supposed to be bounded.

The fix keeps the original hard expiry and refreshes only the idle timeout. The ordinary released sticky-binding code already has a four-lifetime cap. This was a separate fix to the new recovery history, not the first introduction of a cap on sticky sessions.

An old success could hide a newer failure

A shared success said a route worked. Another session then failed on that same route, but the older success still permitted a return. Waiting for the ordinary circuit to open would leave a gap in which that return could happen.

The recovery check now considers the latest relevant negative evidence even below the breaker's failure threshold. The newer failure blocks the return until sufficiently new, matching evidence clears it.

A repeated cache key could conceal a changed prompt

A client can keep the same cache hint while changing the messages. Treating that hint as the actual prefix could make the gateway assume useful cache existed for text the provider had never seen.

Admission and settlement now derive compatible identifiers from the actual prefix and match them to this conversation's successful cache use. A matching hint from another tenant cannot supply that evidence.

A request could use permission that was never committed

A worker could expose permission for a recovery trial before the database transaction committed. Another request could use it, then the transaction could roll back. The trial would have relied on a decision that never became durable.

The worker now publishes that permission only after commit. It belongs to that worker, can be used once, and has a fixed expiry. Reading it or retrying publication does not extend its lifetime.

A live worker could still have missing failure reports

A fresh heartbeat says a worker is alive. It does not say all its outcome reports arrived. If a failure report was missing while an older success remained visible, another worker could incorrectly decide it was safe to return.

The publisher tracks reports that are pending separately from reports that completed. Recovery waits when relevant evidence is incomplete. This includes the first failure for a combination of route and credentials that had no previous record. Waiting for recovery evidence does not stop ordinary serving or make the worker unready.

Clock differences could put success after failure incorrectly

If one worker's clock is one second fast and another's is one second slow, their timestamps can reverse the apparent order of events. A success that actually came first could appear to clear a later failure.

Cross-worker comparisons now require a strict two-second uncertainty margin and use conservative success timestamps. If the order is ambiguous, the success does not justify returning. Missing or expired records, unused expired permissions, and restarts likewise do not manufacture evidence that a cache survived.

Provider status feeds

An official status page can help explain several failures at once. It does not test your key, private deployment, remaining quota, or prompt cache. Our candidate integration is therefore visibility-only: it can show incident information, but cannot reorder routes, open or clear a circuit breaker, or authorize a recovery trial.

Use the feed that actually covers the service

The candidate configures eight fixed sources and explicit regional Bedrock feeds rather than assuming every provider exposes the same status API. A configured source still needs a successful fresh fetch and a verified service mapping before it can describe a route:

We didn't find a machine-readable feed with verified coverage for xAI, Z.ai, Qwen, TokenHub, Meta, Wafer, or Gemini Developer API. A feed may exist, but we haven't verified it for this integration. We don't treat missing coverage as a sign that the service is healthy. Private endpoints need their own checks too.

Say what was reported, and when it was checked

The candidate projection defines four states: an incident was reported; no incident was reported; the information is stale or unavailable; or no verified feed covers the route. It supplies the official link, the service or region covered, and freshness for a future display. Attaching that result to the catalog and customer interface is separate work, not something a parser test proves is live. A failed fetch keeps previous facts marked stale; it does not resolve an incident.

An empty successfully fetched RSS feed has no entries to report, not proof that every resource is healthy. The candidate shows recent RSS entries from the past twenty-four hours as reports, not authoritative verdicts that an outage is still active. Older entries are not treated as current incidents. Parsers must also handle optional fields without losing actual incidents: an omitted empty list is not a reason to discard a nonempty list in a valid response.

Poll in the background

The candidate schedules JSON sources roughly every five minutes and RSS sources every fifteen. Database cron calls an authenticated web relay, which calls the backend poller. Each environment needs its own endpoint and cron-secret configuration. Without those settings, scheduled polling does nothing; an authorized direct poller call is a separate path. Applying the migration alone does not activate hosted polling.

The intervals are not exact promises. Cron checks on fixed five-minute ticks, but a source's next due time starts when its previous fetch finishes. A fetch that finishes just after a tick can miss the next eligible tick. Even successful polls can therefore stretch the five- and fifteen-minute intervals toward ten and twenty minutes; failures and backoff can delay them longer. The candidate projection marks JSON stale after fifteen minutes and RSS after forty-five.

Conditional requests avoid downloading unchanged content. An HTTP 304 means the source has not changed and was checked again; it updates fetch freshness, not the incident's event time. It is not evidence that your particular route recovered.

No status fetch runs during an inference request. Every fetch has a timeout and body-size limit. The poller validates DNS results and redirects and uses a safe XML parser. Only the current poller for a source may save its result, so a delayed older poll cannot overwrite newer information. Backoff and Retry-After limit repeated calls to a failing status service.

A DNS failure affects that source, not the whole batch. A 200 response containing an HTML login page is not valid JSON or RSS evidence. Fetch and read size are bounded: the parser accepts at most 2 MB and 1,000 facts, and reads return at most 1,000 facts per source. That is not a hard thirty-day cap on all stored history. Cleanup deletes facts absent from the latest feed once they have not been seen for thirty days; an entry continuously present in a feed can remain longer. Structured feeds also avoid the need to scrape page layouts or parse subscription emails. No inbox or provider subscription was configured for this implementation. A guessed webhook URL returning 404 would not prove that a page lacks subscriptions; those must be checked separately with the page owner.

Do not reveal a masked provider through its status link

A public route may intentionally show a platform label instead of the underlying provider. Attaching that provider's status URL or a revealing component name would undo the masking. Before this candidate is attached to customer-facing routes, that integration must omit the underlying status link and revealing details on masked public routes. Appropriate status for a customer's own BYOK connection or an unmasked route is a separate display decision. The projection alone does not establish that those UI and masking rules are wired. Internal routing and accounting must keep the actual provider identity either way.

Gateway availability

Provider fallback cannot help if the gateway itself cannot serve the request. Multiple workers allow other workers to accept traffic when one disappears. Readiness checks and graceful draining reduce disruption during releases. They do not recover a stream owned by a dead process, and a single-region deployment is not multi-region disaster recovery.

The serving implementation is native Rust. There is no hidden Python fallback for an incompatible route. A model that cannot be served natively is marked unavailable while other usable models can continue serving. Workers read published catalog snapshots rather than building the whole catalog inside a request callback. Platform-funded attempts reserve money before calling the provider.

Keeping connections warm, reducing database round trips, and moving unnecessary work out of requests also helps reliability. Each request occupies less shared capacity, leaving more room for bursts. Load tests can measure that benefit, but a result with simulated providers does not establish infinite capacity or recovery from a real broken stream.

Failures and fixes

The same visible symptom can have different causes. A provider may throttle while the gateway is otherwise healthy. A database bottleneck may delay requests before any provider is called. The examples below explain what changed and why another provider can help in some cases but not others.

Sol and Astra needed different explanations

These figures summarize roughly six hours of production traffic. They show what happened during that period, not whether a particular fix improved performance. No customer content or identifiers are included.

The request and attempt counts will not match exactly. Request totals include only requests with settled usage records, while attempt totals can include work still running. The two reports also use slightly different time windows. The collection time and release details are in Methodology and test status.

ObservationGPT-5.6 SolGPT-6 Astra
Requests with settled usage records35,6154,370
Requests not marked completed152953
Requests refused by gateway quota rules74396
Provider attempts35,9324,158
Attempts throttled by a provider2380
Attempts after a request's first attempt241161
Later attempts marked completed9377

Sol had 238 provider-throttled attempts. Its decisions included 199 waits before retrying, recorded as throttle_backoff, and 37 decisions to try another route without preserving its cache, recorded as throttle_failover_cold. That is evidence of provider throttling, not evidence that 238 customer requests ultimately failed.

Astra had no provider-throttled attempts in this report. It did have 105 last-resort local overflow decisions, 17 concurrency-bound decisions, and one fair-share rejection. The diagnostic names are saturated_overflow, queue_bound, and fair_share_shed. Changing the provider-throttle threshold alone cannot fix a full worker or an organization quota refusal.

Several labels need care. The API field called failedincludes cancellations and other outcomes not marked completed, so the table does not call them all server errors. The gateway quota row counts quota_exceeded usage events, not every HTTP 429. A later attempt can retry the same provider; 241 later Sol attempts do not mean 241 cross-provider failovers or 93 rescued requests.

Similarly, house_account_exhausted can mean that the selector moved a depleted platform account behind better-funded choices, while retaining it as a last resort. That label can appear on the first actual attempt. It is not a count of upstream balance errors. The report's current provider list is also not a stored record of the order used by each request.

When checked, Sol selected maximize_availability, no cache threshold, and three scheduled retries with a 500 ms base delay and 8,000 ms maximum. Astra had not explicitly set a mode or threshold and had the same retry schedule. Those settings were observed at the time of the read; they may not have been constant for all six hours.

A provider funding error looked like invalid client input

A Sol provider returned HTTP 400 because its account quota was exhausted. The gateway classified that response as invalid request input and stopped. That was the wrong action: another authorized provider could potentially serve the valid request, but the error classification prevented fallback and hid the account-quota signal from health handling.

The correction recognizes specific funding phrases in the provider dialect's error message and records provider_quota. It also unwraps a known reseller error format so the original quota, throttle, refusal, or client error is not replaced by a relay's decoding error. The example and fix are documented in engine #993. That incident description comes from the change review, not a per-request investigation of the aggregate table above.

The fix does not retry every HTTP 400 or every message containing “quota.” A provider account running out of money is provider_quota; the gateway refusing a caller's spending allowance is quota_exceeded. The first can justify fallback on an authorized platform-funded route. The second must not be bypassed. Customer-owned credential errors retain their own rules.

Check whether each destination can fit the prompt

A total context window is not an input budget. Sol's catalog entry, for example, lists a 1,050,000-token total, a 922,000-token maximum input, and a 128,000-token maximum output. Configuring a client to use the total as input space can grow a conversation beyond what the provider accepts. Even compaction may fail if it sends the same oversized conversation back to the model.

Different provider routes can also have different context limits. Engine #995 checks a lower-bound text-size estimate plus the actual output budget against each declared window. It includes any route-specific minimum output requirement, checks again after adapting the request for a route, and prepares the surviving routes again as needed. It does not silently reduce the caller's requested output.

If a route declares no window, this check does not exclude it. If no candidate can fit, it returns context_length_exceeded before reserving an attempt. The text estimate is not a complete guarantee for every tokenizer, image, or audio payload. Both #995 and the quota fix #993 merged after the 0.7.82 release we checked. We can't use that deployment to verify either later fix.

Repair rejected reasoning without repeating visible output

A provider can reject an encrypted reasoning item in a continued Responses conversation. The engine can repair that specific problem once, rather than treating it as an ordinary outage. Recent changes recognize the supported reseller error formats and catch the rejection even when it arrives inside a stream before meaningful output.

Remembering such a rejection needs the right owner and lifetime. It belongs to the admitted caller, not every customer. A cleanup task holding an old expiry record must not erase a rejection learned again more recently. Those fixes are covered by changes from #965 through #990. None permits replay after visible text or a tool call has already reached the client.

Do not confuse empty output with a dead provider

An empty completion is different from a connection failure. Engine #971 gives it its own class, so it does not mark shared infrastructure unhealthy or automatically retry the same route. If the text waterfall is exhausted, the result remains a typed empty completion. Image output has a separate immediate-return rule in #977.

A final provider chunk with no choices can also be usage metadata rather than an empty answer. Treating every such trailer as malformed output creates false failures. That distinction was fixed in #975.

A database scan delayed customers who never reached a provider

The reservation code once locked an organization row and then summed a large history of its attempts to check rate and token limits. Finishing an attempt needed the same organization lock. A busy organization could therefore make reservations and settlements wait behind a slow scan.

Those waiting calls occupied the workers' limited bridge callback threads. Once those threads were full, unrelated customers waited too, before any provider was called. Adding another provider route would not fix this bottleneck.

The fix maintains counters as attempts start and finish instead of repeatedly scanning their history. Platform changes #1811 and #1816 introduced bounded second and minute buckets. Minute token checks still use at most about 61 per-second rows; hour and day token checks use bounded minute buckets. Organization RPM now uses an exact dispatch ordinal and timestamp lookup after its short compatibility warm period, rather than summing those second buckets. Outstanding input reservations, unresolved meter holds, and key daily spend have maintained totals too.

Writers take the organization lock before counter locks in the same order. Without that rule, reservation could hold the organization while waiting for a counter, as settlement held the counter while waiting for the organization. The optimization would replace slow scans with a deadlock.

The remaining token buckets have slightly wider boundaries than exact timestamps: minute token windows can include roughly one extra second and hour token windows up to 59 extra seconds. This is not the boundary of the warmed-up exact RPM check. Changes #1841 and #1842 extend the counter approach to key, identity, team, and model spending checks. This does not mean every budget check is constant-cost; the retained preverification allowance still needs separate scrutiny.

Periodic reconciliation can scan the full ledger to find counter drift, but it belongs outside request handling. It needs its own time budget, a result for each check, and monitoring for a task that stops making progress. Platform #1893 adds that supervision. A check that timed out must not be reported as having found no discrepancy.

Keep money units and cache prices correct through releases

Money now uses integer nano-USD: one dollar is 1,000,000,000 nano-dollars. The change came through engine #915 and platform #1772. Renaming a database argument cannot make an old worker send values in the new unit. If the schema changes before workers roll, old values can still be accepted as integers while pricing attempts at the wrong scale.

Reading old catalog snapshots is a separate compatibility problem. Platform #1791 preserves a published snapshot's identity instead of rejecting it because a newer reader computes a different hash. Being able to read both snapshot formats does not prove both workers interpret money correctly. Both requirements need tests during a mixed old-and-new release.

Missing cache prices can also produce a valid-looking but wrong bill. A provider may report cache reads while the catalog has no cached-input rate, causing them to be priced as full input. The repair is to supply an explicit sourced rate, or an explicit equal-to-input rate when there is no discount, then reconcile affected usage. Guessing one discount for all providers is not safe. Nor should cached input or reasoning output be added twice when already included in the provider's totals.

Keep customer-free and provider-free separate

A promotion can make a request free to the customer even when we pay the provider. That is different from a provider offering a free route. Platform #1771 makes an explicit promotional :free request stop at its allowance instead of silently spending credit. Platform #1881 retires a separate class of provider-free house routes while preserving eligible promotions on paid routes.

Linked-model accounting has to preserve these distinctions. A free request for A cannot automatically take A's subsidy to B. A zero platform charge also cannot be used as proof that a BYOK or promotional attempt cost the provider nothing.

Methodology and test status

The mechanisms above come from three kinds of work: released code, later merged fixes, and a candidate implementation. Tests and production observations answer different questions. This section records those distinctions in one place so the technical explanations do not imply more deployment or performance evidence than we have.

Versions and production observations used in this article

The September 25 source check used published engine 0.7.117 at b7ec2801 and platform f7309626c for the updated serving and accounting explanations. That is a source cutoff, not a claim that every worker runs that combination or a new production acceptance result. Candidate sections and older measurements retain their separate evidence.

The September 16 observations below were reviewed against the release records for platform commit 749f0e7e, which used engine 0.7.82 and native package 0.3.72. That is the historical measurement baseline, not a statement of today's deployed version. Engine fixes #993 and #995 merged after that baseline.

We read the production report at 07:06 UTC on September 16, 2026, looking back six hours. Request totals count settled usage records; attempt totals can also include calls still running. The reports use slightly different cutoff times.

We deployed a release during that six-hour period, so the totals don't describe only version 0.7.82. They also don't show whether a fix improved performance. We'd need a before-and-after comparison for that, and we didn't retrieve the full sequence of events for each request.

Incident examples cited from a pull request are reports from that change review. They are not automatically confirmed by the production aggregate. Similarly, a correct source fix does not prove that a later deployment contains it. The linked-model, recovery, and status sections describe the candidate implementation rather than those released APIs.

What the candidate tests establish

We tested an earlier preview through the actual native worker, using a local server to supply controlled provider responses. We checked healthy routes, child fallback, reciprocal links, the parent's remaining route, both throttle outcomes, broken streams, capability rejection, and root budget rejection. We also checked the resulting attempt and usage records.

Most failure cases in that hosted matrix used Chat. Healthy requests also covered Responses and Messages. Unit and database tests separately exercised concurrent edits, permissions, accounting, recovery clocks, incomplete reports, and cache identity. Those results belong to the exact revisions tested. A reconstructed or rebased branch does not inherit them merely because it has the same feature name.

The original nine-case real-provider suite was prepared for a frozen preview. That run was blocked by a credential-policy check and an incompatible historical test build. Later source and database checks do not turn that blocked run into a pass.

We have not recorded a completed deployed acceptance run for the current nested-fallback and recovery combination here. A direct-provider pilot alone cannot prove a nested path, cache-preserving return, or a fault case it never exercised. We make no measured cache-saving, latency, invoice, or production-activation claim for that candidate. The article's examples explain the rules; they do not replace those tests.

Test the decision, not just the HTTP status

A useful fallback test checks the exact attempt order and the reason for each move. Linked-model tests need reciprocal references, aliases, an exhausted child, each failure policy, deadline expiry, cancellation, and no replay after meaningful output. A successful A-to-B example alone does not cover those behaviors.

Recovery tests need separate workers and customers, controllable clocks, matching and mismatching regions and credentials, old and contradictory observations, changed prefixes, failed trials, and expired one-use permissions. Accounting tests need both root and destination budgets, actual destination prices, promotions, free-only rules, and finalization that does not run twice. Configuration tests need concurrent edits and credential deletion, not just clicking Save once.

Status tests must cover malformed data, empty and historical feeds, stale results, and incorrect region matches. They must also verify that changing an official status record never changes routing or recovery. That is the defining rule of a visibility-only integration.

Report each case as passed, failed, blocked, or untested. A test expecting a capability rejection before dispatch must check that rejection and zero attempts; an unrelated 500 is not a pass. A test of keyed replay must actually send the key and exercise replay. Closing an unkeyed accounting record does not prove it. Unknown usage or an unverified build should be stated as unknown, not filled with an assumed value.

Freeze the exact platform and engine combination

An engine change needs a test through an actual deployment. Claims about real providers need real scoped keys and realistic requests, not only simulated providers or another branch's successful run. Freeze the platform, engine, schema, and relevant settings before testing. Continually rebasing onto moving main changes the thing being tested and requires fresh evidence for the changed combination.

Record the platform commit and tree, engine source, Python and native package hashes, build run and artifact IDs, migration files and hashes, seed/configuration hashes, and relevant flags. A version label alone cannot distinguish a feature package from a published package with the same label. The test must check out that exact source, not an automatically generated merge with a newer main.

The CI receipt must identify the code and workflow that actually ran. A caller-provided file stating a commit hash is not proof by itself. Database verification must likewise record the files applied and the observed schema and privileges. Matching migration numbers alone does not prove identical SQL or access rules.

Compatibility tests need a genuinely supported older release. Choosing a feature ancestor that already imports unpublished APIs can fail before the application starts. That proves the chosen test build is incompatible with its dependencies, not whether the new schema works with production. Keep fixed-candidate evidence separate from checks against current main; both matter, but answer different questions.

Do not reset an existing test database just to make a run pass. Use separate models and keys for automation and human testing. Record the catalog, policy revisions, and settings before and after each phase, and fail on unexpected changes. Seeding human examples after automated failure cases avoids giving people breakers already opened by the test. It also creates a later configuration that needs its own comparison, rather than silently inheriting the earlier result.

Use simulated failures and real requests for different questions

A controlled provider can reliably return two 503s, a chosen Retry-After, an account-error body, or text followed by a severed connection. That makes it useful for testing exact order, bounded retries, and no replay after output. It cannot prove a real cache hit, provider tokenizer, account quota, latency, invoice, or cancellation behavior. A successful real request, in turn, does not exercise every rare failure condition.

The original bounded test plan used nine requests: two turns to remember and recall a fact; one schema-constrained extraction; a request and answer for one validated local inventory tool; cache priming and a repeat with the same sufficiently long prefix; a completed stream; and a cancelled stream. The tool is a fixed local function, not arbitrary code or network access. Inputs are synthetic, not customer conversations. We do not flood a provider to force its rate limit.

That original preview plan allowed twenty physical attempts and one dollar of conservative provider-cost exposure. It used one route, no model references, and no scheduled throttle retries: at most two attempts for each of nine requests, or eighteen in total. Those were limits for that plan, not recorded spending or a standing budget for later pilots. Every run must stay within the campaign's approved cumulative limits; retries and resumed runs count toward the same ceiling. Before every request, the harness checks the published route, authorized key, bounded input and output, and committed usage from the previous requests. Missing meters, unexpected attempts, changed settings, or an incorrect scenario result stop the run before another request.

Cost is budgeted without assuming cache hits. The report separates provider-reported cached input, list-price arithmetic, and a conservative cost ceiling that prices all input. Zero reported cache reads is a miss, not savings. Missing cache-write usage is unknown. A client stopping its stream does not prove the provider stopped generating or charging at that moment, even if gateway accounting has finished.

Test credentials belong to one private test organization and connection, with a short-lived gateway key. Cleanup verifies ownership and removes only the test copy; it must never revoke a production provider key. Nonproduction uses dedicated capped accounts under its deployment policy. Any exception requires explicit approval and must still pass the relevant security checks.

Understand what a staging merge changes

Shared staging follows main and its supported workflows require a reviewed commit already on main. Normal engine dependencies are exact released versions. Separately, the database integration can apply migrations to production at merge, before a manual production app rollout. Merging a schema change to test it on staging is therefore already a production-affecting action, not an isolated experiment.

A safe release tests both the old application and the new application against the new schema. Additive changes keep the old shape usable; removal happens in a later release after old workers are gone. Shared staging also needs a coordinated test window and its approved capped credentials. Testing without production writes instead requires an approved isolated deployment of the fixed combination. Permission to use a feature engine package in a preview does not also authorize it in staging or production.

Sources and further reading

The routing engine is open source at experientiallabs/experiential. We've linked the public engine changes and the release code we checked. Platform change numbers refer to our private repository, so those aren't public links. Provider cache and quota rules can change; the official references below describe their own services, not every cloud that hosts the same model.