← Back to blog

When Your Upstream Runs Out of Quota: Service Failover in Finno

If your provider's daily cap fills up, should your customer get a 429? Finno service failover spills traffic to backup upstreams — same price for the customer, correct settlement for whoever actually delivered.

Finno service failover — spill traffic to fallback upstreams when provider quota is exhausted.

You resell an upstream API. Your customer pays your price for your product name — say, “GPT-4 Turbo via Acme AI.”

Behind the scenes, that product might map to OpenAI today and to a backup provider tomorrow. That part is your operations secret. What the customer should never feel is: “Sorry, our provider ran out of quota, try again later.”

That 429 Too Many Requests is not just an HTTP code. It is lost revenue, an angry customer, and often a manual fire drill in Slack.

Finno service failover exists for that moment.

This is not the same as “the cluster node died”

People hear “failover” and think of high availability: one server dies, another takes over, nothing is lost.

Finno does that too — charges are replicated with Raft so a node failure does not orphan a successful request.

Service failover is a different problem.

Here the gateway is healthy. The cluster is healthy. The customer has balance. But the provider’s shared capacity for that service — a daily call budget, a monthly spend cap, or a concurrency limit you modeled as “service-wide throttle” — is exhausted.

Without failover you have two bad options:

  1. Return 429 and stop serving (revenue stops).
  2. Manually point traffic somewhere else and hope finance can reconcile who delivered what (settlement breaks).

Service failover is the third option: keep serving, bill the customer normally, pay the provider who actually answered.

A simple example

Imagine you sell “Image Generation API” at $0.05 per call.

  • Primary upstream: Provider A, who gives you 10,000 calls per day on a shared key.
  • Fallback upstream: Provider B, configured as a backend-only service in Finno — customers never see it in the catalog, but the gateway knows how to call it.

At 3 p.m. on a busy day, Provider A’s shared budget hits zero. From that point:

  1. The next customer request still hits Image Generation API (the facing service they bought).
  2. Finno sees that the service-wide budget is exhausted.
  3. Instead of rejecting the call, it tries Provider B from your ordered failover list.
  4. The customer gets their image. They are charged $0.05 — the price of Image Generation API, not Provider B’s internal name.
  5. In settlement reports, Provider B gets the wholesale cost for that call, because that is who actually served it.

The customer experience is continuous. Your books stay honest.

How you configure it (in plain terms)

In the admin dashboard, open a consumer-facing service and add failover targets — an ordered list of other services that can stand in when the primary is out of shared capacity.

Each target can carry its own provider cost package, so margin math stays correct even when the backup upstream is cheaper or more expensive.

You can also mark a service as backend-only: it never appears in the consumer catalog, but it can appear in failover chains. That is useful when you want a hidden row whose only job is “spill traffic here when the main upstream is full.”

No new billing model. No special-case Raft command. The same Reserve → Finalize → Ledger flow runs; only the upstream route changes.

What failover will not do (on purpose)

Failover protects provider-side shared limits. It does not let a customer escape their own limits.

If you gave Consumer X a maximum of 500 calls per day on this service, they still get 429 when they hit 500 — even if the failover target has plenty of room. Calls served through failover count toward that consumer cap.

Same for subscription quota: if a quota plan is exhausted and your overage policy says “block,” Finno applies that policy first. Failover is not a back door around a customer’s contract.

That distinction matters for fairness and for fraud: failover is for upstream capacity, not for “free extra calls.”

How this pairs with Survival Mode

Service failover answers: “Our provider’s quota is full — can we still serve?”

Survival Mode answers: “Our machine is under heavy CPU or memory pressure — what do we shed so the host survives?”

They solve different crises. Many production incidents involve both: upstream gets slow, concurrency piles up, the host heats up. Finno can spill to a backup upstream and shed low-value traffic under Survival Mode — without mixing the two concepts in your runbook.

If you want the full story on host-level protection, read Survival Mode. For a product overview of resilience features, see Operational resilience on the features page.

When to turn it on

Service failover is most valuable when you:

  • Resell upstream APIs with hard or soft daily/monthly caps on a shared provider key.
  • Run multiple upstream accounts or providers for the same product tier.
  • Need continuity more than a perfect 1:1 upstream mapping — as long as settlement reflects who delivered.

If you only ever have one upstream with unlimited capacity, you may never configure a failover list. But most resellers eventually hit a provider limit on the worst day of the quarter. That is the day this feature pays for itself.


Questions about failover chains, backend-only services, or settlement with mixed providers? Contact us — we are happy to walk through your traffic shape on a call.