Treat credit-based AI billing as a state machine

A usage ledger, idempotent payment events, and explicit charging rules keep paid AI workflows consistent when requests fail or webhooks retry.

billingStripeSaaSidempotency
Treat credit-based AI billing as a state machine illustration

When a user pays for an AI feature, the payment and the later model request happen at different times. The payment provider confirms money movement; your application decides what that purchase lets the user do. A model request can then time out, return an invalid result, or finish after the browser has disconnected.

Treating all of this as a single balance number makes the product hard to recover and harder to explain. A better starting point is to model billing as a sequence of identifiable state changes.

Separate payment from entitlement

A successful checkout is not the same thing as an application credit. The provider owns payment status; your backend owns the entitlement that follows from it. Keep that boundary on the server and connect the two with stable provider identifiers, such as an invoice or checkout session ID.

In Rezzie, paid usage spans credit grants and renewals, while Stripe webhook state is authoritative for those changes. That gives the application a defined place to translate a confirmed billing event into product access.

Make webhook handling safe to repeat

Webhook delivery is asynchronous. Stripe can retry a failed delivery, and it does not guarantee events arrive in the order they were created. Verify the signature, store the event ID under a uniqueness constraint, and apply the corresponding entitlement change in the same database transaction. A repeated event should return success without granting the same credits again.

There are two different idempotency boundaries here. A Stripe Idempotency-Key protects a request your server sends to Stripe from creating a second object after a retry. Event-ID deduplication protects your application from applying a delivered webhook twice. See Stripe's webhook delivery guidance and idempotent request docs.

Record usage as events

An append-only usage ledger can represent grants, consumption, refunds, and corrections as separate entries. Each entry should identify the user, reason, amount, and related operation. The current balance can be cached for fast reads, but the ledger gives support and recovery code a history it can explain.

Generation workflows may also need a reservation before work starts. If you use reservations, give them explicit states and an expiry path: a completed request consumes the hold, while a failed or abandoned request releases it once. For short operations, a direct debit after successful validation may be simpler. The right rule depends on when the product considers usage to have happened; write that rule down before implementing retries.

Keep a failed request from charging twice

Give each user-initiated generation operation an application-level ID. If the client retries because it lost the response, look up that operation before starting another model call or changing the ledger. This is separate from the provider request ID: one user operation may involve more than one provider call, and a single provider call can succeed even when the browser never receives its response.

Persist enough state to distinguish started, completed, and failed work. If the result is invalid, do not mark the request complete just because the provider returned HTTP 200. Keep the entitlement decision, validation result, and ledger write tied to the same operation.

Reconcile what the product believes

Idempotency protects normal retries, but it does not make billing infallible. Keep the provider event ID and relevant customer or invoice IDs in operational records, expose a support-friendly history, and provide a reconciliation path for missed events or manual corrections. Never repair a balance with an unexplained counter edit.

For a credit-based AI product, a useful baseline is simple: the provider confirms money, the backend grants access once, and every unit of usage can be traced to a user operation. That makes failure handling part of the billing design instead of a support ticket waiting to happen.