Designing provider-agnostic LLM applications
A provider interface should isolate meaningful differences in streaming, tools, errors, and capabilities instead of hiding them behind a lowest-common-denominator request.

Supporting more than one model provider sounds like a small adapter exercise: convert a request, call an endpoint, convert the response. The simple version works for a text prompt. Production applications quickly depend on streaming, tool calls, structured output, cancellation, usage reporting, and provider-specific error behavior.
The design question is not whether providers differ. It is which differences should be normalized, which should remain visible, and how a product can behave honestly when a provider lacks a capability.
Define the boundary around application needs
Start from the workflow. A research product may need structured claims and citations; a coding agent may need streaming tool calls and interruption. Model that application-level request in a provider-neutral contract:
type ModelRequest = {
messages: Message[];
tools?: ToolDefinition[];
responseFormat?: ResponseFormat;
signal?: AbortSignal;
};
type ModelEvent =
| { type: "text_delta"; text: string }
| { type: "tool_call_delta"; index: number; patch: unknown }
| { type: "usage"; inputTokens?: number; outputTokens?: number }
| { type: "completed"; finishReason: FinishReason };This gives the rest of the product a stable conversation shape while leaving each adapter responsible for mapping its provider's event stream and tool-call representation.
Make capability differences explicit
One model may support constrained JSON output; another may support a different tool schema or no streaming tool events. A provider adapter should report its capabilities so the product can choose a supported path, disable a feature with an explanation, or use a deliberate fallback.
Silently dropping a response constraint or translating an unsupported tool call into plain text can create a success-shaped failure. Capability negotiation turns that into an explicit product decision.
Normalize behavior, not just field names
Adapters should normalize the events and errors that the application depends on. They should not erase distinctions the application needs to make. A rate limit, invalid credential, policy refusal, malformed response, and network timeout have different recovery options.
Streaming deserves careful treatment. Tool arguments can arrive in fragments, completion events can have different semantics, and cancellation must stop work in both the application and the provider request. Each adapter should translate these cases into a consistent event lifecycle and preserve provider details for safe diagnostics.
Keep credentials and routing out of the UI
Browser code should not receive server provider secrets. The server selects an entitled provider and owns its credentials, while a local-first application may keep credentials in a host process or user environment. In either case, API keys should not be stored in a shared workspace file or exposed in logs.
Provider selection should be a policy decision with a clear owner. The model can produce output; it should not choose a secret-bearing endpoint or alter account entitlements.
Test the contract at the seams
Contract tests can exercise each adapter against recorded or mocked provider events: text streaming, tool calls, malformed output, timeout, cancellation, authentication failure, and usage metadata. A shared suite checks the interface. Provider-specific tests cover the details that cannot be normalized.
Live tests are still useful, but they should be separate from deterministic adapter tests. They require credentials and can be rate-limited or changed by the provider. A passing mock proves the adapter contract; it does not prove every remote service is reachable today.
Avoid the lowest-common-denominator trap
Provider neutrality does not mean every model can do every task. Keep model and provider capabilities available to the product so it can explain why a feature is unavailable or choose another configured model. If users care about portability, show the operational differences rather than promising that every request behaves identically.
Practical checks
- Define request and event contracts from real application flows.
- Normalize streaming lifecycle, cancellation, usage, and recoverable errors.
- Expose capability metadata instead of silently dropping unsupported features.
- Keep provider credentials in the backend or trusted host process.
- Separate stable contract tests from credentialed live checks.
- Preserve enough provider detail for privacy-safe diagnostics.
The value of a provider boundary is that the product can evolve without rewriting every workflow. That boundary stays useful only when it represents the differences the workflow actually depends on.