Building reliable AI agents with MCP and tool calling
Reliable agent workflows depend on explicit tool contracts, bounded execution, validated results, and traceable state around the model call.

An agent that can call tools has crossed an important boundary. It can now read from, or act on, systems outside the language model. That capability makes the surrounding engineering more important than the prompt that names the tool.
Model Context Protocol helps make tool servers discoverable through a common interface. It does not make every tool safe, every result trustworthy, or every sequence correct. Those properties belong to the application that selects and runs tools.
Treat tool descriptions as contracts
A tool needs more than a name and a sentence. Its input schema should say which fields are required, what values are allowed, and which assumptions the caller must satisfy. Its result should have a stable shape the application can validate.
The application should also attach policy metadata outside the model-visible description. A read-only search and a write operation may both be valid tools, but they do not deserve the same automatic permission. Keep capability, scope, timeout, and approval requirements in a policy layer the model cannot change by rewriting its arguments.
Separate selection from execution
When a model proposes a tool call, the application should validate it before execution:
type ToolRequest = {
name: string;
arguments: unknown;
callId: string;
};
async function dispatch(request: ToolRequest, policy: Policy) {
const tool = registry.get(request.name);
if (!tool) return rejected(request.callId, "unknown_tool");
const input = tool.inputSchema.parse(request.arguments);
const decision = policy.authorize(tool.capability, input);
if (!decision.allowed) return needsApproval(request.callId, decision.reason);
return runWithDeadline(tool, input, decision.deadlineMs);
}Schema validation answers whether the arguments are shaped correctly. Policy answers whether this user, session, and task may perform the action. Time limits and cancellation answer whether it can keep consuming resources indefinitely. Each check solves a different problem.
Treat tool results as untrusted input
Tool output often contains text copied from a page, a repository, a document, or another user's system. That content may be irrelevant, malformed, oversized, or deliberately written to redirect an agent. Validate the response shape, bound its size, label its origin, and keep it separate from trusted instructions.
The agent should receive only the portion of the result needed for the next step. Raw tool output can be retained for a trace or audit, but passing an entire response back to a model by default expands both token cost and the injection surface.
Bound the workflow
An agent loop needs explicit limits: maximum tool calls, iterations, wall-clock time, token or cost budget, and retry count. A timeout should be a normal result state, not an exception that leaves the session half-open. Retries should distinguish transient transport failures from invalid input or denied policy.
For mutations, idempotency keys and operation records help prevent a retried request from applying the same change twice. For long-running work, persist enough state to resume or explain the interruption.
Make the trace useful to a person
A useful execution trace answers: what tool was selected, what policy applied, what was sent, what came back, what failed, and what the model did next? Redact secrets and private content from logs, but keep stable identifiers that let the user follow a claim or action back to its source.
The trace should describe the application’s behavior, not just serialize a chat transcript. That distinction makes debugging and review possible even when the model's text is incomplete.
Put approval at the action boundary
Human approval is most useful before an action with external or destructive effects. The reviewer needs to see the operation, target, relevant input, and expected effect. A generic “allow tool” prompt offers little control if the application hides which resource will change.
Read-only actions can often use a broader policy, but read-only is not automatically risk-free. A search can expose private data to a provider, and a browser fetch can reach an untrusted site. Scope and data handling still matter.
A practical reliability checklist
- Version input and output schemas.
- Keep authorization separate from model-selected arguments.
- Bound size, time, retries, and cost.
- Label and validate external results before model reuse.
- Persist workflow transitions for recovery.
- Keep an auditable trace with secrets redacted.
- Ask a person before consequential mutations.
The language model can help choose a next step. The product is responsible for making that step safe, bounded, observable, and recoverable.