Agent orchestration
ForgeLoop: Making agent-delivered code reviewable
A distributed software-delivery system that coordinates bounded agent work and ties completion to verification evidence and human review.

Overview
ForgeLoop takes a GitHub issue or feature specification through planning, bounded task execution, integration, verification, and review. Its design premise is simple: an agent saying it finished is not evidence that the work satisfies its requirements.
The project has a hosted control plane and a separate runner that works near the repository. It is source-available under its own license; it is not an OSI-approved open-source project.
The problem
Coding agents can write patches. Reliable delivery also has to coordinate dependencies, keep parallel changes from colliding, control repository access and model spend, recover from failures, and show what was checked before a person approves publication.
ForgeLoop explores these as product and system boundaries rather than treating them as details hidden behind a chat panel.
Architecture
- 01Web consoleRepositories, runs, policies, review
- 02Spring Boot APIGraphQL, scheduling, audit, leases
- 03Java runnerWorktrees, Docker, MCP, model calls
- 04VerificationChecks, acceptance evidence, review
The React and TypeScript operator console connects to a Java 21 and Spring Boot GraphQL control plane. PostgreSQL stores workflow state and Flyway manages schema changes; artifacts use filesystem or S3-compatible storage. A separate Java runner pairs with the control plane, creates Git worktrees, coordinates Docker-based execution and verification, and uses runner-local provider credentials. The MCP tool gateway is implemented in TypeScript.
Repository context is sent from the runner to the selected model provider. The control plane stores workflow records, findings, usage, and evidence metadata rather than acting as a proxy for repository context.
Engineering decisions
Split orchestration from execution
The control plane schedules tasks and enforces organization-scoped policy. The runner checks out and operates on repository code near the repository. This keeps the hosted service out of the normal execution path for arbitrary project commands and leaves provider credentials with the runner.
Bound parallel work and make integration deterministic
Task graphs validate dependencies, capabilities, and owned paths. Runners claim leased work and use isolated Git worktrees. Changes enter integration through declared commits rather than agents mutating a shared checkout.
Require evidence before a completion decision
Structured outputs are checked against schemas and repository paths. Versioned verification gates produce criterion-level evidence; bounded repair can address failed checks, and a review stage examines the outcome. GitHub publication remains behind a human approval boundary, with auto-merge disabled unless an administrator opts in.
Technical challenges and solution
Distributed workers can disappear, duplicate requests can arrive, checks can fail mid-run, and provider responses can be malformed. ForgeLoop persists workflow transitions, uses runner heartbeats and expiring leases for recovery, records bounded and redacted evidence artifacts, and tracks usage and retry limits. The operator console exposes progress and failure details so a person can decide what to do next.
These mechanisms make the work observable; they do not imply every operational launch gate is closed. The repository documents a demonstrated issue-to-verified-PR path against a demo repository, while validation across an unrelated second repository and several deployment readiness checks remain open.
What I built
I built the product across the web console, control plane, runner, MCP gateway, workflow contracts, verification and review loop, and deployment infrastructure. The implementation makes policy, recovery, and evidence part of the product instead of a claim attached to a generated patch.
Engineering takeaways
- A verification result needs a clear relationship to the criterion and commit it checked.
- Parallel agents need task ownership and workspace isolation before they need more autonomy.
- Recovery depends on persisted transitions and explicit leases, not only retries.
- A human approval boundary is meaningful only when the reviewer can see the evidence behind it.
Stack
Java 21, Spring Boot, Spring Security, GraphQL, Spring Data JPA, Flyway, PostgreSQL, React, TypeScript, Vite, Playwright, Docker, Git worktrees, Model Context Protocol, Zod, GitHub Apps, GitHub Actions, and S3-compatible artifact storage.
Stack
- Java
- Spring Boot
- GraphQL
- React
- PostgreSQL
- MCP