A sandboxed AI coding agent runtime that autonomously modifies codebases through LLM-driven reasoning and isolated tool execution.
Agent Forge implements the ReAct (Reasoning + Acting) pattern: an agent receives a coding task, iteratively reasons about what to do via an LLM, invokes tools inside ephemeral Docker containers, and loops until the task is complete.
graph TD
CLI["CLI<br/>(Click commands)"]
CLI -->|"agent-forge run"| Core
subgraph Core["Agent Core"]
Loop["ReAct Loop<br/>Observe → Reason → Act"]
State["State Machine<br/>(PENDING → RUNNING → ...)"]
Persist["Persistence<br/>(save/load runs)"]
end
Loop -->|"prompt + tools"| LLM
LLM -->|"function calls"| Loop
Loop -->|"execute tool"| Sandbox
Sandbox -->|"result"| Loop
subgraph LLM["LLM Provider"]
Gemini["Gemini 3.1<br/>(primary)"]
end
subgraph Sandbox["Docker Sandbox"]
Docker["Ephemeral Container<br/>(per-run, isolated)"]
end
subgraph Tools["Tool Registry"]
read_file
write_file
edit_file
list_directory
run_shell
search_codebase
end
Sandbox ---|"workspace<br/>bind-mount"| Tools
- 🔒 Sandboxed Execution — Every tool invocation runs in an ephemeral Docker container (or lightweight bubblewrap namespace) with resource limits — never on the host.
- 🧠 Gemini 3.1 Ready — Full support for thought signatures, exponential backoff with jitter, and
Retry-Afterheader. - 🔌 Multi-Provider LLM — Gemini (primary), OpenAI, and Anthropic adapters, all wired through a pluggable factory.
- 🎭 Agent Profiles — YAML-based personas with configurable prompt scope, LLM overrides, and iteration limits.
- 🌐 Hosted Service Mode — Run Agent Forge as a versioned FastAPI service for external clients with API-key auth, policy controls, and Proof-of-Audit compatibility.
- 🧩 Extension SDK — Scaffold, build, and install domain-specific extensions as separate packages with
agent-forge init-extension. Auto-discovers profiles, tools, and metadata via Python entry points. - 📊 Observability — Structured JSON logs, trace IDs, token/cost tracking on every run.
- 💾 Run Persistence — Every agent run is saved to disk with full conversation history and tool invocations.
- 🔧 Tool Plugins — Add tools via the
ToolABC and ship them as entry-point packages that load automatically.
✅ Verified end-to-end with
gemini-3.1-flash-lite-previewon 2026-03-07 — 17/17 E2E tests passing.
- Python 3.11+
- Docker
- A Gemini API key (or OpenAI/Anthropic)
# Clone the repository
git clone https://github.com/akoita/agent-forge.git
cd agent-forge
# Install in development mode
pip install -e ".[dev]"
# Build the sandbox Docker image
make build-sandbox
# Set your API key
export GEMINI_API_KEY="your-key-here"# Run an agent task (direct mode — default)
agent-forge run \
--task "Add input validation to the /api/users endpoint" \
--repo ./path/to/your/repo
# Run via queue → worker pipeline (in-memory)
agent-forge run \
--task "Fix login bug" \
--repo ./my-app \
--queue memory
# Run via Redis queue (requires Redis)
agent-forge run \
--task "Refactor auth module" \
--repo ./my-app \
--queue redis \
--redis-url redis://localhost:6379/0 \
--max-concurrent-runs 4
# Check run status
agent-forge status <run-id>
# Emit a machine-readable run result and persist it to disk
agent-forge run \
--task "Generate a structured audit report" \
--repo ./my-app \
--output-format json \
--report-file ./artifacts/run-result.json
# Render an existing run as JSON
agent-forge status <run-id> --output-format json
# List recent runs
agent-forge list
# View resolved configuration
agent-forge config
# Run the hosted service
agent-forge serve --host 127.0.0.1 --port 8000Run the demo locally
# Watch the simulated demo
bash scripts/demo.sh
# Record a new asciinema cast
bash scripts/record-demo.shAgent Forge uses a layered configuration system (CLI flags > env vars > project config > user config > defaults).
Create an agent-forge.toml in your project root:
[agent]
max_iterations = 25
default_provider = "gemini"
default_model = "gemini-3.1-flash-lite-preview"
[sandbox]
memory_limit = "512m"
timeout_seconds = 300
network_enabled = falseSee the Configuration Guide for full reference. For hosted deployments, auth, and operations, see the Hosted Service Guide, including the extension-aware image build, multi-instance compose workflow, and GitHub Actions dev deploy pipeline.
# Install dev dependencies
pip install -e ".[dev,redis]"
# Run unit tests
make test-unit
# Run all tests (requires Docker)
make test
# Run e2e tests (requires GEMINI_API_KEY + Docker)
make test-e2e
# Lint & format
make lint
make formatagent_forge/
├── agent/ # ReAct loop, state machine, prompts, persistence
├── extensions/ # Extension SDK (discovery, scaffolding, templates)
├── llm/ # LLM provider adapters (Gemini, OpenAI, Anthropic) + factory
├── profiles/ # Agent profile system (YAML personas + loader)
├── tools/ # Built-in tools + entry-point plugin loader
├── sandbox/ # Sandbox backends (Docker, bubblewrap) + factory
├── service/ # Hosted FastAPI service (API, auth, client harness)
├── orchestration/ # Task queue, event bus, workers
├── observability/ # Structured logging, tracing, cost tracking
├── cli.py # Click-based CLI entry point
└── config.py # Layered configuration system
plugins/
└── proof_of_audit/ # First-party domain extension (audit tools, profiles)
- Architecture — System design, layer responsibilities, ReAct loop sequence.
- Configuration — Full config reference (TOML, env vars, CLI flags, precedence).
- Hosted Service — Hosted architecture, trust boundaries, local-dev, and operations.
- Testing — Running tests, writing new ones, CI workflows, coverage.
- Extending — Adding tools, LLM providers, Extension SDK, custom sandbox configs.
- Technical Spec — Full specification with interface contracts and data models.
Agent Forge is being reset to a harness-first strategy: a small, measurable, high-correctness coding harness that competes on correctness and verified task completion rather than feature breadth. The rationale is recorded in ADR-002: Harness-First Strategic Reset.
The work is now organized into milestones M0–M6 — starting with an evaluation
harness and quality gate (M0), then the agent-computer interface, context engine,
provider API, policy engine, durable execution, and standards interop. Hosted,
multi-tenant, and scaling concerns (web dashboard, RBAC, Kubernetes, fleet
scaling) move to a separate agent-forge-cloud project.
See the Roadmap for milestones, exit criteria, and tracking issues. The former Phase 1–7 plan is superseded.
Contributions are welcome! Please read our Contributing Guide and Code of Conduct before submitting a pull request.
If you discover a security vulnerability, please follow our Security Policy for responsible disclosure.
This project is licensed under the MIT License — see the LICENSE file for details.