A self-evolving agent architecture that allows an LLM agent to synthesise, verify, and retain its own tools at runtime.
Most LLM agents today work off a fixed toolset provided by a developer. If the agent encounters a task that requires a capability it doesn't have, it fails or hallucinates. This project solves that bottleneck by giving the agent the ability to write its own tools, test them in a sandboxed environment, and keep the ones that are broadly useful.
- Verify-Before-Trust Pipeline: The core research insight. The agent writes
pytest-style test cases based only on the task description before it writes the tool code. This prevents the LLM from rubber-stamping its own bad code. - Sandboxed Execution: Synthesised tools and their tests are executed in an isolated
subprocesswith a hard timeout, preventing malicious code or infinite loops from crashing the main agent process. - Autonomous Retention Policy: Tools are evaluated based on usage count, cross-session reuse, description generality, and verification status. They are placed into tiers (
PERMANENT,SESSION, orDISCARD). Useful tools are kept, one-off hacks are discarded. - Cross-Session Persistence: Tools that pass the retention policy are automatically saved to
toolkit.jsonand inherited by future agent sessions. No human copying required. - Rich Terminal Trace: Full, live-updating trace of the agent's thought process, test execution, matching, and retention evaluations, rendered beautifully in the terminal using the
richlibrary.
c:\dev\agentic tool synth\
├── main.py # Main entry point (loads toolkit, runs tasks, prunes)
├── demo_tasks.py # A two-session demo proving cross-session persistence
├── .env # Contains your ANTHROPIC_API_KEY
├── agent/
│ ├── agent.py # The core agent loop (match -> test -> synth -> sandbox -> retain -> execute)
│ ├── console.py # Centralised rich terminal output formatting
│ ├── models.py # Pydantic schemas (Task, Tool)
│ ├── retention.py # Logic for evaluating tool generalisability (PERMANENT / SESSION / DISCARD)
│ ├── sandbox.py # Subprocess isolation for executing synthesised code safely
│ ├── synth.py # The tool generation module (calls Claude)
│ ├── toolkit.py # Registry for tool storage, JSON persistence, and keyword matching
│ └── verifier.py # Test case generation module (calls Claude)
-
Install dependencies: Ensure you have your Python environment set up with
anthropic,pydantic,pytest,rich, andpython-dotenv. -
Setup environment: Create a
.envfile in the root directory and add your Anthropic API key:ANTHROPIC_API_KEY=your_api_key_here
-
Run the Standard Demo:
python main.py
To wipe the inherited toolkit and start completely fresh:
python main.py --reset
-
Run the Two-Session Persistence Demo: This explicitly runs two sessions back-to-back. The first session builds tools, and the second session proves that the agent can inherit and reuse those tools for similar tasks without rewriting them.
python demo_tasks.py
Currently, the agent is running in MOCK_MODE. This means it uses hardcoded tools and tests to demonstrate the architecture without burning API credits.
Once you have added credits to your Anthropic API key, you can enable true autonomous synthesis by disabling MOCK_MODE.
Change MOCK_MODE = True to MOCK_MODE = False in the following files:
agent/synth.py(Line 21)MOCK_MODE = False
agent/verifier.py(Line 18)MOCK_MODE = False
agent/retention.py(Line 48)MOCK_MODE = False
Once these are set to False, the agent will call the claude-sonnet-4-6 model to genuinely write code and test cases on the fly!
Run this command to see the agent start fresh, synthesize tools in Session 1, and then automatically reuse them in Session 2:
python demo_tasks.pyWhat you will see:
- Session 1: The agent encounters new tasks, goes through
[SYNTH], tests the tools in the[SANDBOX], and marks them[VERIFIED]. - Session 2: The agent starts again, but you'll see a banner showing it inherited the tools from Session 1. It will skip synthesis and print a blue
[MATCH]banner showing it successfully reused the tools for different tasks.
If you just want to run the main agent loop from a blank slate, you can use the --reset flag:
python main.py --resetWhat you will see:
- A yellow warning that the
toolkit.jsonwas wiped. - The agent building tools from scratch, testing them, and showing the Retention Report at the end (the table showing which tools are permanently kept based on their score).
If you run python main.py without --reset right after, you'll see it instantly reuse the tools it just built!