Skip to content

About

An agent that writes, tests, and adopts its own tools at runtime, when it hits a task outside its existing toolkit, it synthesizes a new tool, verifies it against self-generated test cases before trusting it, and runs it in a sandboxed environment.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agentic Tool Synth

A self-evolving agent architecture that allows an LLM agent to synthesise, verify, and retain its own tools at runtime.

Most LLM agents today work off a fixed toolset provided by a developer. If the agent encounters a task that requires a capability it doesn't have, it fails or hallucinates. This project solves that bottleneck by giving the agent the ability to write its own tools, test them in a sandboxed environment, and keep the ones that are broadly useful.

Key Features

  • Verify-Before-Trust Pipeline: The core research insight. The agent writes pytest-style test cases based only on the task description before it writes the tool code. This prevents the LLM from rubber-stamping its own bad code.
  • Sandboxed Execution: Synthesised tools and their tests are executed in an isolated subprocess with a hard timeout, preventing malicious code or infinite loops from crashing the main agent process.
  • Autonomous Retention Policy: Tools are evaluated based on usage count, cross-session reuse, description generality, and verification status. They are placed into tiers (PERMANENT, SESSION, or DISCARD). Useful tools are kept, one-off hacks are discarded.
  • Cross-Session Persistence: Tools that pass the retention policy are automatically saved to toolkit.json and inherited by future agent sessions. No human copying required.
  • Rich Terminal Trace: Full, live-updating trace of the agent's thought process, test execution, matching, and retention evaluations, rendered beautifully in the terminal using the rich library.

Project Structure

c:\dev\agentic tool synth\
├── main.py                 # Main entry point (loads toolkit, runs tasks, prunes)
├── demo_tasks.py           # A two-session demo proving cross-session persistence
├── .env                    # Contains your ANTHROPIC_API_KEY
├── agent/
│   ├── agent.py            # The core agent loop (match -> test -> synth -> sandbox -> retain -> execute)
│   ├── console.py          # Centralised rich terminal output formatting
│   ├── models.py           # Pydantic schemas (Task, Tool)
│   ├── retention.py        # Logic for evaluating tool generalisability (PERMANENT / SESSION / DISCARD)
│   ├── sandbox.py          # Subprocess isolation for executing synthesised code safely
│   ├── synth.py            # The tool generation module (calls Claude)
│   ├── toolkit.py          # Registry for tool storage, JSON persistence, and keyword matching
│   └── verifier.py         # Test case generation module (calls Claude)

How to Run

  1. Install dependencies: Ensure you have your Python environment set up with anthropic, pydantic, pytest, rich, and python-dotenv.

  2. Setup environment: Create a .env file in the root directory and add your Anthropic API key:

    ANTHROPIC_API_KEY=your_api_key_here
  3. Run the Standard Demo:

    python main.py

    To wipe the inherited toolkit and start completely fresh:

    python main.py --reset
  4. Run the Two-Session Persistence Demo: This explicitly runs two sessions back-to-back. The first session builds tools, and the second session proves that the agent can inherit and reuse those tools for similar tasks without rewriting them.

    python demo_tasks.py

Switching to the Real Claude API

Currently, the agent is running in MOCK_MODE. This means it uses hardcoded tools and tests to demonstrate the architecture without burning API credits.

Once you have added credits to your Anthropic API key, you can enable true autonomous synthesis by disabling MOCK_MODE.

Change MOCK_MODE = True to MOCK_MODE = False in the following files:

  1. agent/synth.py (Line 21)
    MOCK_MODE = False
  2. agent/verifier.py (Line 18)
    MOCK_MODE = False
  3. agent/retention.py (Line 48)
    MOCK_MODE = False

Once these are set to False, the agent will call the claude-sonnet-4-6 model to genuinely write code and test cases on the fly!

How to check the output?

1. The Two-Session Accumulation Demo

Run this command to see the agent start fresh, synthesize tools in Session 1, and then automatically reuse them in Session 2:

python demo_tasks.py

What you will see:

  • Session 1: The agent encounters new tasks, goes through [SYNTH], tests the tools in the [SANDBOX], and marks them [VERIFIED].
  • Session 2: The agent starts again, but you'll see a banner showing it inherited the tools from Session 1. It will skip synthesis and print a blue [MATCH] banner showing it successfully reused the tools for different tasks.

2. The Standard Run (With Reset)

If you just want to run the main agent loop from a blank slate, you can use the --reset flag:

python main.py --reset

What you will see:

  • A yellow warning that the toolkit.json was wiped.
  • The agent building tools from scratch, testing them, and showing the Retention Report at the end (the table showing which tools are permanently kept based on their score).

If you run python main.py without --reset right after, you'll see it instantly reuse the tools it just built!

About

An agent that writes, tests, and adopts its own tools at runtime, when it hits a task outside its existing toolkit, it synthesizes a new tool, verifies it against self-generated test cases before trusting it, and runs it in a sandboxed environment.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages