A local MCP recovery layer for AI agents that get stuck, retry blindly, or lose control of long-running work.
coffee_report_event . coffee_recommend_recovery . coffee_take_break . coffee_recovery_review
AI agents do not only fail because they lack reasoning ability. They often fail because their control flow degrades:
- the same command is rerun without changing the hypothesis
- the same tool error is retried until the context fills up
- rate limits trigger a noisy retry loop
- a long task continues without a checkpoint
- the agent should ask for help, but keeps acting
AgentNeedCoffee gives the agent a local MCP server that records runtime events, detects loops, and returns a concrete recovery action.
| Detect | Repeated failures, retries, rate limits, and context pressure. |
| Decide | Map run state to actions like backoff, compact_context, or ask_human. |
| Recover | Create checkpoints so the agent resumes from evidence, not habit. |
AgentNeedCoffee is now two pieces:
| Piece | Role | Lives in |
|---|---|---|
| MCP server | Stores state and returns recovery actions | src/agent_need_coffee/mcp_server.py |
| Agent skill | Teaches the agent when and how to call the MCP tools | skills/agent-need-coffee/ |
The MCP server is the runtime. The skill is the operating procedure.
AI agent / MCP client
|
| stdio MCP
v
agent-need-coffee-mcp
|
| record event -> evaluate policy -> return recovery action
v
SQLite state store
AgentNeedCoffee is local-first. It does not need a hosted API, account system, open port, or cloud database.
Default state path:
~/.agent_need_coffee/coffee.sqlite3
Override it:
export AGENT_NEED_COFFEE_DB=/path/to/coffee.sqlite31. Agent runs work
2. Something meaningful happens
3. Agent calls coffee_report_event
4. Server persists the event
5. PolicyEngine calculates run health
6. Server returns a recovery action
7. Agent changes its next move
If the agent calls coffee_take_break, the server records a checkpoint and
resets active loop analysis for future events. Historical events remain in the
timeline.
Requires Python 3.10 or newer. Python 3.9 cannot install the current MCP SDK.
pip install agent-need-coffeeFor local development:
python -m pip install --upgrade pip
pip install -e ".[dev]"Use the installed command:
agent-need-coffee-mcpGeneric MCP configuration:
{
"mcpServers": {
"agent-need-coffee": {
"command": "agent-need-coffee-mcp"
}
}
}Development checkout configuration:
{
"mcpServers": {
"agent-need-coffee": {
"command": "python",
"args": ["-m", "agent_need_coffee.mcp_server"],
"env": {
"PYTHONPATH": "/absolute/path/to/AgentNeedCoffee/src"
}
}
}
}The repository includes a Codex/Agent skill that tells the agent when to call AgentNeedCoffee.
cp -R skills/agent-need-coffee "${CODEX_HOME:-$HOME/.codex}/skills/"Once installed, the skill triggers when the agent repeats failures, accumulates retries, hits rate limits, faces context pressure, or needs a checkpoint. It then instructs the agent to call the MCP tools and follow the returned action.
| Tool | Use it when | What it returns or changes |
|---|---|---|
coffee_report_event |
A command, test, model call, or tool call succeeds or fails. | Stores the event and returns updated state plus recovery advice. |
coffee_get_state |
The agent needs current run health. | Returns scores, counters, status, and dominant failure signature. |
coffee_recommend_recovery |
No new event happened, but the agent needs a decision. | Returns a structured recovery action. |
coffee_take_break |
The agent needs to checkpoint before continuing. | Records a break and resets active loop analysis. |
coffee_clear_run |
The run is intentionally starting over. | Deletes local state for one agent/run pair. |
{
"agent_id": "codex",
"run_id": "fix-tests-42",
"event_type": "tool_error",
"success": false,
"error_kind": "test_failure",
"tool_name": "pytest",
"command": "pytest -q",
"message": "same assertion failed",
"signature": "pytest-same-assertion"
}Example recovery after repeated failures:
{
"state": {
"status": "stuck",
"failure_events": 3,
"repeated_failure_count": 3,
"dominant_failure_signature": "pytest-same-assertion"
},
"recovery": {
"severity": "high",
"action": "create_checkpoint",
"reason": "The agent appears stuck in a failure loop.",
"suggested_next_step": "Create a compact checkpoint, then choose one smaller verification step before continuing."
}
}| Action | Meaning |
|---|---|
continue |
The run is steady. Proceed normally. |
continue_with_smaller_step |
Narrow the next action to one hypothesis. |
backoff |
Pause retries and reduce frequency or request size. |
compact_context |
Summarize the run before continuing. |
create_checkpoint |
Stop broad execution and write a recovery checkpoint. |
ask_human |
Do not keep retrying; ask for direction. |
Resources:
coffee://agents/{agent_id}/state
coffee://runs/{run_id}/timeline
Prompt:
coffee_recovery_review
The prompt asks the agent to write a compact checkpoint containing the current goal, evidence, failed attempts, dominant failure signature, and exactly one next action.
The policy engine computes:
fatigue_score: token volume, retry count, and run durationfriction_score: failure volume, repeated failures, and rate limitsloop_score: how strongly one failure signature is repeating
It then maps those signals to recovery actions:
steady -> continue
minor failures -> continue_with_smaller_step
retries/rate limits -> backoff
context pressure -> compact_context
repeated failure -> create_checkpoint
persistent loop -> ask_human
python -m pip install --upgrade pip
pip install -e ".[dev]"
pytest -qRun the MCP server directly:
AGENT_NEED_COFFEE_DB=/tmp/coffee.sqlite3 agent-need-coffee-mcpThe implementation uses the official Python MCP SDK with FastMCP and the
default stdio transport.