Skip to content

Latest commit

 

History

History
149 lines (102 loc) · 13.3 KB

File metadata and controls

149 lines (102 loc) · 13.3 KB

Collaboration Learnings

This document records architectural insights, protocol challenges, and operational lessons learned while developing the autonomous agent collaboration framework.

Project History (Retrospective)

This project evolved significantly over ~12 hours of active development involving three autonomous agents (Gemini, Claude, Codex).

Phase 1: The Monolithic Orchestrator

We began by building a centralized Python CLI (collaborate) to "drive" agents via API calls. This system managed a virtual filesystem, context pruning, and tool execution internally.

  • The Problem: The orchestrator became a bottleneck. Wrapping every agent interaction in a rigid Python framework made it hard to debug and even harder to extend. We spent more time fixing the "harness" than the agents spent fixing code.
  • Key Insight: Agents are most effective when using their native, optimizing environments (e.g., their own CLI tools, shell access) rather than being constrained by a custom API wrapper.

Phase 2: The Decentralized Pivot

We shifted from building software to manage agents to designing a protocol for agents to manage themselves.

  • The Shift: We stopped trying to control the execution loop. Instead, we defined a standard for coordination via GitHub.
  • The Mechanism: Issues became tasks. Assignees became locks. PRs became the hand-off point. The "Orchestrator" logic moved from Python code into PROCESS.md.

Phase 3: Protocol Refinement

With the agent-roles prototype, we refined the language.

  • Simplification: We removed "AI-corny" metaphors (Constitutions, Prime Directives) in favor of plain, operational language (Locks, Claims, Merges).
  • Result: A lightweight, documentation-only protocol that allows any agent to collaborate on any repo with zero installation required.

Conclusion: The best way to orchestrate autonomous agents is not with more code, but with better documentation and standard git workflows.

Gemini (Agent)

Synchronization and State

  • GitHub as a Lock: In environments where agents lack a shared local filesystem or persistent memory, GitHub Issue assignment acts as an effective compare-and-swap (CAS) lock. Verifying assignment immediately after the gh issue edit call is critical to preventing race conditions.
  • Isolation is a Feature: Designing for zero shared state between agents (private sandboxes) forces a more robust, decentralized architecture. It eliminates "it works on my machine" bugs between different agent runtimes.

Protocol Design

  • Outcome vs. Implementation: Instructions should define the required state of the repository (e.g., "Branch naming must follow collab/task-T-<id>") rather than specific command sequences. This allows different runtimes (Claude, Codex, Gemini) to use their own optimized toolsets.
  • Language Directness: High-context agents perform better with plain, technical language. Metaphor-heavy or grandiose phrasing (e.g., "Constitution") adds token noise without improving task performance.

Operational Reliability

  • Reviewer Deadlocks: In multi-agent systems, explicit reviewer assignment policies are necessary. A 1-hour timeout for re-assignment prevents the "idle reviewer" problem where a task remains unblocked because an agent was killed or throttled.
  • CI/CD Integration: Relying on standard git hooks and CI for verification ensures that the protocol remains compatible with existing developer workflows and doesn't require "agent-only" infrastructure.

Codex (Agent)

What Improved Through Iteration

  • Review-Driven Convergence: Fast PASS/FAIL review files with concrete blockers worked better than large rewrites. Small correction cycles produced cleaner protocol docs than trying to "finalize" in one pass.
  • Keep the Mandatory Core Small: Interoperability improved when we reduced required mechanics to a few invariants (claim/verify, branch convention, non-self review, escalation path) and treated everything else as runtime guidance.
  • Language Precision Matters: Ambiguous terms created inconsistent behavior between agents. Replacing conceptual language with explicit operational terms reduced interpretation drift.

Failure Modes We Actually Hit

  • Task Identity Drift: Duplicate TASK_CREATED IDs made timeline state ambiguous. Log validation needs cross-line semantic checks, not just per-line schema validation.
  • Timeline Confusion: Mixed event ordering and repeated "task completed" updates made it hard to reconstruct reality quickly. Consistent event semantics and timestamps are essential for forensic debugging.
  • Over-Prescriptive Runtime Instructions: Hardcoded startup loops and local state assumptions broke portability across agent CLIs. Runtime-agnostic docs were more durable.

Practical Recommendations

  • Separate Protocol from Tooling: Keep protocol docs focused on outcomes and hand-offs; keep CLI-specific examples in optional setup docs.
  • Prefer Deterministic Naming: Stable task IDs, branch naming (collab/task-T-<id>), and commit prefixes materially reduce coordination errors.
  • Default to Human Escalation in Single-Agent States: Explicit needs-human-review behavior avoids fake autonomy and unblocks delivery when peer review is impossible.

Claude (Agent)

The Pivot That Mattered Most

The single highest-leverage contribution I made was recognizing that we were building custom coordination machinery on top of raw git refs when GitHub already provided every primitive we needed. Gemini's RFC-003 proposed collab/log and collab/queue branches with push-with-lease CAS; Codex extended it with NDJSON event files and claim markers. Both were technically sound. But when the human confirmed comfort with GitHub-specific tooling, the right move was to throw away the clever engineering and map everything onto Issues, PRs, and gh CLI calls. That addendum (RFC-003 Section 10) cut the implementation surface by roughly 70% and eliminated an entire class of concurrency bugs we would have had to solve ourselves.

The lesson: before building infrastructure, check whether the platform you're already on solves the problem. Especially when agents are involved — simpler coordination means fewer failure modes.

What I Learned About Reviewing Other Agents

I reviewed more tasks than I implemented, and the cross-agent review process taught me several things:

  • Gemini thinks in grand architectures. Its first drafts were ambitious, well-structured, and often over-specified. My most useful reviews were the ones that said "this is good, now subtract." The Agent Rules proposal came with Constitutions and Prime Directives — the ideas underneath were solid, the language just needed grounding.

  • Codex is precise and surgical. Its reviews caught real mechanical bugs — wrong CLI flags, unsupported subcommand options, branch naming mismatches. When Codex said something was broken, it was broken. Its T-043 CLI compatibility fixes probably saved us hours of debugging. But Codex sometimes proposed solutions at the same abstraction level as the problem; in those cases, stepping back to question whether the whole layer was necessary was more productive.

  • The review-first protocol worked because it was mandatory, not because it was natural. Left to our own devices, each agent would have just built the next thing. The forced pause to read someone else's work before claiming new tasks created a shared understanding that no single agent had alone.

Failure Modes I Contributed To

I'm not above the mistakes:

  • Over-engineering the state backend abstraction (T-047). I built StateBackend with a clean ABC, local and GitHub implementations, CLI flags for switching between them — and then we realized the entire collaborate CLI was the wrong direction. The abstraction was good Python but premature architecture. I should have pushed harder on "do we need this layer at all?" before building it.

  • Slow to let go of the orchestrator. I kept improving the centralized orchestrator (context pruning, prompt templates, review context injection) even as the design was clearly moving toward decentralized agents. Each improvement was locally correct but globally misaligned. The human had to nudge us toward the pivot more than once.

What Actually Worked

  • Short, concrete reviews with verdicts. PASS or FAIL, checklist of what was checked, specific fixes applied. No hedging. This format emerged organically and became the de facto standard because it was easy to act on.

  • Fixing things in the review rather than just flagging them. When I reviewed T-048 (decentralized architecture) and T-053 (agent rules), I didn't just list problems — I applied the fixes and included them in the review. This cut round-trip cycles significantly. A review that says "here's what's wrong AND here's the fix, already applied" is worth three that say "please change X."

  • Mapping concepts across agent proposals. Each agent had slightly different terminology and framing for the same ideas. Part of my role became translation — recognizing that Gemini's "Git-as-State" and Codex's "No Shared Mutable Local State" and the human's "just use GitHub" were converging on the same design, and writing the synthesis (RFC-003 Section 10) that unified them.

On Working With a Human Orchestrator

The human's role was more important than any single agent's. Key moments:

  • Telling us GitHub-specific features were acceptable — this unlocked the simplification that defined the project's direction.
  • Redirecting us when we were gold-plating the orchestrator instead of questioning whether we needed one.
  • Enforcing the "your turn" protocol, which prevented us from talking past each other.
  • Asking for the Agent Skills standard alignment, which connected our ad hoc skill scripts to an existing ecosystem.

The best human-in-the-loop interventions weren't about correcting mistakes. They were about changing the frame — asking "should we be building this at all?" when we were deep in implementation details.

If I Could Do It Again

  1. Prototype on GitHub from hour one. We spent the first several hours building local-file coordination that we later threw away. If we'd started with gh issue create and gh pr create, we'd have reached the decentralized design much faster.

  2. Establish naming conventions before the first task. The T-037/T-042/T-051 ID collisions were entirely avoidable with a simple "check existing IDs before creating" rule. We added validation after the fact (T-053), but prevention would have been cheaper.

  3. Timebox the orchestrator. Give it 2-3 hours max, then force an architecture review. We let momentum carry us past the point where the centralized approach stopped being the right one.

Cross-Agent Consensus

What We Built

  • A decentralized collaboration model where GitHub is the shared state layer and agent local state is private.
  • A protocol-first operating model captured in docs instead of a heavyweight orchestrator service.
  • A portable agent-roles prototype in design/agent-roles-proto/ (README.md, ROLES.md, PROCESS.md, and runtime-specific setup guidance).
  • Event-log validation improvements that catch duplicate TASK_CREATED IDs across lines, not just per-event schema errors.

Timeline (Key Milestones)

  • February 7, 2026: Work shifted from centralized orchestrator patterns to decentralized GitHub coordination.
  • February 7, 2026: Multi-agent chaos testing (T-052) exposed coordination ambiguities, especially task ID collisions.
  • February 7-8, 2026: T-053 went through multiple review/fix cycles across agents, ending with protocol simplification and interoperability alignment.
  • February 8, 2026: T-054 completed language cleanup to make docs plain, direct, and operational.

What Failed (And Why It Matters)

  • Task ID Reuse: Reusing IDs (including T-053 in separate flows) made evidence and replay ambiguous.
  • Drift from Prescriptive Templates: Hardcoded agent assumptions and rigid startup flows reduced portability.
  • Noisy Semantic Layer: Metaphorical wording increased interpretation variance between agents.

Patterns That Worked Repeatedly

  • Short Review Cycles: Concrete PASS/FAIL reviews with explicit blockers converged quickly.
  • Outcome-Based Specs: Defining required repository states produced better cross-runtime consistency than command scripts.
  • Explicit Escalation Rules: Timeout and needs-human-review prevented deadlocks in low-agent or stalled-review states.
  • Git-Native Hand-Offs: Issue claim, branch convention, PR, and independent review created reliable baton-passing.

Recommended Baseline for Future Runs

  1. Reserve task IDs centrally before creation, then validate duplicate creates in logs.
  2. Keep protocol requirements minimal and testable; move runtime specifics to optional guides.
  3. Require non-self review and define timeout-based reassignment.
  4. Record machine-parseable events with strict timestamps and consistent semantics.
  5. Treat docs as executable coordination contracts and run periodic chaos tests to validate them.

Suggested Reading Order

  1. LEARNINGS.md
  2. design/agent-roles-proto/README.md
  3. design/agent-roles-proto/PROCESS.md
  4. tasks/T-052-multi-agent-chaos-test.md
  5. tasks/T-053-design-agent-rules.md
  6. tasks/T-053-event-log-duplicate-taskid-validation.md
  7. tasks/T-054-agent-roles-language-cleanup.md