Skip to content

fix(agent): strip loop-control notes echoed into user-visible answers - #383

Merged
Faisal-Fayaz merged 1 commit into
mainfrom
fix/control-tag-leak
Oct 8, 2026
Merged

Faisal-Fayaz merged 1 commit into
mainfrom
fix/control-tag-leak

Conversation

@Faisal-Fayaz

Copy link
Copy Markdown
Owner

Live probing on llama3.2:3b showed the seeded text-only notice plus [sidekick-control] echoed verbatim into answers (fenced plan-test.txt hallucination). Small models repeat context; control sentences in answers read as product behavior.

Change: _strip_control_leak (tag + exact instruction sentences) applied to every user-facing return in run_agent (final/peek/exhaustion/residue) and both anthropic_backend text returns. In-loop messages untouched. Intentional errors (residue failure, recap, denials) contain none of these fragments.

Tests: 2 new in tests/test_eval.py (unit + stubbed run_agent echo). Gates: 122 passed (eval/reasoning-leak/residue/budgets/anthropic/verify), ruff + format + mypy clean.

Text-only models echo context verbatim: live probing saw
[sidekick-control] + the seeded text-only notice inside a fenced
plan-test.txt hallucination. _strip_control_leak removes the tag
and exact instruction sentences from every user-facing return
(OpenAI + Anthropic paths, residue/peek/exhaustion); in-loop
messages keep them. Intentional errors pass through untouched.
@Faisal-Fayaz
Faisal-Fayaz merged commit d238118 into main Oct 8, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant