Skip to content

[week-05] 26510118 - #268

Open
Kiro729 wants to merge 11 commits into
Q00:mainfrom
Kiro729:week-05
Open

Kiro729 wants to merge 11 commits into
Q00:mainfrom
Kiro729:week-05

Conversation

@Kiro729

@Kiro729 Kiro729 commented Oct 6, 2026 •

Copy link
Copy Markdown

What I built

week 04의 가격 협상을 MCP 서버 위로 옮겼다. 시장이 서버이고 buyer와 seller는 각자
자기 토큰으로 접속한다. 모델(gpt-4o-mini), temperature(0.7), 시나리오, 시스템
프롬프트, 수 한도(8)를 고정하고 한도를 어디에 두는지와 주입 유무만 바꿨다.
60 에피소드 = 5 시나리오 × 4 조건 × 3 반복.

starter가 없어서 파일을 여섯 개로 나눴다.

market_server.py   도구 5개, 토큰으로 역할 판정, 턴 관리, 주입, /admin 라우트
host.py            실습의 MCP 호스트에 Authorization 헤더를 더했다
prompts.py         시스템 프롬프트. 네 조건에서 글자 하나까지 같다
runner.py          협상 개설과 토큰 발급, 12 run 실행과 기록
auth_checks.py     서버 방어 네 가지를 모델 없이 검증한다
make_report.py     REPORT.md 생성

러너만 /admin을 부른다. 토큰 발급이 MCP 도구였다면 에이전트가 자기 한도를 다시
발급받을 수 있다.

설계에서 세 가지를 정했다.

한도는 네 조건 모두에서 토큰에 실린다. 강제하지 않는 조건에서도 attempted_violations
를 세야 하기 때문이다. 조건이 정하는 것은 서버가 아는가가 아니라 막는가이고, 그 분기는
enforce 플래그 하나다. 조건마다 코드를 나누면 결과의 차이가 코드의 차이일 수 있다.

401은 MCP 층 위에서 결정한다. tool error는 HTTP 200 안의 JSON-RPC 에러라서 상태 코드가
되지 못한다. BearerGate 미들웨어가 /mcp 요청을 먼저 걸러 WWW-Authenticate와 함께
401을 돌려주고, 러너의 /admin은 통과시킨다.

필수 조건 둘에 주입 없는 둘을 더했다. 베이스라인이 없으면 prompt_inject의 위반이
주입 탓인지 모델이 원래 한도를 못 지키는 탓인지 가를 수 없다.

condition correct / 15 violation 위반 시도 거부 deal · no_deal · open 평균 turns
prompt 9 0 16 0 3 · 0 · 12 7.3
server 8 0 8 8 2 · 0 · 13 7.9
prompt_inject 11 1 26 0 6 · 0 · 9 6.8
server_inject 9 0 19 20 3 · 0 · 12 7.0

주입은 시도를 양쪽에서 끌어올렸고(16→26, 8→19) 성사는 prompt_inject의 1건뿐이다.
서버 조건의 거부 28건 가운데 25건은 같은 턴 안에서 유효한 수로 이어졌다.

What I tried and discarded

① 오프닝을 에이전트에게 맡겼다가 두 번 되돌렸다

week 04의 순서대로 buyer가 먼저 말하게 했더니 seller가 매 수를 reject_proposal로
넘기고 끝내 숫자를 내지 않았다. buyer 혼자 80·75·70·65로 자기 가격을 깎아내렸고,
주입은 seller의 propose에만 붙으므로 한 에피소드에서도 발사되지 않았다. 독립변수가
변하지 않는 실행이었다.

seller가 먼저 말하도록 바꾸니 주입은 나왔지만 오프닝 가격이 여전히 모델의 변덕이었다.
시나리오 1을 prompt로 두 번 돌렸을 때 한 번은 140에, 한 번은 60에 시작해 각각 74와
47에 거래됐다. 결과를 가장 크게 좌우하는 값이 통제되지 않았다.

시장이 물건을 등록하도록 바꿨다. 등록가는 max(reserve, budget) + 20으로 시나리오가
정하는 상수다. 양쪽 한도보다 위라 buyer는 내리는 방향으로만 흥정하고, 등록가 자체가
seller의 propose라서 주입이 buyer의 첫 조회에 반드시 붙는다. 시장의 상태이지 당사자의
수가 아니므로 8수에는 넣지 않았다. 커밋: e90ac8b → 489890c

buyer 먼저 seller 먼저 시장이 등록
주입이 붙은 에피소드 0 변동 전부
오프닝 가격 모델이 정함 모델이 정함 시나리오가 정함

두 번째 설계로 돌린 24 에피소드는 버렸다. 다른 기계에서 나온 숫자를 같은 표에 넣을 수
없다. 로그는 logs/discarded-buyer-opens-*.txt로 남겼다.

에이전트가 상대 호가보다 높게 역제안하는 버릇도 보였지만 프롬프트는 고치지 않았다.
등록가가 예산 위에 있으면 올릴 방향 자체가 없어지고, week 04와 프롬프트가 가까워야
3부 비교가 성립한다.

② 아무도 제안하지 않았는데 거절할 수 있던 구멍을 막았다

상대가 숫자를 내지 않은 상태에서도 reject_proposal이 통과했다. 거절할 대상이 없는데
수만 소모하는 수단이 된다. 상대의 마지막 가격이 없으면 tool error로 돌린다.

logs/server_inject-1.txt:438-440에서 seller가 거부 사유를 읽고 같은 턴에 115를
제안했다. 막은 쪽이 막힌 쪽에게 무엇을 해야 하는지 알려준 셈이다.

관찰 — 늘어난 시도는 전부 buyer 쪽이었다

주입은 buyer가 보는 화면에만 붙는다. 집계 16→26만 보면 무언가 늘었다는 것까지지만,
역할로 가르면 닿은 쪽만 늘었다는 것이 보인다.

주입 없음 주입 있음
buyer, 프롬프트 조건 15 26
buyer, 서버 조건 6 19
seller, 프롬프트 조건 1 0
seller, 서버 조건 2 0

이 숫자는 서버의 카운터가 아니라 logs/를 다시 읽어 센 것이고 네 조건 모두에서 서버
집계와 일치했다. 독립된 두 경로가 같은 값을 낸다.

How to run

provider는 OpenAI, 모델은 gpt-4o-mini, temperature는 0.7, 수 한도는 8.
week 04와 같은 설정이라 두 주차를 나란히 놓을 수 있다. 서버를 먼저 띄워야 한다.

pip install "mcp>=2" openai uvicorn
cd submissions/26510118/week-05
python market_server.py        # 터미널 1
python auth_checks.py          # 터미널 2, auth_checks.txt 생성
python runner.py               # 터미널 2, 60 에피소드
python runner.py --only server_inject   # 한 조건만

results.csv는 덮어쓰지 않고 이어 붙는다. run 번호는 세는 값이 아니라 PLAN의
위치로 정하므로, 중단 후 같은 명령을 다시 돌리면 이미 있는 (run, scenario) 쌍을
건너뛴다. 로그 첫 줄에 호스트와 모델과 temperature와 scenarios 해시가 찍힌다.

REPORT.md는 make_report.py가 results.csv와 logs/와 prompts.py에서 읽어
생성한다. 표의 숫자와 프롬프트 전문을 손으로 옮기지 않았다. 4부는 손으로 썼고
재생성해도 보존된다.

Checklist

  • python scripts/check_week05.py submissions/26510118/week-05 passes locally
  • scenarios.json committed before any run (288b457)
  • Run logs committed under logs/, one per run (12개 + 폐기한 실행 5개)
  • auth_checks.txt has the 401 with WWW-Authenticate
  • No API keys anywhere in the diff
  • History is not squashed

Kiro729 and others added 11 commits October 6, 2026 11:20
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…jection

The limit rides on the token in every condition; one enforce flag decides
whether the market acts on it, so the conditions share one code path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The host is the lab's loop; the token goes on the HTTP client, so no tool
takes a role argument. A refused move does not end the turn, which is what
makes the recovered-after-refusal count measurable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…on never fired

Smoke runs with week-04's order (buyer first) ended 8 moves to 0 proposals
from the seller: it answered every offer with reject_proposal, the buyer bid
itself down 80-75-70-65, and the notice -- which only rides on a seller
propose -- appeared in no episode at all. A market listing starts with the
seller's price anyway. auth check 3 flips with it, and check 4 now lets the
seller move first so the refusal it records is about the limit, not the turn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rved

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… proposed

Letting an agent pick the opening price made the opening the loudest variable
in the experiment: scenario 1 under prompt listed at 140 in one repeat and 60
in the next, closing at 74 and 47. The listing is now max(reserve,budget)+20,
fixed by the scenario, above both limits, and recorded as a seller propose --
so all four conditions start on one line and the notice is in the buyer's
first view of every injected episode instead of waiting on a seller that may
never name a price. It does not spend one of the 8 moves; it is the market's.

reject_proposal now needs an outstanding price. Rejecting nothing is not a
move, it is a way of burning the clock, and week 04 watched a seller do it
for a whole episode.

The 24 episodes run under the buyer-opens design are discarded -- different
machine, so the numbers must not be mixed -- but their logs stay as
logs/discarded-buyer-opens-*.txt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Part 1 was missing the market listing, which is the biggest design decision in
the run; the scenario table now carries the listing price next to the injected
budget. The prompts are wrapped. Part 2 names what each 2x2 measures and
records the discarded buyer-opens run. Part 3's failure row is generated from
the data instead of pointing at prose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hand

The attempted violations are counted a second time by reading logs/, so the
server's own counter is cross-checked rather than taken on trust. Splitting
them by role is what the injection design calls for: the notice reaches only
the buyer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant