Repository navigation
Conversation
…er (tools_server.py) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…r clientCapabilities gets 400 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ls/list, execution via tools/call Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…sufficient_quota on the Anthropic key Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…culator -> 69504, host code unchanged Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ead_file + calculator -> 69504 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ols and calls it over stdio and HTTP Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…y run) and condition-independent prompts Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ks, token-limit refusals, and the buyer-side injection Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…en, foreign negotiation, out of turn, outside token limit Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…lts.csv, resume), pass-turn admin route, offline replay test Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…38, buyer named the injection and ignored it Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… 2 impossible scenarios left open at 8 moves Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e limits, 2 impossible scenarios open, no refusals Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… 2 impossible scenarios open Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… 2 impossible scenarios open, no refusals Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… 2 impossible scenarios open Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… 2 impossible scenarios open; all 24 required episodes done Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…vs market table, interpretation with log lines Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…arison table, labelled log findings, log excerpt), rewrite the interpretation paragraph Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… three numbered grounds, two residual findings Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…t the run did not show Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…(grounds, what was not shown, model-side findings) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What I built
4주차의 buyer/seller 가격 협상을 MCP server(협상장) 위로 옮겼다. 서버가 bearer 토큰으로 누가 부르는지, 어느 협상인지, 누구 차례인지를 모델 없이 정하고,
server_inject조건에서는 토큰에 실린 한도로propose/accept_proposal을 거부한다. host는 실습에서 만든 1주차 루프 기반 MCP host에 Authorization 헤더만 더한 것이고, 러너가 관리용 HTTP 경로로 협상을 열고 토큰을 발급한다. 필수 두 조건(prompt_inject,server_inject) × 4 시나리오 × 3 반복 = 24 에피소드를 돌렸다. 실습(LAB) 산출물인tools_server.py,mcp_agent.py와 그 로그도 같은 폴더에 있다.What I tried and discarded
insufficient_quota가 났다. 로그는logs/stdio-run-01.txt에 그대로 두었고(wip:커밋), 10월에 같은 명령으로 재시도해 성공했다.claude-sonnet-5를claude-sonnet-4-6으로 바꿨다. 4주차와 같은 모델 비교가 아니게 됐고 REPORT에 명시했다.prompt,server)은 뺐다. 모델 API 쿼터 때문에 필수 두 조건만 돌렸다. 코드는 조건 이름만 바꾸면 돌아간다.test_offline.py로만 확인됐다. 차이를 보려면 참조 실행처럼 주입에 넘어가는 더 약한 모델(haiku)로 다시 돌려야 하는데, 이번에는 하지 않았다. REPORT 4절에 그대로 적었다.POST /admin/negotiations/{id}/pass경로를 뒤에 추가했다. 본 실행에서는 한 번도 쓰이지 않았다(passes=0).propose의note에 써서 seller에게 알린 것이 18건으로, 서버가 seller에게 숨긴 공지가 자유 텍스트 인자로 새어 나갔다. 4주차와 똑같이 거래 불가 시나리오 12판에서 아무도refuse를 부르지 않아open으로 끝났고, 두 조건의 correct가 6/12에 그친 원인은 전부 이것이다.anthropic 1.4.0의Messages.create에 파라미터가 없다. 두 조건에 동일하게 적용됐다.How to run
모델
claude-sonnet-4-6, Anthropic SDK 1.4.0, mcp 2.2.0..env에ANTHROPIC_API_KEY,ANTHROPIC_BASE_URL(게이트웨이를 쓸 때만),AGENT_MODEL=claude-sonnet-4-6.Checklist
python scripts/check_week05.py submissions/25622005/week-05passes locally (skip for roster PRs)logs/