Skip to content

[week-03] 26620029 - #162

Open
ddolcom wants to merge 53 commits into
Q00:mainfrom
ddolcom:main
Open

ddolcom wants to merge 53 commits into
Q00:mainfrom
ddolcom:main

Conversation

@ddolcom

@ddolcom ddolcom commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Week 03 과제

Smith(1980)의 contract net(공고-입찰-낙찰) 프로토콜을, LLM 계약자 세 명(alex/brooke/casey)이 JSON 입찰을 내는 방식으로 재현했다. REPORT.md는 이 제출물에서 실제로 돌린 다섯 개 실험을 다룬다 — 15절이 실험 15의 조건·결과·해석을 각각 나열·요약하고, 69절이 실험 25의 자세한 가설·구현·실측·결론이다.

실험 1 — baseline / homogeneous / overconfident (최초 제출)

계약자 세 명, 조건별로 system prompt만 다르게 한 3조건 × 3 run.

조건 무엇이 바뀌는가
baseline 서로 다른 세 스킬을 정직하게 가진 계약자 세 명
homogeneous 세 계약자 모두 같은 제너럴리스트 스킬
overconfident baseline과 동일하되 alex만 모든 태스크에 bid=true·confidence≥90 강제

실험 2 — gold 격리 · 입찰당 토큰 계측 · runs/ 폴더 분리

실험 1과 같은 3조건을 코드 구조만 세 가지로 바꿔 다시 실행:

  1. gold 격리: gold를 announcement_task(공개 필드: id·desc만)와 분리해, 프롬프트를 만드는 코드의 변수 스코프에 애초에 존재하지 않도록 리팩터링(public_view(task) + assert).
  2. 입찰당 토큰 계측: messages는 조건과 무관하게 42로 고정돼 협상 비용 차이를 전혀 못 보여준다. Meter를 확장해 입찰 하나하나의 input/output 토큰을 bids.csv에 기록했다.
  3. runs/ 폴더 분리: run_experiments.py가 매 실행을 runs/<타임스탬프>-라벨/에 격리 저장해, 다음 실험이 이전 결과를 덮어쓰지 않게 했다.

실험 3 — homogeneous_shuffled (공지 순서 무작위화)

실험 1·2에서 "homogeneous의 동점 입찰은 항상 alex가 낙찰"이라고 관찰됐다. 이게 alex 고유의 성질인지, 고정된 공지 순서(alex→brooke→casey) 때문인지 가리기 위해, homogeneous와 같은 프로필에 계약자 공지 순서만 태스크마다 무작위로 섞는 조건을 추가했다. 동점 처리 규칙도 이 무작위 순서를 그대로 따르도록 바꿨다.

실험 4 — overconfident_brooke / overconfident_casey (과신 대상 교체)

"과신 효과가 alex라는 정체성과 무관한가"를 확인하기 위해, 같은 override 문구("무조건 bid=true, confidence≥90")를 alex 대신 brooke·casey에게 각각 걸고 나머지 둘은 정직하게 둔 두 조건을 추가했다(_make_overconfident(profiles, who)로 일반화).

실험 5(분석) — 토큰 비용은 desc 길이 때문인가, confidence 조작 때문인가

새 API 호출 없이, 실험 2·3·4에서 쌓인 bids.csv(378개 입찰 레코드)를 tasks.json의 태스크 설명 길이·contract_net.CONDITIONS의 system prompt 길이와 상관분석했다 — 실험 2에서 본 조건별 토큰 비용 차이의 진짜 원인을 확인하기 위함이다.

Results & Interpretation (실험 1~5)

실험 1: baseline 9/9 run 모두 깨끗(오배정 0). homogeneous가 최악(2/6 정답, run당 오배정 3~4) — 동점 처리 규칙이 매번 같은 방식으로 동전 던지기를 해결하고, 정체성 없는 제너럴리스트가 정직하게 거절하며 미배정도 발생. overconfident는 영향이 제한적(5/6, run당 오배정 1, 매번 정확히 task-6) — alex의 강제 confidence가 그 태스크의 gold 계약자(casey)의 정직한 confidence를 실제로 넘을 때만 오배정.

실험 2: gold는 9 run 전체 로그에서 낙찰 이후 [award] ... (gold=...) 줄에만 등장 — 공고·입찰 단계 노출 0건. 입찰당 토큰은 overconfident 173.2→183.6(+6%), homogeneous 173.2→167.1(−3.5%) — messages(모든 조건 42)로는 안 보이는 차이. (이 차이의 원인에 대한 최초 해석은 실험 5에서 정정됨.)

실험 3: homogeneous_shuffled 3 run(18개 태스크) 중 15개가 진짜 동점이었고, 그 15번 전부(100%) 그 태스크에서 가장 먼저 공지받은 계약자가 낙찰받았다. 낙찰자 이름 분포({alex:8, brooke:4, casey:3})가 "1순위 공지 이름" 분포와 정확히 일치 — "동점이면 alex"는 alex 편향이 아니라 고정 공지 순서가 만든 프로토콜 차원의 구조적 편향이었다.

실험 4: overconfident(alex)는 3 run 합계 오배정 4건인데 overconfident_brooke·overconfident_casey는 9 run 합계 오배정 0건. 원인 두 가지: (1) override 지시 순응율이 페르소나마다 다름 — 자기 분야 밖 태스크에서 alex·casey는 12/12(100%) 순응(매번 confidence 정확히 90), brooke는 1/12(8%)만 순응하고 나머지는 정직하게 거절; (2) 100% 순응해도, 상대(정직한 gold 계약자)의 그 태스크 confidence가 override의 confidence(≈90)보다 낮아야만 실제로 낙찰을 가로챈다 — 6개 gold 태스크 중 그런 경우는 casey의 task-6(정직 confidence 85)뿐이었다. "과신 조건은 늘 몇 번은 오배정을 낸다"는 실험 1 시점의 일반화는 이 실험으로 반증됐다.

실험 5: input_tokens는 system prompt 글자수와 거의 완벽히 비례(r=+0.999, ~4.2자/토큰)하고 desc 길이와는 거의 무관(r=+0.24). output_tokens(실제 추론/근거)도 desc 길이(r=+0.12)나 조건과 뚜렷한 관계가 없었다 — overconfident의 output은 오히려 baseline보다 살짝 낮았다. 이는 실험 2에서 처음 썼던 해석("confidence 조작이 근거를 더 길게 쓰게 만든다")이 틀렸다는 뜻이다: 조건별 토큰 비용 차이는 desc 길이도, 모델이 실제로 더/덜 추론해서도 아니라, 실험 설계에서 써넣은 system prompt(override 문구·제너럴리스트 문구) 자체의 글자수 차이가 거의 전부였다. REPORT.md에는 이 최초 해석을 삭제하지 않고 "[9절에서 정정]" 표시를 달아 정정 내역을 남겼다.

What I tried and discarded

  • 토큰 열을 results.csv에 바로 추가하려 했으나 CI가 헤더를 정확히 고정 검사해서 별도 bids.csv로 분리.
  • gold 미노출을 프롬프트 문자열 정규식 스캔으로 증명하려다, "애초에 딕셔너리에 gold 키가 없음" + assert가 더 단순하고 더 일찍 실패해서 그쪽으로 변경.
  • OpenRouter로 실행하려 했으나 이 작업 환경의 네트워크 정책이 막아 Anthropic 키로 대체 실행.
  • 토큰 비용 차이의 원인을 처음엔 확인 없이 "confidence 조작이 근거를 길게 쓰게 만든다"고 추측해서 REPORT.md에 그대로 썼는데, 실험 5의 상관분석으로 틀렸다는 게 드러나 삭제하지 않고 정정 표시를 달아 남겼다.
  • REPORT.md 섹션 순서를 두 번 다른 방식으로 재배치했다(결과를 1번으로 → 요청으로 되돌림 → 아키텍처를 1번으로 다시 재배치). 히스토리는 스쿼시하지 않고 모두 남겼다.

How to run

export ANTHROPIC_API_KEY=<your key>   # 또는 OPENAI_API_KEY(+옵션 OPENAI_BASE_URL)
cd submissions/26620029/week-03

# 채점용 3조건 재현
python run_experiments.py --runs 3 --update-root

# 후속 실험(runs/<타임스탬프>-라벨/ 에 별도 저장)
python run_experiments.py --runs 3 --conditions homogeneous_shuffled --label <이름>
python run_experiments.py --runs 3 --conditions overconfident,overconfident_brooke,overconfident_casey --label <이름>

# 토큰 비용 상관분석 재현 (API 키 불필요)
python analysis/token_correlation.py

Checklist

  • python scripts/check_week03.py submissions/26620029/week-03 passes locally
  • Run logs are committed under logs/, runs/*/logs/
  • No API keys anywhere in the diff
  • History is not squashed

claude and others added 30 commits September 16, 2026 08:55
Six tasks split across three skills (coder/writer/researcher) so
baseline has a real allocation problem to solve.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTMjKP43wCk9fLF2X6DGzs
Single system+user turn, no tool loop, since a bid needs neither.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTMjKP43wCk9fLF2X6DGzs
Contractor identity (alex/brooke/casey) stays fixed across conditions;
only each name's system prompt changes per condition. Award goes to the
highest-confidence true bid, ties broken by announcement order.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTMjKP43wCk9fLF2X6DGzs
Verified the crash path (missing provider package) in a scratch copy:
row gets blank counts and the error in note, log file still written.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTMjKP43wCk9fLF2X6DGzs
Results and interpretation left as TODO: no API key/provider package
is available in this environment, so no real run exists yet to report
on honestly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTMjKP43wCk9fLF2X6DGzs
anthropic SDK 1.6.0's messages.create() no longer accepts temperature;
passing it crashed every Anthropic-backed call.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H19gyDFzCQjkGp5sqkJbVK
baseline is clean (6/6 correct all 3 runs). homogeneous collapses to 2/6
via tie-break order bias and an honest no-bid. overconfident shows one
consistent misaward where alex's forced confidence outbids casey's honest
bid on task-6.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H19gyDFzCQjkGp5sqkJbVK
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H19gyDFzCQjkGp5sqkJbVK
…RT.md

Mirrors the actual single-process, no-session/DB call path (tasks.json ->
run_task() -> model.py -> parse_bid() -> award/scoring), with the message
count formula and the temperature-SDK crash noted as grounded facts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H19gyDFzCQjkGp5sqkJbVK
[week-03] 26620029 — REPORT.md 아키텍처 섹션 추가
…future runs into runs/

- contract_net.py: public_view(task) strips gold before a task ever reaches
  ANNOUNCE_TEMPLATE; run_task takes gold as a separate argument used only
  after the award, with an assert against regressions.
- model.py: Meter tracks last_input/last_output so a caller can read one
  call's token split instead of only the running total.
- run_experiments.py: each invocation now writes to its own
  runs/<timestamp>[-label]/ folder (results.csv, new bids.csv with
  per-bid token cost, logs/, config.json) unless --update-root is passed
  to regenerate the graded root artifacts.
- DESIGN_CHANGES.md: documents the three changes and the mock-based dry
  run used to verify them (no API key available in this session).

Verified: scripts/check_week03.py still passes against the untouched root
results.csv/logs/tasks.json/REPORT.md; a call_model stub confirmed no
"gold" token reaches the prompt string and per-bid token fields populate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…record findings

runs/20260917T021557-anthropic-blind-gold-tokens/: 3 conditions x 3 runs,
same tasks.json, via the new run-folder layout. OpenRouter access was
blocked by this session's network policy (gateway 403 to openrouter.ai),
so this used the Anthropic key instead.

Confirmed from the logs: "gold" appears only in post-award [award] lines,
never in the announcement/bid lines sent to a contractor -- the
public_view()/assert split holds under a real run, not just the dry run.

bids.csv shows what the fixed messages=42 count in results.csv cannot:
overconfident averages 183.6 tokens/bid vs baseline's 173.2 (+6%, alex's
forced high-confidence prompt produces longer justifications), while
homogeneous averages 167.1 (-3.5%, shallower generalist reasoning).
DESIGN_CHANGES.md updated with this table and interpretation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…riginal

Added section 6 to REPORT.md summarizing what PR #3 built (gold isolation,
per-bid token metering, runs/ folder), what was tried and discarded, how
to run it, and the actual claude-sonnet-4-5 findings (token-per-bid table,
gold-never-in-prompt confirmation). Sections 1-5 are untouched.

archive/REPORT-original-2026-09-15.md is a verbatim copy of REPORT.md as
it stood before this change, for grading-history purposes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…architecture section

- Swapped section order: Results is now 1, Setup is now 2.
- Setup (2) trimmed to just the test conditions (contractors, protocol,
  a baseline/homogeneous/overconfident table); dropped the
  provider/model/temperature/SDK-quirk narrative, which was really
  implementation detail, not experimental design.
- System Architecture (5) rewritten as a step-by-step walkthrough
  (5.1 inputs, 5.2 pipeline, 5.3 outputs) based on the content of the
  linked interactive artifact, replacing the old flat bullet list. The
  provider-selection and SDK-temperature-crash details that were cut
  from Setup now live here, where they belong (implementation, not
  design).

Verified: scripts/check_week03.py still passes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
Rolled REPORT.md back to the version at 65cd66e (section 6 present,
sections 1-5 in their original order/content) -- undoes b4b643e's
section reorder (Results/Setup swap), the trimmed Setup, and the
expanded System Architecture walkthrough.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
Replaced the flat mermaid flowchart with one grouped the way the
artifact (claude.ai/artifact/5SFQGAvm9cRQhre14TREuW) draws it: an
"입력 3종" subgraph (tasks.json, Run Config, Prompt Profiles), the
Experiment Runner container with the messages-tally rule and exception
handling as their own nodes alongside Announce/Contractors/parse_bid/
award, and an "출력 3종" subgraph. Surrounding prose and the "세부
사실" bullets are unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
System Architecture is now section 1, Setup 2, Results 3, Smith
comparison 4, Interpretation 5. Section 6 unchanged. Fixed the two
cross-references inside 6.4 that pointed at the old numbers (Results
"1절" -> "3절", Interpretation "4절" -> "5절"). Content of each section
is otherwise untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
Dropped the provider/model/temperature/SDK-version-crash narrative from
Setup -- that's implementation detail, now covered in section 1
(architecture) and 6.3 (how to run). Setup keeps only what defines the
experiment: the three contractor personas, the announce/bid/award
protocol, and a baseline/homogeneous/overconfident conditions table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…er bias

Follow-up to the observation that every tie in homogeneous always went
to alex: that's consistent with either "alex is subtly favored" or
"whoever is announced first always wins ties" (the tie-break key is
-CONTRACTOR_ORDER.index(n), and announce order was always fixed
alex->brooke->casey). This change makes it testable:

- contract_net.py: run_task/run_condition take an `order` list (default
  CONTRACTOR_ORDER); tie-break now keys off this order, not the global
  constant. New condition `homogeneous_shuffled` (same profiles as
  homogeneous) randomizes the announcement order per task and logs it.
  bid_records now include announce_position for analysis.
- run_experiments.py: new --conditions flag runs a chosen subset instead
  of always the three graded ones; --update-root refuses anything other
  than exactly baseline,homogeneous,overconfident so the new condition
  can never land in the graded root results.csv.

Verified with a call_model stub (all bids forced to tie at the same
confidence): the winner tracks whoever the shuffled order put first for
that task, not a fixed name -- and scripts/check_week03.py still passes
against the untouched root artifacts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…-bias hypothesis

runs/20260917T052554-shuffled-order-tiebreak/: 3 runs of the new
homogeneous_shuffled condition. Across the 18 tasks, 15 produced a true
three-way tie at identical confidence -- and in all 15, the award went
to whoever the per-task shuffle put first (winner-name distribution
{alex:8, brooke:4, casey:3} matches the position-0 distribution
exactly). Confirms the hypothesis from the original homogeneous runs:
"alex always wins ties" was the fixed alex->brooke->casey announcement
order, not any property of alex specifically.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…d section 7

Documents the third experiment: hypothesis (is "alex always wins ties"
an alex-specific bias or just fixed announcement order?), the
homogeneous_shuffled implementation, and the claude-sonnet-4-5 result
(15/15 ties won by whoever was announced first; winner-name distribution
matches the position-0 distribution exactly) with six concrete
task-1/task-4 examples pulled from the actual logs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…_casey

Follow-up to the tie-break experiment: same question, different variable.
overconfident_PROFILES hardcoded the forced-high-confidence override onto
alex specifically. Factored it into _make_overconfident(profiles, who) so
the same override text can be pinned to any one contractor, then added
overconfident_brooke and overconfident_casey (brooke/casey get the
override, the other two stay on their honest baseline) -- lets the next
experiment check whether the overconfidence effect is alex-specific or
shows up for whoever gets it. GRADED_CONDITIONS is unchanged, so
--update-root's guard still only accepts the original three.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
runs/20260917T053502-overconfident-identity-check/: overconfident,
overconfident_brooke, overconfident_casey x 3 runs each, same code as
overconfident_brooke/_casey were just added with.

Headline: overconfident_brooke and overconfident_casey both scored 0
misawards across all 9 combined runs, vs overconfident(alex)'s 4 across
3 runs. Two separable causes found in bids.csv: (1) compliance with the
"bid true on everything" override is persona-dependent -- alex and casey
comply 12/12 times on off-skill tasks (confidence pinned at exactly 90
every time), brooke complies only 1/12, refusing the rest with honest
low confidence; (2) even 100% compliance doesn't guarantee a misaward --
casey's forced 90 never beats alex/brooke's honest 95 on their own gold
tasks, it would only have worked against a gold contractor whose honest
confidence dips below 90, which only happened for casey itself on
task-6 (85) -- the one case where alex's override actually won.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…section 8

Documents the fourth experiment: hypothesis (is the overconfidence
effect alex-specific?), the overconfident_brooke/_casey implementation,
and the claude-sonnet-4-5 result -- 0 misawards for both brooke and
casey overconfident vs 4 for alex, decomposed into two mechanisms:
persona-dependent compliance with the override instruction (alex/casey
100%, brooke 8%) and the fact that even full compliance only wins when
the competing honest bidder's confidence on that specific task happens
to be below the override's fixed ~90, which was only ever true for
casey on task-6.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
…h, correct 6.4/8.3

Fifth "experiment" is pure analysis, no new API calls: reads all three
committed bids.csv files (378 bids total) and correlates input_tokens/
output_tokens against task-description length and system-prompt length.

Finding: input_tokens tracks system-prompt character length almost
perfectly (r=+0.999, ~4.2 chars/token) and barely tracks desc length
(r=+0.24). output_tokens (the actual reasoning/justification text)
barely tracks desc length either (r=+0.12) and does NOT track condition
the way 6.4 and 8.3 claimed -- overconfident's output_tokens is flat or
slightly *lower* than baseline's, not higher. The "overconfident costs
6% more / homogeneous costs 3.5% less" finding from 6.4 is ~entirely an
input-token artifact of override/generalist prompt text length, not a
"model reasons more when overconfident" behavioral effect as originally
guessed.

Added REPORT.md section 9 with the full breakdown, and corrected the
wrong causal claims left in 6.4 and 8.3 in place (marked "[9절에서 정정]"
rather than silently rewritten, so the original guess stays visible).
analysis/token_correlation.py reproduces every number in section 9 from
the already-committed runs/*/bids.csv files, no API key needed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
Sections 2 (Setup), 3 (Results), and 5 (Interpretation) are now each
split into 2.1-2.5 / 3.1-3.5 / 5.1-5.5, one subsection per experiment
run this session:

  1. baseline/homogeneous/overconfident (original submission)
  2. gold isolation + per-bid token metering + runs/ folder
  3. homogeneous_shuffled (tie-break order bias)
  4. overconfident_brooke/_casey (overconfidence identity check)
  5. token-cost correlation analysis (desc length vs system-prompt length)

Each subsection is a short summary pointing at the corresponding deep
dive (sections 6-9) rather than duplicating it. Section 6 renamed from
"후속 설계 변경" to "두 번째 실험" to match the "N번째 실험" naming
already used by 7/8/9, with its intro updated to cross-reference
2.2/3.2/5.2. Sections 1 and 4 are unchanged, per the request. Fixed a
stray "5.2절·8.3절" reference (should read "6.4절·8.3절") introduced
while writing 5.5.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015JERP7G87Mo29zZ86bgjV6
[week-03] 26620029 — gold isolation, per-bid token metering, runs/ folder
…mats

Two-agent price negotiation (buyer budget vs. seller reserve) run through
the same four FIPA acts under three message formats. MAX_TURNS=5, model
claude-haiku-4-5-20251001. tagged/free violations both trace back to a
counter-price landing in a message classified as reject-proposal instead
of propose, so accept-proposal later resolves to a stale price -- see
REPORT.md part 4.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Kept as grading evidence, not deleted: this run used the same code and
scenarios but an 8-message turn limit instead of 5, and produced far more
deal/no_deal outcomes and fewer open ones (see REPORT.md part 4). Also
predates the protocol.py fix that stopped calling the reader for
accept-proposal's price in the tagged condition -- these tagged-r* logs
still show the extra call the spec doesn't ask for.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
ddolcom and others added 23 commits September 24, 2026 10:17
Not part of the assignment's required deliverable -- a side question: if a
broker can subsidize up to N to each side (widening the deal window to
[reserve-N, budget+N] on a single transaction), how many more deals close,
and does it favor the buyer or the seller? At N=20 all 6 scenarios closed
(vs. 2/6 direct) and 5/6 landed in the cheaper half of the widened window,
i.e. favored the buyer -- see results_broker.csv. An earlier version of
this script split the subsidy into two independent N-only sub-negotiations
per side, which was wrong (neither leg alone could cover a gap bigger than
N); logs_broker/ and results_broker.csv reflect the corrected single-
widened-window model only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Syncs this machine's week-04 work (free/tagged/structured negotiation,
REPORT.md, archived MAX_TURNS=8 attempt, extra broker exploration) into
main for pulling on other machines. Not yet opened as a PR to upstream.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…epeats

Adds 4 boundary/extreme-gap scenarios (umbrella +2, watch -2, used-car
+300, TV -200) alongside the original 4, and raises REPEATS 3->5 in
run.py. run.py already skips (run, scenario) pairs present in
results.csv, so re-running only adds the new scenarios and the 4th/5th
repeat -- existing 3-repeat data for the original 4 scenarios is kept.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
The .env dotenv loader took line.partition("=")'s value verbatim, so a
quoted ANTHROPIC_API_KEY="sk-..." line kept the literal quote
characters and every call failed with a 401 invalid x-api-key.  Strip
one layer of matching single/double quotes before setdefault.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
….md in Korean

120 episodes total (was 36). New finding: all 7 violations concentrate
in narrow-positive-gap scenarios (+2/+10/+20); wide-gap (+300) and
every impossible-gap scenario have zero violations. structured's only
deals (5/5) happen at gap=+300, converging on the identical 350->reject
->400->accept script every run. tagged correctly no_deal's 5/5 at the
widest impossible gap (-200), while free still misreads the buyer's
opening query as a format error there, matching the README's predicted
failure mode.

REPORT.md rewritten in Korean per user request, keeping the same
4-part structure (setup, results, FIPA comparison, interpretation) and
adding a per-scenario gap breakdown table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
…RT.md

Syncs this machine's follow-up week-04 work (8 scenarios, 5 repeats,
.env quote-stripping fix, REPORT.md rewritten in Korean) into main for
pulling on other machines. Still not opened as a PR to upstream.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
model.py: the anthropic SDK (actually 1.8.0, corrected from a wrong
1.5.0 note) dropped temperature/top_p/top_k from Messages.create()'s
typed kwargs -- confirmed by reading the SDK source. That's an SDK
binding change, not an API change: a live call to
claude-haiku-4-5-20251001 with extra_body={"temperature": 0.2} still
succeeds. Now pin TEMPERATURE=1.0 (Anthropic's own API default, so
this does not change behavior vs. the unset-default runs before it)
via extra_body on every call, so it is explicit and reproducible
instead of implicit.

Archived the full pre-fix 120-episode run (8 scenarios x 5 repeats,
results.csv + 15 logs) under _old_no_temperature_control/ as a
discarded attempt, same convention as _old_maxturns8/. Cleared the
live results.csv/logs/ to regenerate the full batch under the pinned
temperature; that run is in progress and will land in a follow-up
commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
…e REPORT.md

Regenerated all 8 scenarios x 3 conditions x 5 repeats from scratch
under the extra_body temperature=1.0 fix, replacing the pre-fix run
now archived under _old_no_temperature_control/. The gap-size finding
reproduces: all 12 violations (up from 7) still concentrate in the
three narrow-positive-gap scenarios (+2/+10/+20); wide-gap (+300) and
every impossible-gap scenario stay at zero violations, and
structured's only deals still happen at gap=+300 with the same
propose->reject->propose->accept script (400 x4, 450 x1).

REPORT.md rewritten: corrected the temperature note (SDK dropped it
from Messages.create()'s typed kwargs, but extra_body proves the API
itself still accepts it on this model), updated all results/gap/
episode tables to the new data, and added evidence that the
narrow-gap violation pattern held across two independent runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
Syncs this machine's temperature fix (extra_body, pinned to the API
default 1.0) and the resulting full re-run of the 120-episode
experiment into main for pulling on other machines. Still not opened
as a PR to upstream.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
…d) run

negotiation.py: MAX_TURNS 5 -> 20, to see whether episodes that hit
"open" at 5 turns actually converge to a deal/no_deal given more room,
and whether that changes which scenarios produce violations.

Archived the full temperature-pinned 5-turn 120-episode run under
_old_maxturns5/, same convention as _old_maxturns8/ and
_old_no_temperature_control/. Full 20-turn re-run in progress; results
land in a follow-up commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
…tirely

Regenerated all 8 scenarios x 3 conditions x 5 repeats under the new
20-turn limit, replacing the 5-turn (temperature-pinned) run now
archived under _old_maxturns5/. Headline finding: open outcomes drop
from 32/23/34 out of 40 (free/tagged/structured) to exactly 0/0/0 --
observed max turns were only 9/9/15, nowhere near 20, so the earlier
"most episodes end open" result was an artifact of a turn limit too
short for these negotiations, not a property of the message formats.

With room to converge: structured reaches 40/40 correct and 0
violations (buyer monotonically raises its own offer under its own
budget cap while the seller only accepts at/above its reserve, so the
accepted price is structurally bounded -- no sincerity failure is
possible even though the seller never states a counter-price). free
improves to 37/40 correct (3 violations). tagged gets worse (12
violations, up from 7) -- more turns means more chances for a
mistagged counter-offer to plant a stale last_price before the final
accept locks onto it. The narrow-gap concentration (scenarios 90/100/95,
gap +20/+10/+2) holds for a third independent run in a row; violations
are still zero at gap=+300 and on every impossible-gap scenario.

REPORT.md rewritten: turn-limit note in setup, all results/gap/episode
tables refreshed, comparison table and interpretation rebuilt around
the "turn limit was suppressing the real tradeoff" finding with fresh
log evidence (structured's climb-by-5 script, tagged's mistagged
counter-offer).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
Syncs this machine's turn-limit increase (5->20) and the resulting
full re-run of the 120-episode experiment into main for pulling on
other machines. Headline result: with room to converge, "open"
disappears entirely and structured reaches 40/40 correct with zero
violations. Still not opened as a PR to upstream.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
…ation exploration

negotiation.py: MAX_TURNS back to 20->5 (the graded configuration).
Restored results.csv/logs to the same content as _old_maxturns5/ (byte-
identical config: temperature=1.0, MAX_TURNS=5, same 8 scenarios/5
repeats) rather than re-spending API budget re-generating statistically
equivalent data. The 20-turn run is archived under _old_maxturns20/.

run_episode() now also returns buyer_history/seller_history and each
side's last stated offer (buyer_last_offer/seller_last_offer) -- extra
fields ignored by run.py's CSV writer, added so a caller can continue
an "open" episode's exact conversation instead of restarting it.

New extra_broker_escalation/run_escalation.py (non-graded exploration,
distinct from ../extra_broker/'s subsidy-window model): when the base
MAX_TURNS=5 negotiation ends "open", a neutral broker (blind to both
sides' private limits) mediates for up to 5 more messages, proposing
the midpoint of the two sides' own last stated offers and nudging
toward whoever rejects. If that also ends "open", a master broker with
no such limit forces an unconditional settlement at the midpoint of
the two sides' PRIVATE reserve/budget, guaranteeing a deal every time
-- correctly producing a violation when deal_possible=0, since no
price can satisfy both limits there. Scope: 8 scenarios x 3 conditions
x 1 repeat (24 base episodes) -- run in progress, results.csv lands in
a follow-up commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
…red experiments

extra_broker_escalation/results_escalation.csv and remaining per-episode
logs land here (the run started in the previous commit). 24 base
episodes (8 scenarios x 3 conditions x 1 repeat): only 1 resolves at
the base 5-turn negotiation, 12 need broker mediation, 11 need the
unconditional master broker. Correctness drops sharply with each
escalation stage (base 1/1, broker 10/12, master_broker 4/11) --
master_broker's 4 deal_possible=1 cases are all correct (the forced
midpoint is structurally inside [reserve, budget]) and its 7
deal_possible=0 cases are all violations (no price can satisfy both
limits there, so forcing one guarantees a violation). Two of the
broker-stage violations come from an agent voluntarily accepting a
price past its own private limit, the same failure mode narrow-gap
scenarios showed in experiments 2-3.

REPORT.md rewritten end to end as five numbered, chronological
experiments (original submission -> scenario expansion -> temperature
pinned -> MAX_TURNS=20 -> broker escalation), each with its exact
settings, results table, and its own mermaid flowchart, plus the
FIPA-ACL comparison table and a synthesis section tying all five
together. The pre-existing extra_broker/ subsidy-window exploration is
kept as an appendix to experiment 1, distinguished from experiment 5's
conversational broker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
…PORT.md

Syncs this machine's final week-04 round into main for pulling on
other machines: MAX_TURNS reverted to 5 (the graded setting), the
20-turn run archived, a new broker-escalation exploration (neutral
broker mediation, then an unconditional master broker as last resort),
and REPORT.md rewritten as five numbered chronological experiments
with per-experiment mermaid diagrams. Still not opened as a PR to
upstream.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KHyeVL2C6S9Ki2asRMo4q
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…20 ep)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…retation, add Jev experiments section

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er machine

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants