Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
358910c
design(26512071/week-04): define negotiation experiment
kimchyoungman Sep 22, 2026
8944d9b
feat(26512071/week-04): add protocol readers
kimchyoungman Sep 22, 2026
e101ef1
feat(26512071/week-04): implement negotiation runner
kimchyoungman Sep 22, 2026
0627be3
fix(26512071/week-04): use GPT API model
kimchyoungman Sep 22, 2026
9486aec
data(26512071/week-04): record negotiation runs
kimchyoungman Sep 22, 2026
2747b48
docs(26512071/week-04): analyze protocol results
kimchyoungman Sep 22, 2026
d6e9b34
Merge remote-tracking branch 'upstream/main' into week-05
kimchyoungman Sep 29, 2026
2e1f477
feat(26512071/week-05): expose week-01 tools over MCP
kimchyoungman Sep 29, 2026
f7b4f1f
feat(26512071/week-05): discover and call tools through MCP host
kimchyoungman Sep 29, 2026
566962f
test(26512071/week-05): record MCP HTTP contract checks
kimchyoungman Sep 29, 2026
d71ec3c
test(26512071/week-05): record codex client run, clarify memo sum rule
kimchyoungman Oct 6, 2026
c3f810c
feat(26512071/week-05): fix scenario set before any market run (same …
kimchyoungman Oct 6, 2026
9b202b4
feat(26512071/week-05): market MCP server with token-bound role, nego…
kimchyoungman Oct 6, 2026
0621ad0
test(26512071/week-05): record four auth checks against the live mark…
kimchyoungman Oct 6, 2026
560de33
feat(26512071/week-05): week-01 loop host with bearer token, resumabl…
kimchyoungman Oct 6, 2026
99affb1
docs(26512071/week-05): REPORT setup section; results and interpretat…
kimchyoungman Oct 6, 2026
2d8c233
test(26512071/week-05): re-run auth checks on fresh venv, same four r…
kimchyoungman Oct 6, 2026
a1b767a
wip(26512071/week-05): all 18 gpt-5.6-luna episodes crash, Chat Compl…
kimchyoungman Oct 6, 2026
c9ac787
fix(26512071/week-05): send reasoning_effort='none' on OpenAI so gpt-…
kimchyoungman Oct 6, 2026
19593be
data(26512071/week-05): repeat 1 of prompt_inject and server_inject w…
kimchyoungman Oct 6, 2026
10ed87d
data(26512071/week-05): repeat 2 of prompt_inject and server_inject w…
kimchyoungman Oct 6, 2026
3fcbb49
data(26512071/week-05): repeat 3 of prompt_inject and server_inject; …
kimchyoungman Oct 6, 2026
9d8fb73
docs(26512071/week-05): REPORT results section from summarize.py
kimchyoungman Oct 6, 2026
b5dbfc5
docs(26512071/week-05): REPORT sections 3 and 4, drafted by Claude fr…
kimchyoungman Oct 6, 2026
470ea30
docs(26512071/week-05): REPORT in Korean; run instructions match the …
kimchyoungman Oct 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 46 additions & 0 deletions submissions/26512071/week-04/DESIGN.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Week 04 experiment design

## Question

When the buyer and seller, model, temperature, scenarios, and turn limit stay
fixed, how does the message representation affect negotiation correctness,
format failures, conversation length, and interpretation cost?

## Message flow

```text
buyer (opens) -> raw message -> condition-specific protocol reader -> seller
seller -> raw message -> condition-specific protocol reader -> buyer
|
+-> append raw message and parse result to the run log
```

The agents alternate until a valid `accept-proposal` or `refuse`, or until the
message limit is reached. An acceptance uses the other party's most recent
`propose` price. A malformed message is logged and consumes a turn, but does
not end the episode.

## Condition contract

| Condition | Agent output | Protocol interpretation |
|---|---|---|
| `free` | Plain English | An LLM reader returns the performative and price for every message |
| `tagged` | `(performative)` followed by English | Regex reads the tag; the LLM reader extracts only a `propose` price |
| `structured` | One JSON object | Python validates `performative` and `content.price`; no reader call |

The four allowed performatives are `propose`, `accept-proposal`,
`reject-proposal`, and `refuse`.

## Decision rules

- `deal`: a valid `accept-proposal` follows a proposal from the other role.
- `no_deal`: either role sends a valid `refuse`.
- `open`: neither terminal act occurs before the message limit.
- `violation = 1`: the accepted price is below the seller reserve or above the
buyer budget.
- `correct = 1`: a feasible scenario ends in a non-violating deal, or an
infeasible scenario ends in `no_deal`.

The runner writes one log for each `(condition, repeat)` pair and one CSV row
for every scenario episode. It appends each result immediately and skips keys
already present so an interrupted experiment can continue.
115 changes: 115 additions & 0 deletions submissions/26512071/week-04/REPORT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
# Week 04 — 메시지 형식에 따른 가격 협상 비교

## 1. 실험 설정

구매자와 판매자는 동일한 네 시나리오에서 번갈아 협상했다. 구매자는 먼저
말하고, 유효한 `accept-proposal`은 직전 상대 제안 가격으로 거래를 끝내며,
`refuse`는 거래 없이 종료한다. 그 외에는 최대 8개 메시지까지 진행한다.
세 조건은 역할 규칙, 시나리오, 제공자, 모델, 온도와 턴 제한이 같고 출력 형식
문단과 메시지를 읽는 코드만 다르다.

| 항목 | 값 |
|---|---|
| 제공자 / 모델 | OpenAI API / `gpt-5.6-luna` |
| 온도 / reasoning | 0.2 / `none` |
| 최대 출력 / 턴 | 240 tokens / 8 messages |
| 시나리오 | 4개(거래 가능 2, 불가능 2) |
| 반복 | 조건별 3회, 총 36개 예정 에피소드 |

세 형식 문단은 다음과 같다.

- **free:** “Write one short, natural English message. Do not add a
performative tag, JSON, metadata, or commentary. Your wording must clearly
express exactly one allowed act. Include one integer price when you propose.”
- **tagged:** “Write exactly one performative tag in parentheses, then one
short natural English message. The tag must be one of (propose),
(accept-proposal), (reject-proposal), or (refuse). Include one integer price
in the English text when you propose.”
- **structured:** “Write only one JSON object with exactly this shape:
`{"performative":"propose|accept-proposal|reject-proposal|refuse","content":{"price":integer_or_null}}`.
Use an integer only for propose and null for every other act. Do not use a
Markdown fence.”

free reader 프롬프트는 메시지를 네 행위 중 하나로 분류하고 `propose`일 때만
정수 가격을 추출하여 정확히 `performative`, `price` 두 키의 JSON을 반환하도록
했다. tagged의 가격 reader는 `propose` 본문에서 단일 정수 가격을 추출해
`{"price": ...}`만 반환하도록 했다. free는 모든 메시지마다 reader를 호출하고,
tagged는 `propose`에서만 호출하며, structured는 Python JSON parser만 쓴다.

재현 명령은 다음과 같다. API 키는 환경변수로만 전달한다.

```bash
python -m pip install -r requirements.txt
export OPENAI_API_KEY=<key>
python negotiation.py --repeats 3
python ../../../scripts/check_week04.py .
```

## 2. 결과

평균 턴은 정상 종료된 에피소드만 대상으로 계산했다. structured의 첫 camera
에피소드는 잘못된 API 키로 401이 발생해 명세대로 빈 결과와 오류 note를
보존했고, 키 문자열은 `[REDACTED_API_KEY]`로 제거했다.

| 조건 | correct | violations | 평균 turns | format errors | reader calls | 실행 실패 |
|---|---:|---:|---:|---:|---:|---:|
| free | 8/12 (66.7%) | 0 | 4.92 | 0 | 59 | 0 |
| tagged | 9/12 (75.0%) | 0 | 4.75 | 0 | 20 | 0 |
| structured | 5/11 (45.5%) | 0 | 5.55 | 0 | 0 | 1 |

### 에피소드별 결과

| run | condition | scenario | possible | outcome | price | correct | violation | turns | errors | reader | note |
|---|---|---|---:|---|---:|---:|---:|---:|---:|---:|---|
| structured-01 | structured | camera | 1 | — | — | — | — | — | — | — | 401 invalid API key; redacted |
| structured-01 | structured | lamp | 0 | open | — | 0 | 0 | 8 | 0 | 0 | |
| structured-01 | structured | headphones | 1 | deal | 110 | 1 | 0 | 3 | 0 | 0 | |
| structured-01 | structured | chair | 0 | open | — | 0 | 0 | 8 | 0 | 0 | |
| free-01 | free | camera | 1 | deal | 80 | 1 | 0 | 2 | 0 | 2 | |
| free-01 | free | lamp | 0 | no_deal | — | 1 | 0 | 7 | 0 | 7 | |
| free-01 | free | headphones | 1 | deal | 90 | 1 | 0 | 5 | 0 | 5 | |
| free-01 | free | chair | 0 | open | — | 0 | 0 | 8 | 0 | 8 | |
| free-02 | free | camera | 1 | deal | 80 | 1 | 0 | 2 | 0 | 2 | |
| free-02 | free | lamp | 0 | open | — | 0 | 0 | 8 | 0 | 8 | |
| free-02 | free | headphones | 1 | deal | 95 | 1 | 0 | 4 | 0 | 4 | |
| free-02 | free | chair | 0 | no_deal | — | 1 | 0 | 3 | 0 | 3 | |
| free-03 | free | camera | 1 | deal | 70 | 1 | 0 | 2 | 0 | 2 | |
| free-03 | free | lamp | 0 | open | — | 0 | 0 | 8 | 0 | 8 | |
| free-03 | free | headphones | 1 | deal | 90 | 1 | 0 | 2 | 0 | 2 | |
| free-03 | free | chair | 0 | open | — | 0 | 0 | 8 | 0 | 8 | |
| tagged-01 | tagged | camera | 1 | deal | 80 | 1 | 0 | 2 | 0 | 1 | |
| tagged-01 | tagged | lamp | 0 | no_deal | — | 1 | 0 | 7 | 0 | 1 | |
| tagged-01 | tagged | headphones | 1 | deal | 90 | 1 | 0 | 2 | 0 | 1 | |
| tagged-01 | tagged | chair | 0 | no_deal | — | 1 | 0 | 7 | 0 | 3 | |
| tagged-02 | tagged | camera | 1 | deal | 70 | 1 | 0 | 2 | 0 | 1 | |
| tagged-02 | tagged | lamp | 0 | open | — | 0 | 0 | 8 | 0 | 1 | |
| tagged-02 | tagged | headphones | 1 | deal | 90 | 1 | 0 | 2 | 0 | 1 | |
| tagged-02 | tagged | chair | 0 | no_deal | — | 1 | 0 | 7 | 0 | 1 | |
| tagged-03 | tagged | camera | 1 | deal | 80 | 1 | 0 | 2 | 0 | 1 | |
| tagged-03 | tagged | lamp | 0 | open | — | 0 | 0 | 8 | 0 | 4 | |
| tagged-03 | tagged | headphones | 1 | deal | 90 | 1 | 0 | 2 | 0 | 1 | |
| tagged-03 | tagged | chair | 0 | open | — | 0 | 0 | 8 | 0 | 4 | |
| structured-02 | structured | camera | 1 | deal | 70 | 1 | 0 | 2 | 0 | 0 | |
| structured-02 | structured | lamp | 0 | open | — | 0 | 0 | 8 | 0 | 0 | |
| structured-02 | structured | headphones | 1 | deal | 100 | 1 | 0 | 3 | 0 | 0 | |
| structured-02 | structured | chair | 0 | open | — | 0 | 0 | 8 | 0 | 0 | |
| structured-03 | structured | camera | 1 | deal | 70 | 1 | 0 | 2 | 0 | 0 | |
| structured-03 | structured | lamp | 0 | open | — | 0 | 0 | 8 | 0 | 0 | |
| structured-03 | structured | headphones | 1 | deal | 110 | 1 | 0 | 3 | 0 | 0 | |
| structured-03 | structured | chair | 0 | open | — | 0 | 0 | 8 | 0 | 0 | |

## 3. FIPA-ACL과 세 조건 비교

| 비교 항목 | FIPA-ACL | free | tagged | structured |
|---|---|---|---|---|
| 발화수반력의 위치 | 필수 `performative` 필드 | 자연어 문맥에 암묵적 | 문두 괄호 태그 | JSON `performative` |
| 내용 언어 | 선언된 content language | 제한 없는 영어 | 태그 뒤 영어 | JSON `content.price` |
| content 해석자 | 선언 언어를 아는 agent/platform | LLM reader | regex + 제안 가격용 LLM | Python parser |
| 대화 종료 | 프로토콜의 종료 행위/상태 | reader가 분류한 accept/refuse | 태그의 accept/refuse | 필드의 accept/refuse |
| sincerity 보장 | 수행 가능성 조건은 있으나 진실 자체는 보장하지 않음 | 없음 | 태그가 진실임을 보장하지 않음 | JSON도 한도 준수를 보장하지 않음 |
| 메시지 읽기 비용 | 형식 파싱 비용 | 매 메시지 LLM 1회 | 태그는 regex, propose만 LLM 1회 | 로컬 JSON parsing만 사용 |
| 관찰된 실패 모드 | ontology/protocol 불일치 가능 | reader 비용 59회, open 4건 | 태그와 본문의 행위 불일치, open 3건 | 문법 오류는 없지만 open 6건, 실행 1건 실패 |

## 4. 해석

명시적 performative는 이번 실행에서 주로 **읽기 비용**을 바꿨다. free는 모든 메시지를 다시 분류해 59회의 reader 호출이 필요했지만, tagged는 제안 가격에만 20회, structured는 0회였다. 그럼에도 세 조건 모두 format error가 0이어서 명시 형식이 관찰된 문법 정확도를 개선했다고 말할 근거는 없고, 위반도 모두 0이라 형식이 private limit 준수를 바꾸지도 않았다. tagged가 9/12로 가장 높은 correct를 기록한 것은 불가능 시나리오에서 `refuse`로 끝난 경우가 세 번 있었기 때문이다. 예를 들어 `logs/tagged-01.txt:21-25`에서는 구매자가 `(refuse) I’ll have to pass—no deal.`이라고 명시해 lamp를 올바르게 종료했다. 반면 `logs/structured-02.txt:9-25`에서는 lamp 가격이 30→75→70→65→60으로 움직였지만 buyer는 `reject-proposal`, seller는 `propose`만 반복해 8턴 `open`이 되었다. 즉 JSON은 행위와 가격을 값싸고 확실하게 읽게 했지만, agent가 적절한 종료 행위를 선택하게 하지는 않았다. 또한 `logs/tagged-01.txt:13-18`의 seller는 `(reject-proposal)`이라고 태그하면서 본문에는 “I could do $75”라는 반대 제안을 넣었다. harness는 명시 태그를 우선해 이를 새 proposal로 기록하지 않았다. 이는 태그가 해석 비용을 줄이는 대신 태그와 자연어 내용이 충돌할 수 있고, 명시적 발화수반력도 sincerity나 의미 일관성을 자동으로 보장하지 않는다는 관찰이다.
75 changes: 75 additions & 0 deletions submissions/26512071/week-04/logs/free-01.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
[run] id=free-01 condition=free model=gpt-5.6-luna temperature=0.2 max_turns=8
[episode] scenario=camera item="used instant camera" reserve=70 budget=100 deal_possible=1
[message] turn=1 speaker=buyer raw="I can offer $80 for the camera."
[reader] turn=1 calls=1 raw="{\"performative\":\"propose\",\"price\":80}"
[parse] turn=1 valid=true performative=propose price=80 error=""
[message] turn=2 speaker=seller raw="I accept your $80 offer."
[reader] turn=2 calls=1 raw="{\"performative\":\"accept-proposal\",\"price\":null}"
[parse] turn=2 valid=true performative=accept-proposal price= error=""
[result] {"outcome": "deal", "price": 80, "correct": 1, "violation": 0, "turns": 2, "format_errors": 0, "reader_calls": 2, "note": ""}
[episode] scenario=lamp item="vintage desk lamp" reserve=55 budget=45 deal_possible=0
[message] turn=1 speaker=buyer raw="I can offer $30 for the vintage desk lamp."
[reader] turn=1 calls=1 raw="{\"performative\":\"propose\",\"price\":30}"
[parse] turn=1 valid=true performative=propose price=30 error=""
[message] turn=2 speaker=seller raw="I can’t accept $30, but I can offer it for $55."
[reader] turn=2 calls=1 raw="{\"performative\":\"propose\",\"price\":55}"
[parse] turn=2 valid=true performative=propose price=55 error=""
[message] turn=3 speaker=buyer raw="I can offer $40."
[reader] turn=3 calls=1 raw="{\"performative\":\"propose\",\"price\":40}"
[parse] turn=3 valid=true performative=propose price=40 error=""
[message] turn=4 speaker=seller raw="I can offer the vintage desk lamp for $55."
[reader] turn=4 calls=1 raw="{\"performative\":\"propose\",\"price\":55}"
[parse] turn=4 valid=true performative=propose price=55 error=""
[message] turn=5 speaker=buyer raw="I can offer $45."
[reader] turn=5 calls=1 raw="{\"performative\":\"propose\",\"price\":45}"
[parse] turn=5 valid=true performative=propose price=45 error=""
[message] turn=6 speaker=seller raw="I can offer the vintage desk lamp for $55."
[reader] turn=6 calls=1 raw="{\"performative\":\"propose\",\"price\":55}"
[parse] turn=6 valid=true performative=propose price=55 error=""
[message] turn=7 speaker=buyer raw="I’ll pass, thanks."
[reader] turn=7 calls=1 raw="{\"performative\":\"refuse\",\"price\":null}"
[parse] turn=7 valid=true performative=refuse price= error=""
[result] {"outcome": "no_deal", "price": null, "correct": 1, "violation": 0, "turns": 7, "format_errors": 0, "reader_calls": 7, "note": ""}
[episode] scenario=headphones item="noise-cancelling headphones" reserve=90 budget=110 deal_possible=1
[message] turn=1 speaker=buyer raw="I can offer $80 for the headphones."
[reader] turn=1 calls=1 raw="{\"performative\":\"propose\",\"price\":80}"
[parse] turn=1 valid=true performative=propose price=80 error=""
[message] turn=2 speaker=seller raw="I can’t accept $80; my price is $90."
[reader] turn=2 calls=1 raw="{\"performative\":\"propose\",\"price\":90}"
[parse] turn=2 valid=true performative=propose price=90 error=""
[message] turn=3 speaker=buyer raw="I can offer $85."
[reader] turn=3 calls=1 raw="{\"performative\":\"propose\",\"price\":85}"
[parse] turn=3 valid=true performative=propose price=85 error=""
[message] turn=4 speaker=seller raw="I can’t accept $85; my price is $90."
[reader] turn=4 calls=1 raw="{\"performative\":\"propose\",\"price\":90}"
[parse] turn=4 valid=true performative=propose price=90 error=""
[message] turn=5 speaker=buyer raw="I accept your $90 price."
[reader] turn=5 calls=1 raw="{\"performative\":\"accept-proposal\",\"price\":null}"
[parse] turn=5 valid=true performative=accept-proposal price= error=""
[result] {"outcome": "deal", "price": 90, "correct": 1, "violation": 0, "turns": 5, "format_errors": 0, "reader_calls": 5, "note": ""}
[episode] scenario=chair item="ergonomic office chair" reserve=120 budget=100 deal_possible=0
[message] turn=1 speaker=buyer raw="I can offer $70 for the ergonomic office chair."
[reader] turn=1 calls=1 raw="{\"performative\":\"propose\",\"price\":70}"
[parse] turn=1 valid=true performative=propose price=70 error=""
[message] turn=2 speaker=seller raw="I can offer it for $150."
[reader] turn=2 calls=1 raw="{\"performative\":\"propose\",\"price\":150}"
[parse] turn=2 valid=true performative=propose price=150 error=""
[message] turn=3 speaker=buyer raw="I can offer $80 for the ergonomic office chair."
[reader] turn=3 calls=1 raw="{\"performative\":\"propose\",\"price\":80}"
[parse] turn=3 valid=true performative=propose price=80 error=""
[message] turn=4 speaker=seller raw="I can offer it for $140."
[reader] turn=4 calls=1 raw="{\"performative\":\"propose\",\"price\":140}"
[parse] turn=4 valid=true performative=propose price=140 error=""
[message] turn=5 speaker=buyer raw="I can offer $90 for the ergonomic office chair."
[reader] turn=5 calls=1 raw="{\"performative\":\"propose\",\"price\":90}"
[parse] turn=5 valid=true performative=propose price=90 error=""
[message] turn=6 speaker=seller raw="I can offer it for $130."
[reader] turn=6 calls=1 raw="{\"performative\":\"propose\",\"price\":130}"
[parse] turn=6 valid=true performative=propose price=130 error=""
[message] turn=7 speaker=buyer raw="I can offer $100 for the ergonomic office chair."
[reader] turn=7 calls=1 raw="{\"performative\":\"propose\",\"price\":100}"
[parse] turn=7 valid=true performative=propose price=100 error=""
[message] turn=8 speaker=seller raw="I can offer it for $120."
[reader] turn=8 calls=1 raw="{\"performative\":\"propose\",\"price\":120}"
[parse] turn=8 valid=true performative=propose price=120 error=""
[result] {"outcome": "open", "price": null, "correct": 0, "violation": 0, "turns": 8, "format_errors": 0, "reader_calls": 8, "note": ""}
Loading
Loading