Skip to content

[week-05] 25512085 - #261

Open
MJforge wants to merge 15 commits into
Q00:mainfrom
MJforge:week-05
Open

MJforge wants to merge 15 commits into
Q00:mainfrom
MJforge:week-05

Conversation

@MJforge

@MJforge MJforge commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

What I built

I built a Streamable HTTP MCP market server (market_server.py) with Bearer token authentication, turn/session enforcement, server-side price limit validation, and indirect prompt injection ([market notice]), along with an asynchronous multi-turn negotiation runner (runner.py) using a local LLM (qwen/qwen3.8-27b on AMD Radeon 8060S GPU via Vulkan) that evaluated 24 negotiation episodes across prompt_inject and server_inject conditions.

What I tried and discarded

  • LM Studio GUI/CLI daemon instability: Initially attempted to drive the model using LM Studio CLI (lms.exe), but the service would unload the model when idle or disconnect during long multi-turn reasoning sequences. Discarded running through the LM Studio frontend and directly spawned LM Studio's embedded Vulkan binary (llama-server.exe) as a dedicated local daemon (-ngl 99 --port 1234 -c 8192) on AMD GPU Vulkan, which maintained stable low-latency generation (~11.5 tokens/s) throughout all 24 episodes.
  • Python package import incompatibility (httpx vs httpx2): The Python 3.13 virtual environment uses httpx2 / httpcore2 instead of standard httpx. The initial runner failed immediately with ModuleNotFoundError: No module named 'httpx'. Discarded global pip reinstallation and adapted runner.py to construct an httpx2.AsyncClient with custom Authorization: Bearer <token> headers passed into mcp.client.streamable_http.streamable_http_client.
  • Discarding Week 04's 20% exception discretion: In Week 04, agents were given a 20% discretion rule for tough negotiations. For Week 05, this would have caused false-positive violations and violated the CI invariant that server_inject must strictly have violation = 0. Discarded the 20% exception entirely and switched to strict grounded Option B prompts where limits are non-negotiable hard ceilings/floors based on realistic item facts and personal constraints.
  • Unhandled tool exceptions on limit refusal: An early version of market_server.py threw a raw JSON-RPC protocol error when a token limit was exceeded, which abruptly aborted the client host turn loop. Discarded raw error throwing and instead returned structured tool error messages (refused by the market: {price} is above the maximum your token allows ({limit})), enabling the LLM to inspect the refusal reason and successfully recover with a valid offer within the same turn (demonstrated in server_inject-run1, S03).

How to run

  • Model: qwen/qwen3.8-27b (Q4_K_M GGUF, local Vulkan server)
  • Env vars:
    export OPENAI_BASE_URL=http://127.0.0.1:1234/v1
    export OPENAI_API_KEY=not-needed
    export AGENT_MODEL=qwen/qwen3.8-27b
    
  1. Start market server
    uv run python submissions/25512085/week-05/market_server.py 8001

  2. Run negotiation experiment
    uv run python submissions/25512085/week-05/runner.py

  3. Verify checks
    python scripts/check_week05.py submissions/25512085/week-05

Checklist

  • python scripts/check_week01.py submissions/<student-id>/week-01 passes locally (skip for roster PRs)
  • Run logs are committed under logs/
  • No API keys anywhere in the diff
  • History is not squashed

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant