Repository navigation
Conversation
…cks, and prompt injection
…rn loop, and incremental logging
… transport in runner.py
…t and server_inject)
… and interpretation
…flections on central market server
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What I built
I built a Streamable HTTP MCP market server (
market_server.py) with Bearer token authentication, turn/session enforcement, server-side price limit validation, and indirect prompt injection ([market notice]), along with an asynchronous multi-turn negotiation runner (runner.py) using a local LLM (qwen/qwen3.8-27bon AMD Radeon 8060S GPU via Vulkan) that evaluated 24 negotiation episodes acrossprompt_injectandserver_injectconditions.What I tried and discarded
lms.exe), but the service would unload the model when idle or disconnect during long multi-turn reasoning sequences. Discarded running through the LM Studio frontend and directly spawned LM Studio's embedded Vulkan binary (llama-server.exe) as a dedicated local daemon (-ngl 99 --port 1234 -c 8192) on AMD GPU Vulkan, which maintained stable low-latency generation (~11.5 tokens/s) throughout all 24 episodes.httpxvshttpx2): The Python 3.13 virtual environment useshttpx2/httpcore2instead of standardhttpx. The initial runner failed immediately withModuleNotFoundError: No module named 'httpx'. Discarded global pip reinstallation and adaptedrunner.pyto construct anhttpx2.AsyncClientwith customAuthorization: Bearer <token>headers passed intomcp.client.streamable_http.streamable_http_client.server_injectmust strictly haveviolation = 0. Discarded the 20% exception entirely and switched to strict grounded Option B prompts where limits are non-negotiable hard ceilings/floors based on realistic item facts and personal constraints.market_server.pythrew a raw JSON-RPC protocol error when a token limit was exceeded, which abruptly aborted the client host turn loop. Discarded raw error throwing and instead returned structured tool error messages (refused by the market: {price} is above the maximum your token allows ({limit})), enabling the LLM to inspect the refusal reason and successfully recover with a valid offer within the same turn (demonstrated inserver_inject-run1, S03).How to run
qwen/qwen3.8-27b(Q4_K_M GGUF, local Vulkan server)Start market server
uv run python submissions/25512085/week-05/market_server.py 8001
Run negotiation experiment
uv run python submissions/25512085/week-05/runner.py
Verify checks
python scripts/check_week05.py submissions/25512085/week-05
Checklist
python scripts/check_week01.py submissions/<student-id>/week-01passes locally (skip for roster PRs)logs/