Repository navigation
feat(scripts): hackathon race monitor - #266
cardosofede wants to merge 5 commits into
Conversation
…ume each minute
scripts/hackathon_monitor.py reads the race's agent ids (strategy slugs) from
the hackathon API, measures each one across every configured Condor server
with the same aggregator the dashboard uses (bots under the strategy's
namespace, stopped instances included, plus the standalone executors its
sessions tagged), and POSTs { timestamp, agents: [{ agent_id, pnl_quote,
volume_quote }] } once a minute with a bearer token.
--start-race records a baseline that every later post subtracts, so the race
counts from its start. A server that stops answering keeps contributing its
last figures, persisted in the state file so a restart does not lose them; a
server never heard from counts as nothing and blocks taking the baseline.
|
| totals[agent_id] = Totals( | ||
| pnl=sum(perf[tag].total_pnl for tag in owner_tags if tag in perf), | ||
| volume=sum(perf[tag].volume for tag in owner_tags if tag in perf), |
There was a problem hiding this comment.
Mixed currencies distort race totals. If an executor trades in a non-USD quote currency, its PnL and volume remain in that currency while bot figures are restated in USD. Summing them here posts a mixed-currency score; a strategy with only non-USD executors is never converted at all.
Knowledge Base Used: Portfolio performance and rates
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| if not args.url: | ||
| parser.error("set HACKATHON_API_URL or pass --url") |
There was a problem hiding this comment.
Bearer token crosses plaintext connection. If
--url or HACKATHON_API_URL uses http://, this check accepts it and the monitor sends its bearer token with both GET and POST requests. An on-path observer can then read the credential. How this was verified: The configured URL has no scheme check and is used by the session that carries the bearer token.
| if self.args.start_race: | ||
| if unknown: | ||
| raise RuntimeError( | ||
| "Cannot take the baseline: no figure from " | ||
| + ", ".join(sorted(unknown)) | ||
| + ". Fix the server or leave it out with --servers." | ||
| ) | ||
| save_baseline(self.state_path, totals, now) |
There was a problem hiding this comment.
Stale figures become race baseline. If a server was read previously but fails during
--start-race, combine substitutes its last figure without marking it unknown. This guard then saves that stale figure as the baseline, so activity between the last successful read and the race start is counted as race activity when the server recovers.
| return { | ||
| agent_id: t.minus(baseline.get(agent_id, Totals())) | ||
| for agent_id, t in totals.items() |
There was a problem hiding this comment.
| names = distinct_servers(get_config_manager().list_servers(), self.args.servers) | ||
| answers = await asyncio.gather( | ||
| *(self._read_server(name, matched) for name in names) | ||
| ) | ||
| totals, unknown = combine( | ||
| self.agent_ids, dict(zip(names, answers)), self.last_good | ||
| ) |
| for agent_id, owner_tags in tags.items(): | ||
| totals[agent_id] = Totals( | ||
| pnl=sum(perf[tag].total_pnl for tag in owner_tags if tag in perf), | ||
| volume=sum(perf[tag].volume for tag in owner_tags if tag in perf), | ||
| ) | ||
| if failed.intersection(owner_tags): | ||
| degraded.add(agent_id) |
There was a problem hiding this comment.
Missing bots overwrite reliable totals. If a previously measured bot has no discoverable live or archived instance, the aggregator marks its base unresolved but does not add its tag to
failed_ids. This code treats the incomplete total as clean, overwrites the last reliable figure, and posts an apparent drop in PnL or volume.
Knowledge Base Used:
| owner_tags = list(owner.agent_ids) or [owner.run_key] | ||
| tags[agent_id].extend(owner_tags) | ||
| bases[owner_tags[0]] = [owner.namespace, *owner.declared_bots] |
There was a problem hiding this comment.
Session tags overwrite bot ownership. Strategies named
grid and grid_1 can produce the same key when grid has session alice.grid_1 and grid_1 has no sessions. The latter's run-key fallback overwrites the former's bot-base entry, so bot PnL and volume can be omitted or assigned to the wrong race ID without a degradation signal.
| known = last_good.setdefault(server, {}) | ||
| for agent_id in agent_ids: | ||
| if agent_id in fresh and agent_id not in degraded: | ||
| known[agent_id] = fresh[agent_id] | ||
| if agent_id in known: | ||
| totals[agent_id] = totals[agent_id].plus(known[agent_id]) |
There was a problem hiding this comment.
…ST result
The site wraps its answers in {"data": ...}. parse_agent_ids looked for a
top-level list and returned [] for the real response, so the monitor
stopped with "listed no agents". It now reads data.agents and raises on any
other shape. The POST result is read too: the accepted count is logged and
agent ids the site does not know are warned about.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| try: | ||
| return [agent["agent_id"] for agent in payload["data"]["agents"]] |
There was a problem hiding this comment.
Supported rosters now fail. If the race API returns a bare list or a list under
agents, agent_ids, or data, as the PR description says it may, this parser rejects it. The first tick exits without posting; if the response shape changes after a successful post, the monitor keeps the old roster and misses new entrants.
| session: Any, url: str, payload: dict[str, Any] | ||
| ) -> tuple[int, list[str]]: | ||
| async with session.post(url, json=payload) as response: | ||
| if response.status >= 400: |
There was a problem hiding this comment.
Successful posts stop monitoring. If the race API accepts a POST but returns no JSON body, or omits either
data.accepted or data.unknown, the new response parser raises an error. Because posted is set only after parsing succeeds, this exits the monitor on its first tick even though the figures may have been accepted, preventing all later updates.
The leaderboard names agents by botcamp strategy slug, which rarely equals the Condor strategy slug, so every agent read zero. HACKATHON_AGENT_MAP (or --agent-map) points at a YAML of race id -> run keys, re-read every minute; an id in it is measured by those keys only, any other id by slug as before. scripts/hackathon_agents.example.yml lists the 14 agents of agent-builders-cup-1 with the keys their submissions declare. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| if agent_id in agent_map: | ||
| keys = agent_map[agent_id] | ||
| matched[agent_id] = [owner for owner in owners if owner.run_key in keys] | ||
| else: | ||
| matched[agent_id] = [ | ||
| owner | ||
| for owner in owners | ||
| if agent_id in (owner.strategy_slug, owner.run_key) | ||
| ] |
There was a problem hiding this comment.
| from condor.agents.fleet_map import build_fleet_map | ||
| from config_manager import get_config_manager | ||
|
|
||
| agent_map = load_agent_map(self.args.agent_map) if self.args.agent_map else {} |
There was a problem hiding this comment.
Map edits corrupt race totals. If an agent’s map is edited after
--start-race, the next tick measures the newly assigned strategy but subtracts the baseline saved for the old one. If a server is down, it can also reuse the old strategy’s last-good figure. Correcting an assignment mid-race therefore posts inaccurate PnL and volume.
… agent map The builder resubmitted as a Condor agent; the race now lists the new strategy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s and closed executors The dock's agent rows folded only the running population, so an agent's PnL dropped as soon as it stopped a bot or an executor closed. Each agent row now folds its live controllers plus its finished records (stopped bots' final controller snapshots, realized only, and closed executors), with one History row under it so the rows still add up. The live half keeps updating over the socket. The terminated-leaf construction moves from PerfBrowser to lib/perf-population (terminatedLeaves) so /bots and the dock share it, and useFleetData gains a `terminated` option to load runs and finished controllers alongside the live fleet.
| terminatedLeaves({ | ||
| executors: fleet.executors, | ||
| terminatedControllers: fleet.terminatedControllers, | ||
| runs: fleet.runs, | ||
| owners: fleet.owners, | ||
| deeds: fleet.deeds, | ||
| botByController: botsByController(fleet.controllers), | ||
| }), |
There was a problem hiding this comment.
Stopped bot totals misconverted. If a stopped bot trades in a non-USD quote currency that no live controller or loaded executor uses, the dock includes that bot in the agent’s history but does not request its exchange rate. Conversion leaves its PnL and volume unchanged while displaying them in the selected currency, so the agent’s lifetime totals are wrong.
Knowledge Base Used: Frontend application
| onOpen={() => | ||
| navigate( | ||
| `/bots?population=terminated&scope=${encodeURIComponent(row.parentId ?? "")}`, | ||
| ) |
There was a problem hiding this comment.
History link omits older activity. The history row shows an agent’s lifetime total, but clicking it opens the terminated browser with its default three-month window. For agents with older stopped bots or closed executors, the detail view omits figures included in the row, making the total difficult to verify.
Knowledge Base Used: Frontend application
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
@david-hummingbot this is the script that feeds the Agent Builders Cup race board.
Summary
scripts/hackathon_monitor.pyreports every hackathon agent's cumulative PnL and volume to the race API once a minute./api/hackathons/agent-builders-cup-1/race-datareturns the agent ids, which are strategy slugs.config.ymland summed. It usesfetch_agent_performance_batch, the aggregator behind the dashboard: bots under the{agent}-{strategy}namespace (stopped instances included) plus the standalone executors its sessions tagged. Non-USD quotes are restated in USD.{ timestamp, agents: [{ agent_id, pnl_quote, volume_quote }] }withAuthorization: Bearer <token>.Running it
From the repo root of the checkout that holds
.condor/agentsandconfig.yml:Env (names in
.env.example):HACKATHON_API_URL,HACKATHON_API_TOKEN, optionalHACKATHON_SLUG, and optionalHACKATHON_SERVERSto restrict the server set.Behaviour worth reviewing
--start-racesaves each agent's figures at that moment, and every later post subtracts them, so earlier test runs don't count. It refuses to run while any server has never answered; restrict with--serversin that case..condor/hackathon/<slug>.jsonso they survive a restart. A server never reached counts as nothing and is named in the log every minute.Not verified yet
agents,agent_idsordata. The timestamp is ISO-8601 UTC (2026-10-02T16:39:22Z).Test plan
tests/test_hackathon_monitor.py: 18 testslocalandbrigado_2(moneymakerandcornellunreachable from here): bearer header, lifetime post, baseline refusal, baseline, restartbrigado_2, live and stopped, matched a by-hand sum of the raw snapshots, after the USD conversion