Skip to content

feat(scripts): hackathon race monitor - #266

Open
cardosofede wants to merge 5 commits into
mainfrom
feat/hackathon-race-monitor
Open

cardosofede wants to merge 5 commits into
mainfrom
feat/hackathon-race-monitor

Conversation

@cardosofede

Copy link
Copy Markdown
Contributor

@david-hummingbot this is the script that feeds the Agent Builders Cup race board.

Summary

scripts/hackathon_monitor.py reports every hackathon agent's cumulative PnL and volume to the race API once a minute.

  1. GET /api/hackathons/agent-builders-cup-1/race-data returns the agent ids, which are strategy slugs.
  2. Each id is matched to the Condor strategy (or strategies) with that slug.
  3. Each strategy is measured on every server in config.yml and summed. It uses fetch_agent_performance_batch, the aggregator behind the dashboard: bots under the {agent}-{strategy} namespace (stopped instances included) plus the standalone executors its sessions tagged. Non-USD quotes are restated in USD.
  4. POST { timestamp, agents: [{ agent_id, pnl_quote, volume_quote }] } with Authorization: Bearer <token>.

Running it

From the repo root of the checkout that holds .condor/agents and config.yml:

uv run python scripts/hackathon_monitor.py --dry-run --once   # print the payload, post nothing
uv run python scripts/hackathon_monitor.py --start-race       # at the start: record the baseline, keep posting
uv run python scripts/hackathon_monitor.py                    # after a restart: reuses the saved baseline

Env (names in .env.example): HACKATHON_API_URL, HACKATHON_API_TOKEN, optional HACKATHON_SLUG, and optional HACKATHON_SERVERS to restrict the server set.

Behaviour worth reviewing

  • Race start: --start-race saves each agent's figures at that moment, and every later post subtracts them, so earlier test runs don't count. It refuses to run while any server has never answered; restrict with --servers in that case.
  • Server outages: a server that stops answering keeps contributing its last figures, saved in .condor/hackathon/<slug>.json so they survive a restart. A server never reached counts as nothing and is named in the log every minute.
  • Duplicate servers: two server names pointing at the same host and port are read once, so nothing is double counted.
  • Stopping never takes anything off the board: a paused controller in a running bot, a stopped and archived bot (its realized PnL; its last open-position value is not counted), and closed or open standalone executors all stay in the agent's total. This is pinned by a test that goes through the real aggregator.
  • Unknown ids: an id with no matching strategy posts 0 / 0 and is logged.

Not verified yet

  • Not yet run against the real race API (no URL or token on hand). The GET response is read as a bare list, or a list under agents, agent_ids or data. The timestamp is ISO-8601 UTC (2026-10-02T16:39:22Z).
  • The dashboard's Running view and the agent's Execution dock still drop stopped bots and closed executors from an agent's on-screen number; this PR does not change that.

Test plan

  • tests/test_hackathon_monitor.py: 18 tests
  • Full suite: 6424 passed, 23 skipped (21 of the skips are the web-route tests that need a frontend build, which this worktree does not have)
  • End to end against a local stand-in for the race API, with real reads from local and brigado_2 (moneymaker and cornell unreachable from here): bearer header, lifetime post, baseline refusal, baseline, restart
  • Totals for real bot groups on brigado_2, live and stopped, matched a by-hand sum of the raw snapshots, after the USD conversion
  • Dry run against the real race API once the URL and token are available

…ume each minute

scripts/hackathon_monitor.py reads the race's agent ids (strategy slugs) from
the hackathon API, measures each one across every configured Condor server
with the same aggregator the dashboard uses (bots under the strategy's
namespace, stopped instances included, plus the standalone executors its
sessions tagged), and POSTs { timestamp, agents: [{ agent_id, pnl_quote,
volume_quote }] } once a minute with a bearer token.

--start-race records a baseline that every later post subtracts, so the race
counts from its start. A server that stops answering keeps contributing its
last figures, persisted in the state file so a restart does not lose them; a
server never heard from counts as nothing and blocks taking the baseline.
@greptile-apps

greptile-apps Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 0/5

[Medium risk] Adds a hackathon race monitor script and supporting UI changes.

The PR is not safe to merge while the race monitor’s outstanding scoring defects and the dock’s incorrect converted totals remain.

Findings

  1. P1 Stopped Bot Totals Misconverted ▶
  2. P1 Mixed currencies distort race totals. ▶
  3. P1 Security Bearer token crosses plaintext connection. ▶
  4. P1 Stale figures become race baseline. ▶
  5. P1 Late entrants inherit zero baseline. ▶
  6. P1 Empty server list posts zeros. ▶
  7. P1 Missing bots overwrite reliable totals. ▶
  8. P1 Session tags overwrite bot ownership. ▶
  9. P1 Server changes reuse old figures. ▶
  10. P1 Supported rosters now fail. ▶
  11. P1 Successful posts stop monitoring. ▶
  12. P1 One strategy counted twice. ▶
  13. P1 Map edits corrupt race totals. ▶
  14. P2 History Link Omits Older Activity ▶

Summary

The PR adds a race monitor that collects strategy performance across configured servers, persists a race-start baseline and last-good server figures, and posts periodic totals. It also adds agent-mapping examples and extends the execution dock to include stopped bots and closed executors in agent totals. The new dock history can use incomplete currency rates, and its detail link defaults to a shorter period than the total it displays.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Servers[Configured servers] --> Monitor[Race monitor]
  Monitor --> State[Baseline and last-good state]
  Monitor --> API[Race API]
  Fleet[Live fleet and terminated records] --> Dock[Execution dock]
  Dock --> Browser[Terminated performance browser]
Loading

Reviews (5) · Last reviewed commit: "feat(web): the execution dock credits ea..."

Comment on lines +196 to +198
totals[agent_id] = Totals(
pnl=sum(perf[tag].total_pnl for tag in owner_tags if tag in perf),
volume=sum(perf[tag].volume for tag in owner_tags if tag in perf),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Mixed currencies distort race totals. If an executor trades in a non-USD quote currency, its PnL and volume remain in that currency while bot figures are restated in USD. Summing them here posts a mixed-currency score; a strategy with only non-USD executors is never converted at all.

Knowledge Base Used: Portfolio performance and rates

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +469 to +470
if not args.url:
parser.error("set HACKATHON_API_URL or pass --url")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Bearer token crosses plaintext connection. If --url or HACKATHON_API_URL uses http://, this check accepts it and the monitor sends its bearer token with both GET and POST requests. An on-path observer can then read the credential. How this was verified: The configured URL has no scheme check and is used by the session that carries the bearer token.

Comment on lines +396 to +403
if self.args.start_race:
if unknown:
raise RuntimeError(
"Cannot take the baseline: no figure from "
+ ", ".join(sorted(unknown))
+ ". Fix the server or leave it out with --servers."
)
save_baseline(self.state_path, totals, now)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Stale figures become race baseline. If a server was read previously but fails during --start-race, combine substitutes its last figure without marking it unknown. This guard then saves that stale figure as the baseline, so activity between the last successful read and the race start is counted as race activity when the server recovers.

Comment on lines +315 to +317
return {
agent_id: t.minus(baseline.get(agent_id, Totals()))
for agent_id, t in totals.items()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Late entrants inherit zero baseline. If the race API adds an agent ID after the baseline was saved, this code assumes its starting PnL and volume were zero. When the newly listed strategy has pre-race trading history, its full lifetime results are posted as race activity.

Comment on lines +376 to +382
names = distinct_servers(get_config_manager().list_servers(), self.args.servers)
answers = await asyncio.gather(
*(self._read_server(name, matched) for name in names)
)
totals, unknown = combine(
self.agent_ids, dict(zip(names, answers)), self.last_good
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Empty server list posts zeros. If no servers are configured, this gathers no answers, yet combine returns zero totals with no unknown figures. A normal run then posts zero PnL and volume for every agent each minute despite having read no trading data.

Comment on lines +195 to +201
for agent_id, owner_tags in tags.items():
totals[agent_id] = Totals(
pnl=sum(perf[tag].total_pnl for tag in owner_tags if tag in perf),
volume=sum(perf[tag].volume for tag in owner_tags if tag in perf),
)
if failed.intersection(owner_tags):
degraded.add(agent_id)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Missing bots overwrite reliable totals. If a previously measured bot has no discoverable live or archived instance, the aggregator marks its base unresolved but does not add its tag to failed_ids. This code treats the incomplete total as clean, overwrites the last reliable figure, and posts an apparent drop in PnL or volume.

Knowledge Base Used:

Comment on lines +183 to +185
owner_tags = list(owner.agent_ids) or [owner.run_key]
tags[agent_id].extend(owner_tags)
bases[owner_tags[0]] = [owner.namespace, *owner.declared_bots]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Session tags overwrite bot ownership. Strategies named grid and grid_1 can produce the same key when grid has session alice.grid_1 and grid_1 has no sessions. The latter's run-key fallback overwrites the former's bot-base entry, so bot PnL and volume can be omitted or assigned to the wrong race ID without a degradation signal.

Comment on lines +295 to +300
known = last_good.setdefault(server, {})
for agent_id in agent_ids:
if agent_id in fresh and agent_id not in degraded:
known[agent_id] = fresh[agent_id]
if agent_id in known:
totals[agent_id] = totals[agent_id].plus(known[agent_id])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Server changes reuse old figures. If a configured server name is pointed at a different host or port and the new endpoint is unavailable, last_good still contains figures from the old endpoint under that name. The monitor posts those old figures as though they came from the replacement server.

…ST result

The site wraps its answers in {"data": ...}. parse_agent_ids looked for a
top-level list and returned [] for the real response, so the monitor
stopped with "listed no agents". It now reads data.agents and raises on any
other shape. The POST result is read too: the accepted count is logged and
agent ids the site does not know are warned about.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment on lines +95 to +96
try:
return [agent["agent_id"] for agent in payload["data"]["agents"]]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Supported rosters now fail. If the race API returns a bare list or a list under agents, agent_ids, or data, as the PR description says it may, this parser rejects it. The first tick exits without posting; if the response shape changes after a successful post, the monitor keeps the old roster and misses new entrants.

session: Any, url: str, payload: dict[str, Any]
) -> tuple[int, list[str]]:
async with session.post(url, json=payload) as response:
if response.status >= 400:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Successful posts stop monitoring. If the race API accepts a POST but returns no JSON body, or omits either data.accepted or data.unknown, the new response parser raises an error. Because posted is set only after parsing succeeds, this exits the monitor on its first tick even though the figures may have been accepted, preventing all later updates.

The leaderboard names agents by botcamp strategy slug, which rarely equals the
Condor strategy slug, so every agent read zero. HACKATHON_AGENT_MAP (or
--agent-map) points at a YAML of race id -> run keys, re-read every minute;
an id in it is measured by those keys only, any other id by slug as before.
scripts/hackathon_agents.example.yml lists the 14 agents of
agent-builders-cup-1 with the keys their submissions declare.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment on lines +197 to +205
if agent_id in agent_map:
keys = agent_map[agent_id]
matched[agent_id] = [owner for owner in owners if owner.run_key in keys]
else:
matched[agent_id] = [
owner
for owner in owners
if agent_id in (owner.strategy_slug, owner.run_key)
]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 One strategy counted twice. If one race ID maps to a Condor strategy and another race ID matches that strategy’s slug, both IDs receive its full PnL and volume. The matching test permits this overlap, and the monitor posts both totals, crediting the same trading twice on the leaderboard.

from condor.agents.fleet_map import build_fleet_map
from config_manager import get_config_manager

agent_map = load_agent_map(self.args.agent_map) if self.args.agent_map else {}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Map edits corrupt race totals. If an agent’s map is edited after --start-race, the next tick measures the newly assigned strategy but subtracts the baseline saved for the old one. If a server is down, it can also reuse the old strategy’s last-good figure. Correcting an assignment mid-race therefore posts inaccurate PnL and volume.

fengtality and others added 2 commits October 3, 2026 20:05
… agent map

The builder resubmitted as a Condor agent; the race now lists the new strategy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s and closed executors

The dock's agent rows folded only the running population, so an agent's PnL
dropped as soon as it stopped a bot or an executor closed. Each agent row now
folds its live controllers plus its finished records (stopped bots' final
controller snapshots, realized only, and closed executors), with one History
row under it so the rows still add up. The live half keeps updating over the
socket.

The terminated-leaf construction moves from PerfBrowser to
lib/perf-population (terminatedLeaves) so /bots and the dock share it, and
useFleetData gains a `terminated` option to load runs and finished
controllers alongside the live fleet.
Comment on lines +270 to +277
terminatedLeaves({
executors: fleet.executors,
terminatedControllers: fleet.terminatedControllers,
runs: fleet.runs,
owners: fleet.owners,
deeds: fleet.deeds,
botByController: botsByController(fleet.controllers),
}),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Stopped bot totals misconverted. If a stopped bot trades in a non-USD quote currency that no live controller or loaded executor uses, the dock includes that bot in the agent’s history but does not request its exchange rate. Conversion leaves its PnL and volume unchanged while displaying them in the selected currency, so the agent’s lifetime totals are wrong.

Knowledge Base Used: Frontend application

Comment on lines +503 to +506
onOpen={() =>
navigate(
`/bots?population=terminated&scope=${encodeURIComponent(row.parentId ?? "")}`,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 History link omits older activity. The history row shows an agent’s lifetime total, but clicking it opens the terminated browser with its default three-month window. For agents with older stopped bots or closed executors, the detail view omits figures included in the row, making the total difficult to verify.

Knowledge Base Used: Frontend application

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants