Repository navigation
Conversation
Consolidate coding, memory, capacity, and skill suites under evaluation and keep each suite's commands, tests, documentation, and execution tools local. Include the work-continuity benchmark from the latest master branch. Fix issues found by bounded live smoke tests: Scope reuse, configured database identity, HTTP model endpoints, and isolated OceanBase SWE runs. Preserve backend evidence through report loading and align console retry history with the current API and automatic retry behavior.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
| task = load_tasks(_REPOSITORY / "e2e" / "bub" / manifest)[0] | ||
| protected = [_REPOSITORY / "e2e" / "bub" / name for name in ("harbor-tasks", "paired-tasks", "tasks")] | ||
| protected.append(_REPOSITORY / "benchmark") | ||
| protected.append(_REPOSITORY / "evaluation") |
There was a problem hiding this comment.
[P2] Heads with #1853: e2e/bub/tests/test_harbor_job_config.py conflicts, and one side is a security-relevant path
This file is the only conflict between the two branches, but it is worth resolving deliberately rather than by taking one side. git merge-tree 673d44d6 <this head> 799ea12c reports exactly one changed in both:
base 100644 b9fe8124... e2e/bub/tests/test_harbor_job_config.py
our 100644 293e50a6... e2e/bub/tests/test_harbor_job_config.py
their 100644 06a40cbe... e2e/bub/tests/test_harbor_job_config.py
This PR changes one line there, and it is the one that matters:
- protected.append(_REPOSITORY / "benchmark")
+ protected.append(_REPOSITORY / "evaluation")That is the protected list handed to the agent container in test_agent_container_cannot_read_workload_answers. Since this PR moves benchmark/ to evaluation/, dropping that line would leave the answers directory readable by the agent under test — the assertion would still be testing something, but not the protection.
#1853 rewrites the same file (about 120 added lines in test_pi_runs_without_a_saved_session, plus imports and five other tests), so a textual conflict is unavoidable; the two changes do not conflict semantically. Whichever side lands first, please keep both: the protected entry pointing at evaluation, and #1853's new tests.
Also worth re-checking after the merge: #1853 has an open review comment on this file (4192361330) whose line anchors were computed against its own head.
Which issue or RFC does this PR close?
No linked issue; this is a repository evaluation-layout cleanup.
Rationale for this change
Benchmarks and evaluation tools were split between
benchmark/and a shared evaluation package. Group them by what they measure, with each suite owning its commands, tests, documentation, and execution tools.What changes are included in this PR?
evaluation/coding,evaluation/memory,evaluation/performance, andevaluation/skills. Keep the SWE console, worker, and deployment files inside SWE-bench Pro. Include the work-continuity benchmark from current master under coding.settings.databasewith a local SQLite fallback; persist backend identity and reject cross-database resumes.Are there any user-facing changes?
Evaluation paths, Python imports, and commands change. Replace
powercontext-eval ...with the documentedswebench-pro,longmemeval-v2, orwork-continuitycommand; LoCoMo and capacity modules now run underevaluation.*. There is no compatibility wrapper. Existing dataset contents and result schema identifiers are preserved. LoCoMo resumes require a valid saved Scope mapping, and LoCoMo-Plus resumes must match the recorded backend. OceanBase resume identity uses the effective database target and excludes authentication values. Saved OceanBase runs with an older fingerprint format remain replayable but require a new output directory for execution.How was this change tested?
On the merged branch:
make check: passed, including hooks, lock consistency, type checks, generated contracts, and 30 integration-manifest tests.pytest -m "not live": 1,530 passed.make evaluation-unit-testalso passed all 705 cases with GitHub Actions ANSI colors enabled; CLI assertions normalize formatting while retaining error and credential-redaction checks.run_benchmarkpassword rotation, effective target changes, and replay/resume guards using the installed OceanBase dialect. These checks do not connect to OceanBase or call an external model.make harness-check(134 tests), andmake harness-compose-checkpassed.make docs-testwith Node 22: passed, including static export and internal links across 875 public pages.Bounded live samples used runtime base
5cf66be6plus a separate runtime-fix snapshot. The successful Alpine OceanBase startup samples used lazy SQLite extension imports; this PR retains upstream top-level imports and requires the complete locked dependency set, so those samples do not establish Alpine compatibility of this branch. Memory-capacity operations passed on real OceanBase. LoCoMo/Plus persisted real data; LongMemEval completed 10/10 FTS retrieval and 7/10 Reader answers. SWE Gold passed and both OFF and a separate ON-only attempt executed real model-generated commands against isolated OceanBase databases; ON ingestion persisted a Source. Provider account refusals and other upstream response errors prevented full scoring and a completed SWE pair. These samples do not establish complete benchmark accuracy or a successful OFF/ON comparison. Skill-up used a real host/model with mocked MCP. Temporary databases and credentials were cleaned up.AI usage statement
Implemented and validated with OpenAI Codex (GPT-6), including parallel agents for migration and validation.