Skip to content

[framework, tasks] feat: retain runner scoring diagnostics in trajectories - #237

Merged
yyDing1 merged 1 commit into
verl-project:mainfrom
tongyx361:pr/runner-reward-context
Oct 1, 2026
Merged

yyDing1 merged 1 commit into
verl-project:mainfrom
tongyx361:pr/runner-reward-context

Conversation

@tongyx361

@tongyx361 tongyx361 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Preserve runner scoring evidence in trajectory output so task failures can be diagnosed independently of the final training reward.

Changes

  • Framework writes the runner reward, accuracy, and task context to extra_fields.runner_reward_info.
  • SWE-bench retains evaluator exit codes, per-test results, and explicit agent errors.
  • Document the added fields and cover TransferQueue output, training masks, and SWE resolution behavior.
  • Attach the test log-capture handler to Uni-Agent's non-propagating namespace so existing logging assertions capture the intended records.

Validation

  • PYTHONPATH=.:verl uv run --no-project --python /usr/bin/python python -m pytest -q tests/uni_agent/tasks/test_swe_agent_error_on_cpu.py tests/uni_agent/tasks/test_swe_reward_on_cpu.py tests/uni_agent/framework/test_generate_sequences_on_cpu.py: 75 passed. Tests used the runtime interpreter with Uni-Agent and the pinned public verl gitlink explicitly selected.
  • pre-commit run --all-files --show-diff-on-failure: passed.

Internal validation and observed benefit

A historical internal GPU smoke evaluated four tasks across four model-version evaluations, producing 16 completed sessions. Persisted results contained runner reward, evaluator exit code and parsed test-status evidence for 16/16 sessions, and runner rewards matched the corresponding evaluation results in 16/16 cases. This made the task-side scoring evidence inspectable alongside the final result.

This is diagnostic coverage from an integrated implementation of the feature, not a before/after reward-quality comparison or an end-to-end validation of this split PR's exact head. No accuracy or throughput gain is claimed. The isolated PR regression tests are listed above.

Compatibility

The diagnostic fields are additive. Reward-worker scores, training masks, and SWE-bench resolution criteria retain their existing behavior. No migration is needed. This is a standalone observability improvement without a linked issue.

Checklist

  • The PR is focused and explains why no issue is needed.
  • The title follows the owning-layer format.
  • Behavior is covered by tests; skipped validation is stated above.
  • User-facing configuration and behavior changes are documented where applicable.
  • Compatibility and migration requirements are documented.
  • Logs, fixtures, and examples contain no credentials or private data.
  • pre-commit run --all-files --show-diff-on-failure passes.

@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yyDing1
yyDing1 merged commit 00b20d7 into verl-project:main Oct 1, 2026
6 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants