Skip to content

Bound automatic compaction input to model capacity - #1160

Open
KeenWill wants to merge 1 commit into
agent/daemon-live-restore-call-lexingfrom
agent/daemon-live-model-bounded-compaction
Open

Bound automatic compaction input to model capacity#1160
KeenWill wants to merge 1 commit into
agent/daemon-live-restore-call-lexingfrom
agent/daemon-live-model-bounded-compaction

Conversation

@KeenWill

Copy link
Copy Markdown
Owner

Summary

  • cap an automatic context-compaction summary prefix at the smaller of half the visible transcript content and the selected model's context window after its configured output reservation
  • preserve the existing first-safe-boundary rule, including closing a tool exchange that crosses the derived target
  • reject malformed automatic preparation requests whose model-derived target is absent or attached to an explicit request
  • update the model-call execution contract for the model-bounded selection rule

This prevents a large accumulated transcript from making its own automatic summary request exceed the selected model's capacity.

Meaningfully changed lines: 112 (excluding lockfiles).

Numeric bounds added: none. The target is derived from the already configured per-model context window and maximum output reservation.

Validation

  • cargo fmt --all -- --check
  • workspace check and clippy with all targets/features
  • targeted automatic-compaction unit and process-protocol tests
  • complete ignored PostgreSQL integration suite with postgres-integration
  • workspace doctests and warning-denied rustdoc
  • documentation, domain-spine, migration, numeric-bound, panic, and invariant checks

CI is the authoritative complete-bar backstop.

Absorption ledger

Stack base: #1159 (agent/daemon-live-restore-call-lexing).

@coderabbitai

coderabbitai Bot commented Aug 23, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: a9ff5ce0-14af-4d64-893a-64970f78d40e

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Your free Security trial is over. An organization admin can activate billing to continue.


Comment @coderabbitai help to get the list of available commands.

@KeenWill
KeenWill marked this pull request as ready for review August 23, 2026 06:43
Copilot AI lite review requested due to automatic review settings August 23, 2026 06:43
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@KeenWill

Copy link
Copy Markdown
Owner Author

@codex review

Exact reviewed head: e3c854e

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e3c854e298

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +7266 to +7268
let automatic_content_byte_target = u64::from(definition.context_window_tokens())
.checked_sub(u64::from(definition.max_output_tokens()))
.and_then(NonZeroU64::new)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Bound the rendered compaction request, not raw content

When a frontier contains many small entries—or content requiring JSON escaping—this target does not actually bound the model input: bounded_safe_boundary counts only stored payload bytes, while load_context_compaction_range serializes every selected entry with IDs, type metadata, field names, delimiters, and escaping, and the request also adds the nonempty compaction system prompt. The selected prefix can therefore render far beyond context_window_tokens - max_output_tokens, causing the provider to reject the sole automatic compaction attempt and leaving the oversized queued turn unrecoverable. Compute the boundary from the rendered/token-counted request or reserve all serialization and prompt overhead.

AGENTS.md reference: AGENTS.md:L116-L120

Useful? React with 👍 / 👎.

total.saturating_add(member.content_bytes.max(1))
});
let midpoint_weight = total_weight.div_ceil(2);
let target_weight = total_weight.div_ceil(2).min(content_byte_target);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Ensure one compaction can make the successor fit

When a session history is sufficiently larger than twice the model's input capacity, capping the summarized prefix at content_byte_target leaves a suffix that exceeds the same capacity even if the summary call itself fits and returns a minimal summary. The counted activation path then renders the successor, detects that it is still oversized, and returns ContextStillExceeded because this turn has already used its single automatic compaction; the queued turn is permanently stalled. The boundary strategy must either guarantee that the retained suffix plus summary can fit or support multiple bounded chunks before consuming the turn's sole attempt.

AGENTS.md reference: AGENTS.md:L252-L257

Useful? React with 👍 / 👎.

@github-actions

Copy link
Copy Markdown
Contributor

Rust coverage (report only)

Report only. This measurement has no threshold, gates no merge, and
fails no check; it exists so untested code stays visible.

Measured suite Outcome
workspace (--all-targets --all-features) success
persistence PostgreSQL (--ignored) success
signalboxd PostgreSQL (--ignored) failure
terminal-client PostgreSQL (--ignored) success
What this number does not measure
  • Doctests. cargo llvm-cov --doctests needs a nightly
    toolchain; this workspace pins stable, so the compile-fail
    sealing proofs and every other doctest are outside the
    denominator.
  • Live smokes, which spend real network requests or
    credentials and are never run here: the tools-github and
    tools-web smokes stay --ignored, and the whole-daemon and
    real-provider terminal smokes are skipped by name above.
  • The Swift native client, which Xcode measures separately.
  • One environment-clearing test in signalbox-tools-exec,
    which instrumenting its supervisor process perturbs. The
    workflow comment on the workspace step states why; rust.yml
    runs that test uninstrumented and gates on it.
  • Dedicated test files. cargo-llvm-cov excludes tests/
    and benches/ targets and *tests.rs modules from the
    report by default, so that test code is counted neither
    covered nor uncovered here. Inline #[cfg(test)] modules
    inside a source file are the exception: they are
    instrumented, and they land on both sides. A test body
    that ran counts as covered, which makes every percentage
    below optimistic; the body of an #[ignore]d test no
    measured suite runs counts as uncovered, which puts test
    lines into the file ranking. Read both tables as close,
    not exact.
Measure Covered Total Percent
Lines 254023 311240 81.62%
Functions 20713 25166 82.31%
Regions 321800 402653 79.92%

Per crate, least-covered first

Crate Line % Lines Function % Region %
crates/program-runtime 7.57% 38/502 5.77% 6.13%
crates/runner-wire 60.41% 644/1066 62.67% 59.14%
apps/signalboxd 63.67% 35477/55724 70.84% 63.54%
crates/tools-sessions 65.06% 378/581 63.29% 61.68%
apps/signalbox-runner 69.47% 1784/2568 70.93% 70.29%
crates/approval-judge-eval 73.87% 492/666 76.12% 76.41%
crates/tool-schema-derive 74.20% 279/376 88.00% 72.42%
crates/tools-basic 75.37% 771/1023 66.15% 79.16%
crates/model-runtime-claude-cli 76.51% 1655/2163 73.10% 77.37%
crates/tools-github 77.06% 2408/3125 74.19% 74.82%
crates/tools-exec 77.75% 5195/6682 75.04% 74.08%
apps/client 78.88% 13283/16840 89.58% 75.76%
crates/persistence 80.08% 53982/67414 78.02% 75.48%
crates/tools-code-host 80.39% 7089/8818 80.46% 76.60%
crates/blob-store-filesystem 80.50% 1371/1703 70.17% 80.45%
crates/model-provider-runtime 81.37% 2463/3027 85.71% 79.97%
crates/blob-store 82.57% 308/373 75.41% 83.68%
crates/tools-conversations 83.01% 508/612 83.78% 77.65%
crates/egress-transport 85.35% 134/157 83.33% 77.61%
crates/conversation-import-claude-code 85.45% 740/866 66.09% 84.67%
crates/tools-plan 85.71% 768/896 88.89% 81.74%
crates/conversation-import-codex 85.76% 873/1018 64.00% 83.56%
crates/tools-workspace 86.36% 2977/3447 82.14% 87.30%
crates/tools-web 86.59% 2978/3439 87.93% 85.03%
crates/tools-git 87.40% 8750/10011 86.50% 83.24%
crates/application 88.65% 19184/21640 87.86% 88.81%
crates/model-runtime-codex-cli 90.01% 1550/1722 90.98% 89.96%
crates/test-bin 90.91% 10/11 100.00% 70.00%
crates/domain 92.28% 60165/65196 91.37% 94.13%
crates/web-contract 92.31% 408/442 82.61% 82.09%
crates/conversation-import-json 92.51% 284/307 96.88% 91.36%
crates/model-runtime-openai 93.55% 3812/4075 97.50% 91.72%
crates/process-protocol 93.67% 8469/9041 95.89% 86.68%
crates/model-runtime-anthropic 93.88% 3910/4165 97.51% 91.26%
crates/model-runtime 93.96% 9221/9814 93.29% 94.69%
crates/expect-table 95.60% 977/1022 100.00% 95.75%
crates/tool-contract 97.18% 688/708 96.47% 95.54%

25 files with the most uncovered lines

File Uncovered lines Line %
apps/signalboxd/src/process_runtime.rs 6580 53.40%
apps/signalboxd/src/repo_watch_runtime.rs 2235 69.20%
apps/signalboxd/src/runner_protocol_runtime.rs 2212 34.75%
crates/persistence/src/submit_input.rs 2071 71.81%
apps/client/src/lib.rs 1697 79.51%
apps/signalboxd/src/review_orchestration_runtime.rs 1333 2.56%
apps/signalboxd/src/main.rs 1303 47.46%
crates/persistence/src/model_execution.rs 1237 84.68%
crates/domain/src/turn_eligibility.rs 1223 90.33%
crates/persistence/src/runner_protocol.rs 1157 82.16%
crates/tools-code-host/src/code_host/github.rs 1052 76.26%
crates/persistence/src/review_workflow.rs 997 77.35%
crates/tools-exec/src/bin/signalbox-exec-supervisor.rs 885 45.61%
crates/persistence/src/tool_loop.rs 749 76.07%
crates/tools-github/src/lib.rs 717 76.47%
apps/signalboxd/src/convergence_sweep_runtime.rs 683 22.21%
crates/persistence/src/process_read.rs 681 82.33%
apps/client/src/presentation.rs 680 80.55%
apps/signalboxd/src/lib.rs 647 72.20%
crates/domain/src/submit_input.rs 640 88.03%
apps/signalboxd/src/daemon_tools.rs 573 89.68%
crates/process-protocol/src/lib.rs 572 93.67%
crates/persistence/src/review_orchestration.rs 569 75.66%
crates/persistence/src/session_delegation.rs 543 77.63%
crates/application/src/model_execution.rs 541 86.15%

Measured at f7b55ca823cffab252bedba9bb9b26ec267a4af8, the merge commit this pull request builds, whose head is e3c854e2980d937a26bbcdf2e4b4c5aaef41bad2, by run 32622752279, which uploads the HTML report and LCOV as an artifact.

@codecov

codecov Bot commented Aug 23, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 74.57627% with 15 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.27%. Comparing base (6541cd6) to head (e3c854e).

Files with missing lines Patch % Lines
apps/signalboxd/src/process_runtime.rs 0.00% 13 Missing ⚠️
crates/persistence/src/context_compaction.rs 95.65% 2 Missing ⚠️
Additional details and impacted files
Flag Coverage Δ
rust 82.85% <74.57%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants