Repository navigation
spec: opt-in decision-model gate for subagent delegation and (model, effort) tier selection #6236
Closed
Yeachan-Heo
started this conversation in
Ideas
Replies: 1 comment
|
Closing per maintainer decision: not actionable at this time. Can be reopened if the need resurfaces. — |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
Spec-only issue. Requested by @probe4206 for local prototyping — no implementation is being started here.
Two related asks:
(model, effort)dynamically instead of pinning it per role.Both reduce to one cheap classification call per decision point, which is what a small decision model (Jev / the open-source Kev clone) is built for.
Why effort must be decided at the subagent boundary, not mid-session
Changing effort inside a live session invalidates prompt cache on both major providers:
output_config.effortis documented the same way.reasoning.effortat its original value as changing that setting can rewrite instructions in the hidden system instructions." Measured in Changing reasoning level results in cache miss openai/codex#35416 as an effectively full miss; opencode observed 100%.A subagent starts its own session and shares no prefix, so per-subagent
(model, effort)selection costs nothing in cache terms. That is the safe place to be dynamic.(OpenAI GPT-6+ also exposes a
configuration_updateinput item that changes effort mid-conversation while preserving the cached prefix. Out of scope here, but relevant if in-session dynamic effort is revisited.)Existing surfaces to build on
packages/coding-agent/src/config/autorouting-tier-map.ts— tier assignments already validate an optionalassignment.effortagainst anEFFORTSset, enforce at most one assignment per tier, and check provider/tier rank collisions. The(model, effort)pair is already the unit of routing.packages/coding-agent/src/hooks/events.ts—HOOK_EVENT_SCHEMAS,HookAuthority, andHookErrorBehavior(isolate | fail-closed) already distinguish advisory from blocking hooks.Proposed shape
Decision model exposes three primitives — noul (yes/no probability), choice (named options), score (ordered scale) — so one request covers both asks:
(model, effort)pair.choice scores arbitrary option sets against the question, so adding or renaming a tier needs no retraining — unlike a fixed-label classifier.
Tiers stay a small fixed set (e.g. fast / balanced / strong). Continuous per-request effort is explicitly rejected: it produces a unique cache lineage per request.
Role defaults follow call frequency — executor is called most often and defaults low; planner/architect are called once or twice per task and can default high.
Three modes, default off
offhintisolateenforcefail-closedenforcemust stay closed until a threshold is validated on our own data. Kev's published limitation: Kev-4B assigns >=90% confidence to a wrong answer on 8.2% of new-source development questions (Kev-9B: 7.5%), and its README states plainly: "Test it on your own data before choosing a probability threshold." At that rate a blocking gate would reject correct direct edits often enough to be worse than no gate.Decision criteria must be observable
Whatever triggers the noul question has to be computable from turn state — files touched, consecutive edit count, diff size. "Looks complex" is not implementable.
Domain adaptation
Public checkpoints will not match this repo's categories. Kev documents
kev.train --init_from, which continues from a released checkpoint's LoRA adapter and pointer head. Their reported before/after on 836 labeled examples: training from the base model collapsed to 0.33 on the held-out set, while--init_fromheld 0.83 there and reached 0.88 on the new domain. A few hundred labeled decisions is the realistic budget.Risks to size before building
127.0.0.1. Fine for single-machine prototyping, must be closed before anything wider.Acceptance for a future implementation
off;hintandenforceare explicit opt-in.Sources: Anthropic prompt-caching / extended-thinking docs, OpenAI prompt-caching guide, openai/codex#35416, and explainx.ai's write-up of jaredpalmer/kev.
—
[repo owner's gaebal-gajae (clawdbot) 🦞]
—
[repo owner's gaebal-gajae (clawdbot) 🦞]
All reactions