Skip to content

Publish DeepSeek V4 Flash 0731 MMBT campaign - #44

Merged
Lightheartdevs merged 29 commits into
mainfrom
deepseek-v4-flash-extended
Aug 1, 2026
Merged

Publish DeepSeek V4 Flash 0731 MMBT campaign#44
Lightheartdevs merged 29 commits into
mainfrom
deepseek-v4-flash-extended

Conversation

@Lightheartdevs

Copy link
Copy Markdown
Contributor

What this publishes

This PR publishes the complete verified DeepSeek V4 Flash 0731 MMBT campaign:

  • the optimized Tower2 deployment recipe for 2x RTX PRO 6000 (1M served context, FP8 KV cache, 500 W/GPU);
  • canonical microbench results and corrected grader overlay;
  • extended single-PR, finance, board-material, and frozen 75-PR results;
  • compact audit JSON, hashes, completion audit, and verified-results synthesis;
  • reproducible runners, telemetry, validators, and harness/grader hardening;
  • root README, scorecard, claims registry, limitations, and findings-index integration.

Headline results

  • Canonical microbench: 23/36 raw; 35/36 after audited grader correction across N=3. The remaining miss is a genuine word-limit failure.
  • Single-PR audit: 3/3 complete artifacts, 2/3 expected verdicts, with one over-strict REVISE.
  • Wall Street finance tasks: 0/2 substantive success.
  • Board-material tasks: shipped outputs, but material deck-quality defects.
  • Frozen 75-PR audit: 0/3 strict success. Two runs were scaffold-heavy/incomplete; one became an 815,279-token runaway. Earlier artificially capped v2/v3 attempts are preserved but excluded from valid outcomes.
  • Production/performance validation: 1,048,576-token context, about 7,255 tok/s prefill, 272 tok/s single-stream decode, and 1,373 tok/s aggregate at concurrency 8, with restart/recall plus Sanctuary and Pixel validation.

Interpretation guardrails

  • DeepSeek used its documented model-specific sampling, context, and FP8 KV operating point; this is the best validated configuration for this hardware, not a uniform leaderboard setting.
  • DeepSeek canonical N=3 is less statistically mature than the older Qwen N=10 entries.
  • Corrected scores use the hardened grader and retain the raw score for auditability.
  • Cloud-model comparisons are categorical where deployment and telemetry are not comparable.
  • Full raw run archives are intentionally externalized under MMBT repository-space policy (the V5 archive alone is 137 MB). Compact audits, manifests, hashes, and verified summaries are committed.

Validation

  • Publication integrity validator: 10 JSON files parsed, 34 unique claims, 3 valid 75-PR outcomes.
  • All 17 changed Python files compile.
  • Targeted regression suite: 4 passed.
  • Full PR diff passes git diff --check.
  • No large GitHub blobs; largest blob in the branch is about 7.7 MB.

User Name added 29 commits August 1, 2026 01:31
The live backlog grew to 272 PRs, so a source-only fixture and explicit
pinned-baseline task note are required to preserve the original benchmark
semantics. Add a validated input_path mount without changing other suite
behavior.
@Lightheartdevs
Lightheartdevs marked this pull request as ready for review August 1, 2026 22:52
@Lightheartdevs
Lightheartdevs merged commit dcd9431 into main Aug 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant