Skip to content

Publish Gemma 4 31B Q4 verified Tower2 results - #45

Merged
Lightheartdevs merged 32 commits into
mainfrom
gemma4-31b-q4-mmbt
Aug 2, 2026
Merged

Publish Gemma 4 31B Q4 verified Tower2 results#45
Lightheartdevs merged 32 commits into
mainfrom
gemma4-31b-q4-mmbt

Conversation

@Lightheartdevs

Copy link
Copy Markdown
Contributor

What changed

  • Adds the complete Gemma 4 31B QAT Q4_0 Tower2 campaign: preregistration, pinned model/runtime/topology, canonical N=3 and N=10 scorecards, correction overlays, grader manifests, evidence audits, and extended-suite audits.
  • Documents the selected two-GPU operating point: one independent four-slot llama.cpp replica per RTX PRO 6000, Q8 KV, native 262,144-token context per slot, and 500 W per-GPU caps.
  • Publishes direct comparisons to Qwen3.6-27B, Qwen3-Coder-Next, Qwen3.5-397B-A17B Q3, Qwen3.6-35B-A3B, and DeepSeek V4 Flash 0731 without treating unlike operating points as a controlled leaderboard.
  • Hardens the 75-PR and completion gates so missing artifacts fail closed and a qualifying completion tag must be annotated, point at HEAD, and belong to a clean candidate repository.
  • Records the exact post-campaign restoration to the proven DeepSeek V4 Flash deployment, including fresh Sanctuary and Pixel tool-use checks with no fallback.

Key results

  • Canonical N=3: 29/36 raw, 32/36 corrected.
  • Canonical N=10: 89/120 raw, 99/120 corrected; 116/120 ordinary completions.
  • The only correction is the project-management family (0/10 raw to 10/10 corrected), with every changed cell tied to immutable grader, report, archive, and correction-script hashes.
  • Versus Qwen3.6-27B: Gemma leads the comparable N=3 bounded-quality cohort (29/36 raw versus 20/36), while short-context single-stream speed is effectively tied (70.3 versus 72.1 tok/s).
  • Qwen3.6-27B/vLLM remains the high-concurrency serving winner (1,336.5 aggregate tok/s at C32 on one GPU versus Gemma's 290.3 aggregate tok/s at total C8 across two GPUs; operational comparison, not a controlled model-speed A/B).
  • Extended suite: 0/12 strict substantive passes. The model often produced polished but incomplete finance/presentation work and failed every frozen 75-PR marathon attempt under strict audit.

Why

This supplies the missing Gemma 31B data needed to compare a strong dense Q4 model with the existing Qwen and DeepSeek Tower2 results. It also records the practical boundary between bounded task quality and unattended marathon reliability.

Validation

  • python3 -m pytest -q tooling/test_*.py tooling/scripts/test_*.py tooling/deployments/gemma4-31b-q4-tower2/test_*.py — 62 passed.
  • python3 tooling/validate_gemma4_comparison_sources.py — passed.
  • python3 tooling/deployments/gemma4-31b-q4-tower2/validate_preregistration.py — valid.
  • sha256sum -c benchmarks/gemma4-31b-q4/SHA256SUMS — all files OK.
  • All published JSON and claims.yaml parse; Markdown local-link audit and git diff --cached --check pass.
  • Post-restore: DeepSeek health 200, served model DeepSeek-V4-Flash-0731, 1,048,576-token context, 500 W caps, both portals healthy, and fresh Sanctuary/Pixel exec traces pass with fallbackUsed=false.

User Name added 30 commits August 1, 2026 19:11
@Lightheartdevs
Lightheartdevs marked this pull request as ready for review August 2, 2026 07:13
@Lightheartdevs
Lightheartdevs merged commit eaaa8ca into main Aug 2, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant