Skip to content

Commit 0a3b387

Browse files
Fatih Dincclaude
authored andcommitted
docs: honest positioning after Ollama comparison
Ollama qwen2.5:14b scores 40/40 with --format json. Updated tagline, comparison table, and blog post to position on speed (10-25x faster) and no server dependency instead of compliance exclusivity. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1 parent 41c71e3 commit 0a3b387

3 files changed

Lines changed: 43 additions & 40 deletions

File tree

‎TODOS.md‎

Lines changed: 9 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -12,21 +12,15 @@ Generated by /plan-eng-review on 2026-04-30
1212
- **Context:** Design-Doc spezifiziert API-Kontrakt bereits. ~100 Zeilen Code + Tests. Blocked by: llmjson Package fertig.
1313
- **Depends on:** llmjson Package (Step 1-3 des Sprint)
1414

15-
### TODO-002: Ollama-Vergleichstest
16-
- **What:** Teste Ollama `--format json` gegen die 4 BNR-002-Schemas (simple_object, nested_object, financial_transfer, medical_record). Dokumentiere Fehlerrate.
17-
- **Why:** Differenzierungsmerkmal für Marketing. Wenn Ollama 100% schafft, hat llmjson kein USP.
18-
- **Pros:** Harter Beweis für den Blogpost, nicht nur Behauptung
19-
- **Cons:** Falls Ollama ebenfalls 100% liefert, ist das Marketing-Argument weg
20-
- **Context:** Ollama bietet GBNF-Grammar-Support, operiert aber auf Grammar-Ebene, nicht Logit-Ebene. Erwartung: niedrigere Compliance bei komplexen Schemas.
21-
- **Depends on:** —
22-
23-
### TODO-003: Blogpost-Titel und Claims ehrlich formulieren
24-
- **What:** "100% valid JSON from local LLMs" ersetzen durch "100% valid JSON for supported schema types from local LLMs" oder äquivalent
25-
- **Why:** Outside Voice: HN findet Edge Cases in Stunden. Schema-Einschränkungen müssen prominent stehen, nicht als Fußnote.
26-
- **Pros:** Ehrlich, baut Vertrauen auf, vermeidet HN-Backlash
27-
- **Cons:** Weniger catchy Headline
28-
- **Context:** Sicherer Nutzungsbereich schließt $ref, patternProperties, additionalProperties aus. Das muss in README, Blogpost, und Landing Page klar kommuniziert werden.
29-
- **Depends on:** —
15+
### TODO-002: Ollama-Vergleichstest ✅ DONE (2026-04-30)
16+
- **Ergebnis:** Ollama qwen2.5:14b --format json: 40/40 = 100% schema compliance
17+
- **Konsequenz:** "100% valid JSON" ist kein USP. Differenzierer ist Speed (10-25x schneller) und keine Server-Abhängigkeit.
18+
- **Test:** tests/test_ollama_comparison.py
19+
- **Marketing angepasst:** Tagline, Vergleichstabelle, Blogpost aktualisiert.
20+
21+
### TODO-003: Blogpost-Titel und Claims ehrlich formulieren ✅ DONE (2026-04-30)
22+
- **Ergebnis:** Schema-Support-Tabelle auf Landing Page, Qualifier im Blogpost, Ollama-Vergleich ehrlich dokumentiert.
23+
- **Tagline:** "Valid JSON from local LLMs. No server, no retries, 10x faster than Ollama."
3024

3125
## Post-Launch (Woche 3+)
3226

‎site/blog/how-i-got-100-percent-valid-json.html‎

Lines changed: 10 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -274,22 +274,24 @@ <h2>Limitations (honest ones)</h2>
274274

275275
<p><strong>Token alignment.</strong> Most tokenizers split text into subword tokens that don't align perfectly with JSON structural characters. A string like <code>"name":</code> might be one token or four. The tracker handles this by processing decoded text character by character, but edge cases exist with unusual tokenizers.</p>
276276

277-
<h2>Why local matters</h2>
277+
<p><strong>Ollama does this too.</strong> I ran the same 4 schemas through Ollama's <code>--format json</code> with Qwen2.5:14b. It scored 40/40. Ollama uses GBNF grammars to constrain output and it works. If you already use Ollama and don't care about speed, it's a fine choice. llmjson is 10-25x faster on the same hardware because it talks to MLX directly instead of going through a server.</p>
278278

279-
<p>OpenAI's Structured Outputs mode is good. If you're using their API and you're fine sending your data to their servers, it might be all you need.</p>
279+
<h2>Why llmjson over Ollama?</h2>
280280

281-
<p>But there are reasons to want local inference:</p>
281+
<p>Both guarantee valid JSON. The differences are practical:</p>
282282

283283
<ul>
284-
<li><strong>Privacy:</strong> The data never leaves your machine. Medical records, financial data, proprietary content.</li>
285-
<li><strong>Cost:</strong> After the one-time model download, generation is free. No per-token charges.</li>
286-
<li><strong>Latency:</strong> No network round-trip. On an M3 Max, generation starts in milliseconds.</li>
287-
<li><strong>Control:</strong> You pick the model, the parameters, the schema. No API limits, no deprecation notices.</li>
284+
<li><strong>Speed:</strong> llmjson generates in ~1-2 seconds. Ollama takes 10-25 seconds for the same schema on the same Mac. Direct MLX access vs. HTTP intermediary.</li>
285+
<li><strong>No daemon:</strong> llmjson is a Python library. Import it, call it. No background server to manage.</li>
286+
<li><strong>Privacy:</strong> The data never leaves your process. No local network socket, no shared memory.</li>
287+
<li><strong>Logit access:</strong> You get the raw probability distribution. Useful if you're building something beyond basic JSON generation.</li>
288288
</ul>
289289

290+
<p>If you need an HTTP API, use Ollama. If you want a library you can embed in your Python code with maximum speed, use llmjson.</p>
291+
290292
<h2>How to try it</h2>
291293

292-
<pre><code>pip install jsongate</code></pre>
294+
<pre><code>pip install llmjson</code></pre>
293295

294296
<p>You'll need Python 3.9+ and an Apple Silicon Mac (M1 or later). The MLX dependencies install automatically.</p>
295297

‎site/index.html‎

Lines changed: 24 additions & 17 deletions
Original file line numberDiff line numberDiff line change
@@ -262,8 +262,8 @@
262262

263263
<section class="hero">
264264
<h1><span>llmjson</span></h1>
265-
<p class="tagline">100% valid JSON from local LLMs. Zero retries. Zero post-processing. Constrained decoding on Apple Silicon.</p>
266-
<div class="install-box" onclick="navigator.clipboard.writeText('pip install jsongate')" title="Click to copy">pip install jsongate</div>
265+
<p class="tagline">Valid JSON from local LLMs. No server, no retries, 10x faster than Ollama. Constrained decoding on Apple Silicon.</p>
266+
<div class="install-box" onclick="navigator.clipboard.writeText('pip install llmjson')" title="Click to copy">pip install llmjson</div>
267267
</section>
268268

269269
<section class="stats">
@@ -334,49 +334,56 @@ <h3>Valid JSON, guaranteed</h3>
334334
</section>
335335

336336
<section class="compare">
337-
<h2>Why not just retry?</h2>
337+
<h2>How does it compare?</h2>
338338
<table>
339339
<thead>
340340
<tr>
341341
<th>Approach</th>
342342
<th>100% valid</th>
343-
<th>No retries</th>
343+
<th>Speed</th>
344+
<th>No server</th>
344345
<th>Local/private</th>
345-
<th>Schema-aware</th>
346346
</tr>
347347
</thead>
348348
<tbody>
349349
<tr class="highlight-row">
350350
<td>llmjson</td>
351351
<td class="yes">Yes</td>
352-
<td class="yes">Yes</td>
352+
<td class="yes">~1-2s</td>
353353
<td class="yes">Yes</td>
354354
<td class="yes">Yes</td>
355355
</tr>
356356
<tr>
357-
<td>Retry + parse</td>
358-
<td class="no">No</td>
359-
<td class="no">No</td>
360-
<td class="partial">Depends</td>
361-
<td class="no">No</td>
357+
<td>Ollama --format json</td>
358+
<td class="yes">Yes</td>
359+
<td class="partial">~10-25s</td>
360+
<td class="no">No (daemon)</td>
361+
<td class="yes">Yes</td>
362362
</tr>
363363
<tr>
364-
<td>OpenAI JSON mode</td>
365-
<td class="partial">Mostly</td>
364+
<td>OpenAI Structured Outputs</td>
366365
<td class="yes">Yes</td>
366+
<td class="yes">~1-3s</td>
367+
<td class="no">No (API)</td>
367368
<td class="no">No</td>
368-
<td class="partial">Partial</td>
369369
</tr>
370370
<tr>
371371
<td>Outlines / Guidance</td>
372372
<td class="yes">Yes</td>
373+
<td class="yes">~1-3s</td>
373374
<td class="yes">Yes</td>
374375
<td class="yes">Yes</td>
375-
<td class="yes">Yes</td>
376+
</tr>
377+
<tr>
378+
<td>Retry + parse</td>
379+
<td class="no">No</td>
380+
<td class="no">Varies</td>
381+
<td class="partial">Depends</td>
382+
<td class="partial">Depends</td>
376383
</tr>
377384
</tbody>
378385
</table>
379-
<p style="color: var(--text-muted); margin-top: 16px; font-size: 14px;">llmjson is purpose-built for Apple Silicon with MLX. Outlines and Guidance target CUDA/CPU. Pick the right tool for your hardware.</p>
386+
<p style="color: var(--text-muted); margin-top: 16px; font-size: 14px;">llmjson is a Python library for Apple Silicon (MLX). No server process, no API keys. Outlines and Guidance target CUDA/CPU. Ollama works but runs 10-25x slower on the same hardware. Pick the right tool for your setup.</p>
380387
</section>
381388

382389
<section class="compare" style="padding-top: 0; border-top: none;">
@@ -420,7 +427,7 @@ <h2>CLI included</h2>
420427
<h2>Get started in 30 seconds</h2>
421428
<p>Requires Python 3.9+ and Apple Silicon (M1/M2/M3/M4).</p>
422429
<a href="https://github.com/fatdinhero/llmjson" class="btn">GitHub</a>
423-
<a href="https://pypi.org/project/jsongate/" class="btn btn-secondary">PyPI</a>
430+
<a href="https://pypi.org/project/llmjson/" class="btn btn-secondary">PyPI</a>
424431
<a href="blog/how-i-got-100-percent-valid-json.html" class="btn btn-secondary">Blog Post</a>
425432
</section>
426433

0 commit comments

Comments
 (0)