Skip to content

Sync HarfBuzz mark and grapheme ordering - #496

Merged
behdad merged 3 commits into
mainfrom
fix/mark-order-sync
Oct 5, 2026
Merged

behdad merged 3 commits into
mainfrom
fix/mark-order-sync

Conversation

@behdad

@behdad behdad commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Port existing HarfBuzz mark/grapheme-ordering behaviors missing from HarfRust, found while testing the merged 104,404-glyph Noto font. This PR is independent of the ISO OFF extended-layout work and is based on main.

Three focused commits:

  • Telugu length marks use modified combining classes 4 and 5 rather than zero. Their canonical Unicode classes remain 84 and 91. This makes them reorder before virama, matching HarfBuzz.
  • USE includes symbol clusters in its syllable-reordering filter. The existing machine already recognizes these clusters with pre-base vowel marks; the reorder filter omitted them.
  • When no script was detected, native direction defaults to LTR, matching HarfBuzz. Explicitly directionless scripts still retain that behavior. This preserves emoji base/modifier ordering in RTL runs.

Reduced differential cases:

  • Tamil shaper: U+0C4D U+0C55.
  • Sharada shaper: U+111CD U+111CE.
  • Nandinagari shaper: U+119E3 U+119E4.
  • No detected script, RTL direction: U+1F44D U+1F3FD.

Direct regression tests cover the modified combining classes, USE cluster classification/reordering, emoji grapheme ordering, and explicitly directionless scripts. The Telugu and USE tests fail before their fixes. Validation includes full workspace tests, strict all-features/all-targets Clippy, and Rust 1.85 no-std/libm checks.

These changes are handwritten; the Ragel source and generated machines are unchanged.

With these fixes integrated with the pending Fontations CAPS/extended-layout readers and HarfRust extended-layout branch, all 4,374 multilingual and exhaustive Unicode-batch shaping comparisons match HarfBuzz on the original and Fontations-roundtripped mega font. Before the fixes, the sweep reproduced the reduced cases above.

Test font: https://drive.google.com/file/d/1vMQI9mZMzfKIMHD96Iie40Oyo9NnnweH/view

behdad added 3 commits October 4, 2026 14:38
Port HarfBuzz's modified combining classes 4 and 5 for Telugu length
marks instead of zeroing their classes. This preserves reordering
before virama and keeps the two length marks ordered relative to
each other, without changing their canonical Unicode classes.

The mega Noto sweep exposed the difference on U+0C4D U+0C55 under
the Tamil shaper. A direct regression test fails with the old values.
Full workspace tests, strict Clippy, and the MSRV/no-std check pass.

Assisted-by: OpenAI Codex
Match HarfBuzz by including symbol clusters among the USE syllables
eligible for reordering. The machine already permits vowel marks in
these clusters, but the reorder filter omitted their syllable type.

The mega Noto sweep exposed this on Sharada U+111CD U+111CE and
Nandinagari U+119E3 U+119E4. Test OTHER, BASE_OTHER, and SB bases
with both pre-base vowel categories and verify merged clusters.
The regression fails before the fix. Full workspace tests, strict
Clippy, and the MSRV/no-std check pass.

Assisted-by: OpenAI Codex
Match HarfBuzz's LTR fallback for an unspecified script. Previously
None became Invalid, suppressing conversion to native direction and
reversing emoji bases and their modifiers in RTL runs.

Preserve Invalid for scripts whose direction is explicitly unspecified.
Test emoji grapheme order with absent/common/unknown scripts and
preserve behavior for Old Hungarian, Old Italic, Runic, and Tifinagh.
Full workspace tests, strict Clippy, and the MSRV/no-std check pass.

Assisted-by: OpenAI Codex
@behdad
behdad requested a review from dfrg October 4, 2026 20:39
@behdad
behdad merged commit b5bcc43 into main Oct 5, 2026
3 checks passed
@behdad
behdad deleted the fix/mark-order-sync branch October 5, 2026 03:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant