Repository navigation
Sync HarfBuzz mark and grapheme ordering - #496
Merged
Merged
Conversation
Port HarfBuzz's modified combining classes 4 and 5 for Telugu length marks instead of zeroing their classes. This preserves reordering before virama and keeps the two length marks ordered relative to each other, without changing their canonical Unicode classes. The mega Noto sweep exposed the difference on U+0C4D U+0C55 under the Tamil shaper. A direct regression test fails with the old values. Full workspace tests, strict Clippy, and the MSRV/no-std check pass. Assisted-by: OpenAI Codex
Match HarfBuzz by including symbol clusters among the USE syllables eligible for reordering. The machine already permits vowel marks in these clusters, but the reorder filter omitted their syllable type. The mega Noto sweep exposed this on Sharada U+111CD U+111CE and Nandinagari U+119E3 U+119E4. Test OTHER, BASE_OTHER, and SB bases with both pre-base vowel categories and verify merged clusters. The regression fails before the fix. Full workspace tests, strict Clippy, and the MSRV/no-std check pass. Assisted-by: OpenAI Codex
Match HarfBuzz's LTR fallback for an unspecified script. Previously None became Invalid, suppressing conversion to native direction and reversing emoji bases and their modifiers in RTL runs. Preserve Invalid for scripts whose direction is explicitly unspecified. Test emoji grapheme order with absent/common/unknown scripts and preserve behavior for Old Hungarian, Old Italic, Runic, and Tifinagh. Full workspace tests, strict Clippy, and the MSRV/no-std check pass. Assisted-by: OpenAI Codex
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Port existing HarfBuzz mark/grapheme-ordering behaviors missing from HarfRust, found while testing the merged 104,404-glyph Noto font. This PR is independent of the ISO OFF extended-layout work and is based on main.
Three focused commits:
Reduced differential cases:
U+0C4D U+0C55.U+111CD U+111CE.U+119E3 U+119E4.U+1F44D U+1F3FD.Direct regression tests cover the modified combining classes, USE cluster classification/reordering, emoji grapheme ordering, and explicitly directionless scripts. The Telugu and USE tests fail before their fixes. Validation includes full workspace tests, strict all-features/all-targets Clippy, and Rust 1.85 no-std/libm checks.
These changes are handwritten; the Ragel source and generated machines are unchanged.
With these fixes integrated with the pending Fontations CAPS/extended-layout readers and HarfRust extended-layout branch, all 4,374 multilingual and exhaustive Unicode-batch shaping comparisons match HarfBuzz on the original and Fontations-roundtripped mega font. Before the fixes, the sweep reproduced the reduced cases above.
Test font: https://drive.google.com/file/d/1vMQI9mZMzfKIMHD96Iie40Oyo9NnnweH/view