Skip to content

Support ISO OFF expanded GSUB, GPOS, and GDEF shaping - #495

Open
behdad wants to merge 27 commits into
mainfrom
feat/extended-layout
Open

behdad wants to merge 27 commits into
mainfrom
feat/extended-layout

Conversation

@behdad

@behdad behdad commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Implement expanded GSUB/GPOS/GDEF shaping from ISO Open Font Format fifth edition (ISO/IEC 14496-22:2026), using the ISO spec for field sizes and HarfBuzz as an implementation reference.

Implemented in small tested commits.

Implemented:

  • GSUB/GPOS 1.2 lookup-list selection and full-width lookup offsets.
  • GDEF 1.4 class and mark-filtering table selection.
  • Coverage and ClassDef formats 3/4 in shared helpers, digests, and GDEF caches.
  • Full-width coverage indices/classes, retaining the existing compact caches and bypassing them for unrepresentable values.
  • SingleSubst formats 3/4, including signed 24-bit wrapping and full-width replacement indices.
  • MultipleSubst, AlternateSubst, and ReverseChainSingleSubst format 2 application.
  • LigatureSubst format 2, preserving bounded digest caching and ligature bookkeeping.
  • SinglePos formats 3/4 and CursivePos format 2, using the existing value and attachment evaluators.
  • PairPos formats 3/4 with full-width glyph-pair and class caches, plus a separate fix preventing invalid class indices from selecting another matrix row.
  • MarkBasePos, MarkMarkPos, and MarkLigPos format 2, preserving attachment selection and testing their different count and offset widths.
  • SequenceContext formats 4/5/6 with actual nested GSUB/GPOS lookup tests, preserving optimized rule matching and testing cached/uncached class handling.
  • ChainedSequenceContext formats 4/5 with nested GSUB/GPOS tests, preserving optimized glyph/class matching and cached/uncached behavior.
  • Feature variation conditions 1–5 using the shared Fontations evaluator, with GDEF deltas and bounded evaluation.
  • LookupVariations application in GSUB/GPOS, including lookup-set union/deduplication, alternate-feature defaults, empty replacement sets, and shape-plan cache compatibility.
  • Regression tests for wide counts, glyph IDs/classes, offset precedence, null fallback, and malformed data.

The extended-layout shaping formats and LookupVariations consumer are implemented. The binary reader/writer support for LookupVariations comes from googlefonts/fontations#2216. Closure/subsetting is implemented separately in Fontations PR #2218 (googlefonts/fontations#2218).

Depends on the read-fonts APIs in googlefonts/fontations#2217. The dependency is temporarily pinned to that branch's tested revision; it should return to a crates.io release before landing.

Validation so far: full workspace tests (including 6,069 shaping regressions), strict all-features/all-targets workspace Clippy, Rust 1.85 build, and thumbv7em no-std/libm build pass.

Test font: https://drive.google.com/file/d/1vMQI9mZMzfKIMHD96Iie40Oyo9NnnweH/view

Mega Noto validation (104,404 glyphs): all 4,374 comparisons match HarfBuzz, including 26 multilingual samples, four directions, four feature configurations covering all 82 feature tags, three scales, and exhaustive Unicode batches. Both the original font and the Fontations-repacked font are compared. This uses a temporary integration of pending Fontations feature branches (including CAPS reader/metrics support in googlefonts/fontations#2208) and the mark/grapheme-ordering sync fixes from #496, now landed on main. The integration branches are for testing only; these PRs remain separate. Full combined-workspace tests and strict Clippy pass.

Additional regression coverage actually reuses cached shape plans while switching between positive, negative, and default locations. Reused GSUB/GPOS plans match fresh-plan glyphs and positions; coordinates resolving to the same lookup set share a plan. The full workspace suite, strict Clippy, and formatting pass.

Rebased onto main after #496 landed. Full all-feature workspace tests,
strict workspace Clippy, Rust 1.85 no-std/libm checks, and formatting pass.
A new end-to-end Skera subset sweep covers 26,767 mapped characters in
287 batches across 144 scripts. HarfRust and HarfBuzz agree for every
batch on the original, full subset, and all-feature/no-hints subset.
The Fontations integration uses #2218 plus CAPS subsetting in
googlefonts/fontations#2225; all 104,404 glyph
metrics and retained outline data are preserved, with FontTools round trips.

@behdad
behdad requested a review from dfrg October 4, 2026 17:54
@behdad
behdad marked this pull request as ready for review October 4, 2026 19:33
behdad added 26 commits October 4, 2026 21:02
Use the read-fonts expanded-layout implementation from Fontations #2213
while it awaits release. Resolve lookup and GDEF offsets through the
selected table, preserving wide-offset precedence and null fallback.

Support Coverage and ClassDef formats 3/4 in shared accessors, digests,
mark filtering, and cached GDEF properties. Keep coverage indices and
class values full-width; values outside the compact caches bypass them.
Legacy lookup application remains supported while expanded subtables
will be added in follow-up commits.

Add width-boundary, large-offset, null-fallback, and malformed-data tests.
The workspace test suite and Rust 1.85/std-free builds pass.

Reference: googlefonts/fontations#2213
Assisted-by: OpenAI Codex
Current derive-macro dependencies use both syn 2 and 3. Allow this
build-time duplication without disabling checks on other crate names
or changing runtime dependencies to satisfy the lint.

Tested with strict all-features/all-targets workspace Clippy.

Assisted-by: OpenAI Codex
Dispatch the expanded single-substitution formats through the lookup
cache and would-apply checks. Format 3 wraps signed deltas in 24 bits;
format 4 keeps full-width coverage indices and replacement glyph IDs.

Add application tests for signed wrapping, high coverage indices,
unmatched glyphs, and extension lookup dispatch.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share the existing multiple-substitution implementation with format 2
to preserve expansion, deletion, and ligature-component behavior.
Dispatch the expanded format with 24-bit sequence offsets and glyph IDs.

Add regressions for wide output glyphs, component bookkeeping,
single-glyph replacements, deletion, and offsets above 64K.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share alternate selection between the narrow and expanded formats,
including feature values, out-of-range handling, and random selection.
Wire format 2 through lookup dispatch and would-apply checks.

Add application tests for high glyph IDs and 24-bit set offsets.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share reverse-chain application with the expanded format, retaining
backtrack/lookahead matching and the no-nested-application guard.
Dispatch wide replacement IDs and 24-bit coverage offsets through
ordinary and extension lookups.

Add regressions for wide context glyphs, failed matches, extension
dispatch, and right-to-left application order.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share ligature matching and application with the expanded format,
retaining fast-path filtering, mark skipping, and concat hazards.
Reuse the compact caches without narrowing coverage offsets, and keep
the optional second-component digest walk bounded for wide tables.

Test wide components/results, large ligature offsets and set counts,
high coverage indices, and the small-cache and traversal-budget paths.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Dispatch expanded single positioning through the existing value-record
evaluator. Keep format 4 record counts and coverage indices full-width,
and reject an index outside the declared value array.

Test scaled horizontal/vertical placement and advances, high record
indices, extension dispatch, and invalid coverage indices.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share the existing cursive attachment algorithm with format 2, including
direction-dependent advances, attachment chains, and safety checks.
Resolve 24-bit nullable entry/exit anchors for wide coverage glyphs.

Test all four writing directions, lookup attachment direction, offsets
above 64K, and null anchors.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
PairPos format 2 manually indexes its class matrix without checking the
class counts. An invalid second class can therefore alias a valid
record in the next row instead of failing to match.

Check both class indices before indexing and use checked arithmetic for
the matrix offset, including on 32-bit targets. Add a regression that
previously applied a 100-unit advance from the wrong row, plus a wide
class that cannot fit the declared matrix.

Testing: reproduced the wrong advance before the fix; full workspace
tests and strict workspace Clippy pass with the fix.

Assisted-by: OpenAI Codex
Share class-matrix application with the expanded format, including
value evaluation, mark skipping, and the class bounds checks. Keep
coverage and ClassDef offsets full-width in the small cache.

Test wide glyph IDs, large offsets, both cache layouts, and classes
that must not truncate to valid matrix indices.

Testing: full workspace tests, strict workspace Clippy, Rust 1.85
workspace build, and thumbv7em no-std/libm build.

Assisted-by: OpenAI Codex
Share pair application with format 3 using 24-bit second glyphs, pair
counts, and set offsets. Preserve value-record offset bases and concat
hazards, and keep coverage offsets full-width in the small cache.

Bound the optional wide pair-digest walk and reuse its result for the
lookup's second-glyph digest. Large tables still use binary search.

Test wide pair/set counts, coverage indices, high offsets, both value
records, and cache fallback beyond the traversal budget.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share base selection with format 2 and extract the existing stateless
mark-attachment helper for both mark-array encodings. Preserve mark
skipping, scaling, cross offsets, and attachment-chain checks.

Test wide mark/base counts and coverage indices, large array and anchor
offsets, null base anchors, and skipped marks.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Remove the unnecessary module qualification from the skipped-mark
regression, satisfying the strict unused-qualifications lint.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share mark selection and ligature-component matching with format 2,
using the common attachment helper and expanded mark/anchor arrays.
Keep Mark2Array2's count 16-bit as specified by ISO OFF.

Test high glyph IDs, large offsets, component matching, and a wide
mark2 coverage index that must not alias the first record.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share ligature selection and component matching with format 2, using
24-bit array/attachment/anchor offsets and the common mark helper.
Retain 16-bit component counts and the existing last-component fallback.

Test component selection, wide ligature counts and coverage indices,
large offsets from all three bases, and empty component arrays.

Testing: full workspace tests and strict workspace Clippy.

Assisted-by: OpenAI Codex
Share coverage-based context matching with format 3, retaining nested
lookup application and unsafe-to-break/concat handling. Dispatch format 6
in both GSUB and GPOS and use its coverage for lookup filtering.

Test actual nested substitutions and positioning with full-width glyph
IDs, 24-bit coverage offsets, and unmatched input sequences.

Tested with the full workspace suite and strict all-targets Clippy.

Assisted-by: OpenAI Codex
Parameterize the optimized context rule parser and pre-match probes for
24-bit input glyph IDs. Share application and would-apply behavior with
format 1 without changing the legacy fast path or boundary handling.
SequenceRuleSet2's rule offsets remain 16-bit, as specified by ISO OFF.

Test nested GSUB and GPOS lookups, 24-bit rule-set offsets and counts,
coverage indices above 65535, failed-rule skipping, and malformed data.

Tested with the full workspace suite and strict all-targets Clippy.

Assisted-by: OpenAI Codex
Share class-context matching with format 2 and support 24-bit rule
offsets and rule-set counts plus 32-bit coverage and ClassDef offsets.
Keep classes within ClassSequenceRule at their specified 16-bit width.

Enable the existing hot class cache for format 5, bypassing its compact
entries for larger classes. Test actual nested GSUB/GPOS application with
and without the hot cache, wide offsets and counts, failed-rule skipping,
and classes that must not alias their low 16 bits.

Tested with the full workspace suite and strict all-targets Clippy.

Assisted-by: OpenAI Codex
Extend the optimized chained-rule parser and pre-match probes to 24-bit
glyph IDs and offsets. Share format 1's matching, nested lookup behavior,
and unsafe boundary handling without changing its legacy specialization.

Test nested GSUB/GPOS application with wide backtrack, input and lookahead
glyphs, high coverage indices and rule offsets, failed-rule skipping, and
single-glyph inputs whose pre-match probes use the lookahead sequence.

Tested with the full workspace suite and strict all-targets Clippy.

Assisted-by: OpenAI Codex
Share format 2's optimized chained class matching and caches with format
5. Retain 16-bit classes and the 16-bit rule-set count while resolving
32-bit coverage/ClassDef offsets and 24-bit rule and rule-set offsets.

Test nested GSUB/GPOS lookup application with separate backtrack, input
and lookahead classes, cached and uncached matching, high class indices,
large offsets, single-glyph inputs, and rejection of truncated aliases.

Tested with the full workspace suite and strict all-targets Clippy.

Assisted-by: OpenAI Codex
Use shared condition evaluation when preparing shaping data and matching
shape-plan keys. Support variable values and AND/OR/NOT conditions in
addition to axis ranges, including default locations and null conditions.

Preserve fractional variation deltas and the strict positive-value test.
Bound recursion and repeated condition visits, and propagate malformed
conditions without allowing negation to turn an error into a match.

Test all condition formats, endpoints, missing axes, fractional deltas,
null children, malformed data, and recursive/repeated condition trees.
The full workspace suite and strict all-targets Clippy pass.

Assisted-by: OpenAI Codex
Resolve legacy FeatureVariations with read-fonts' shared condition
evaluator and supply GDEF variation deltas. Remove the duplicate tree
evaluator and its tests, which now live in Fontations.

Update the pinned Fontations revision to include the shared evaluator.
The full workspace tests and strict all-feature Clippy pass.

Assisted-by: OpenAI Codex
Resolve FeatureVariations 1.1 lookup records independently of the legacy
feature-variation selection. Evaluate every condition with read-fonts,
union and sort matching lookup sets, and take ADD_DEFAULT_LOOKUPS from
the current feature, including any selected alternate.

Use the resolved sets when compiling GSUB/GPOS lookups. Preserve the
difference between an absent record and an empty replacement set, and
include resolved lookup state in shape-plan keys so coordinate changes
cannot reuse an incompatible plan.

Add synthetic GSUB/GPOS shaping tests for defaults, overlapping records,
duplicates, alternate features, fractional GDEF deltas, malformed data,
plan reuse, and full-width condition counts and offsets.

The full workspace tests (including 6,069 shaping regressions), strict
all-feature Clippy, formatting, Rust 1.85 workspace checks, and the
thumbv7em no-std/libm build pass.

Assisted-by: OpenAI Codex
Use the tested Fontations revision containing default-location feature
selection and coordinate-aware conditional autohint lookups. Keep this
temporary git dependency until the reader APIs are released.

Full workspace tests, strict Clippy, and the Rust 1.85 no-std/libm
check pass.

Assisted-by: OpenAI Codex
Cache plans by resolved lookup sets and reuse them while moving between
positive, negative, and default locations. Compare GSUB glyphs and GPOS
positions with freshly constructed plans, including empty default coords.
Ensure that different coordinates resolving to the same set share a plan.

Tested with the full workspace suite, strict all-feature Clippy, and fmt.

Assisted-by: OpenAI Codex
Keep the format-specific coverage and anchor-offset readers in their
wrappers, and share anchor resolution and attachment adjustment through
a non-generic, non-inlined helper. Reuse child and parent position
references instead of repeating bounds-checked indexing.

Exercise both CursivePos formats in the all-direction attachment test,
including the right-to-left lookup flag. No allocations or trait-object
dispatch are added.

On x86-64 Rust 1.89 release/fat-LTO builds, ELF text plus data shrinks by
920 bytes for hr-shape and 928 bytes for the C API. Stripped file sizes
are unchanged because of segment alignment. Alternating pinned shaping
benchmarks on Arabic, Urdu and Latin samples remain within about 1%.

Tested with the complete all-feature workspace suite (including 6073
shaping regressions), strict Clippy, formatting, and a no-std/libm check.
Shaping output and flags match the baseline in all four directions.

Assisted-by: OpenAI Codex
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant