Skip to content

[ot] Share contextual rule matchers across class-cache modes - #502

Open
behdad wants to merge 1 commit into
feat/extended-layoutfrom
perf/share-contextual-matching
Open

behdad wants to merge 1 commit into
feat/extended-layoutfrom
perf/share-contextual-matching

Conversation

@behdad

@behdad behdad commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Stacked on #495; this PR contains only the contextual-matcher sharing change.

Replace matcher/first-value closure types with concrete borrowed glyph/class
matchers. This keeps the existing parsing, rule probes, malformed-rule fallback,
and cache behavior, but stops duplicating the large matching loops for every
closure/cache mode. No allocations or trait-object dispatch are added.

In the measured release/LTO build, the large chained-rule matcher goes from
eight instances to three, saving about 52 KiB of ELF text in both hr-shape
and the C API library.

Performance tradeoff

The Nastaliq sample is about 6–8% slower. This is an intentional size/speed
tradeoff, not a performance-neutral refactor. The first measurement was +6.1%;
after rebasing onto the cursive-sharing change (#498), a fresh comparison measured
+8.4%. A partially specialized prototype removed the initial regression but saved
only about 24 KiB; we chose the larger saving.

CPU-pinned, alternating before/after runs on Linux x86-64, Rust 1.89, release
optimization with fat LTO and one codegen unit. The latest comparison against the
base including #498 uses eight samples per binary:

Sample Shaping time change
Amiri -0.1%
Noto Nastaliq Urdu +8.4%
Noto Naskh Arabic +0.2%
Noto Sans Devanagari -0.8%
Noto Sans Bengali -0.9%
Noto Sans Latin +1.1%

These are individual text samples, not a claim about all fonts or workloads.
Glyphs, positions, and flags match the baseline in all four directions for the
six samples.

Validation

  • Workspace all-feature tests, including 6073 shaping cases.
  • Direct regression tests for wide glyph/class values, cache sentinels, and
    preservation of the other cached nibble.
  • Strict workspace/all-target/all-feature Clippy and rustfmt.
  • Rust 1.85 no-std build for thumbv7em-none-eabihf with libm.

Matcher and first-value closures multiply the large contextual rule
loops for glyph matching and each class-cache mode. Replace them with
borrowed glyph/class matchers while retaining typed rule parsing and
the existing cache sentinels, probes, and malformed-rule handling.

This reduces chained-rule matcher instances from eight to three and
saves about 52 KiB of text in release/LTO shaping binaries. The measured
tradeoff is about 6% slower Nastaliq shaping; other samples remain close
to the baseline.

Add direct tests for wide glyph/class values and class-cache sentinels
and nibble preservation. Workspace all-feature tests (including 6073
shaping cases), strict Clippy, formatting, and Rust 1.85 no-std ARM
build pass.

Assisted-by: OpenAI Codex
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant