Repository navigation
perf(binding): compact Layout codec and cover Wasm in ordinary CI - #51
Merged
Merged
Conversation
hyfdev
marked this pull request as ready for review
August 22, 2026 06:02
hyfdev
force-pushed
the
codex/layout-output-transport-01a02590
branch
from
August 22, 2026 14:12
84b0ba6 to
e7b9f63
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Every public Layout read currently makes napi-rs construct one root object, seven nested objects, and 28 properties. This PR keeps the same complete ordinary JavaScript result while moving its 21 numeric values through a reusable private buffer, reducing Native read time by about 90%.
Why
getLayout()andgetUnroundedLayout()return fixed numeric records, but their binding cost is much larger than the Rust layout lookup. The optimization must preserve the existing signatures, complete object shape, errors, special numeric values, and detached-snapshot behavior on every supported runtime.This PR does not add a public typed array, narrow getter, bulk API, or per-call typed-array allocation. Bulk reads remain a separate API decision in #41.
Implementation
The public methods still return the same value:
The binding now writes the 21 Layout numbers into one module-owned
Float64Array, and TypeScript immediately reconstructs a fresh complete object. Rust validates the exact buffer length, writes every slot synchronously, and retains neither the view nor its pointer.api/layout-codec.jsonis the only maintained Layout field inventory. The existing API codegen pipeline emits the public declaration, TypeScript decoder, Rust slot constants, and Rust writer from that model.Every supported runtime shares one write path, but only if the scratch buffer is allocated over an explicit
ArrayBuffer:JavaScriptCore materializes the backing buffer of a length-constructed typed array lazily, and Bun loses the first pointer write into any such buffer. A buffer built by length therefore reports a zero size for a laid-out 120x80 node on its first read and the true size afterwards, on Bun 1.2 and 1.3 alike. The allocation happens once at module load, so the per-call path is unchanged.
The Wasm target fills the same 21 slots through
napi_set_elementand never takes the pointer. Node-API promises thatnapi_get_typedarray_infoyields a pointer into the typed array's own storage, and a Wasm module cannot be given one, because its linear memory cannot address the JavaScript heap. Both targets expose one method name and share one buffer, so the tree wrapper and the generated decoder are runtime-independent, and the Wasm loaders keep the transformation they already had.Slot order
Performance
Base:
ee5c93396e89dbfba2bf36d81e53d93108b96ca0. Node v22.20.0, Linux x64. Medians cover 501 complete public calls over an already laid-out persistent tree, including NodeId validation, the binding call, Rust lookup and write, public object reconstruction, and caller field reads.getUnroundedLayout()getLayout()A second machine reproduces the same shape on every supported runtime. macOS arm64, 500 public calls, base and head measured alternately:
Wasm keeps its previous read time rather than gaining one: a Wasm read crosses the boundary 21 times instead of the 57 Node-API calls it replaces, and at that price the saved calls do not pay for the JavaScript-side reconstruction.
The unchanged read-sensitive Native scenarios also improved:
The complete local benchmark suite passed cross-target output equivalence. No current TaffyJS scenario became faster than
yoga-layout.benchmarks/results/published.jsonremains unchanged.Validation
vp run check:codegenvp run check@taffyjs/nodetestsvp run check:wasmvp run benchmarkThree new native binding tests pin the output buffer contract: a wrong length, a wrong element type, and a non-typed-array all fail, and a successful write leaves no stale slot.
The shared runtime smoke asserts the first read into the scratch buffer before the steady-state read, because a runtime that fails to expose the buffer's backing store loses exactly that write. It was red on Bun before the explicit
ArrayBufferallocation.Ordinary CI gains a Wasm job that runs the whole
check:wasmgraph, because the binding compiles target-specific code forwasm32that no other ordinary job builds or runs..agents/docs/tooling-decisions.mdrecords the revised platform boundary; macOS and the release targets stay at publication time.No benchmark website data was updated.
Related to #41.