Skip to content

perf(binding): compact Layout codec and cover Wasm in ordinary CI - #51

Merged
hyfdev merged 1 commit into
mainfrom
codex/layout-output-transport-01a02590
Aug 22, 2026
Merged

hyfdev merged 1 commit into
mainfrom
codex/layout-output-transport-01a02590

Conversation

@hyfdev

@hyfdev hyfdev commented Aug 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

Every public Layout read currently makes napi-rs construct one root object, seven nested objects, and 28 properties. This PR keeps the same complete ordinary JavaScript result while moving its 21 numeric values through a reusable private buffer, reducing Native read time by about 90%.

Why

getLayout() and getUnroundedLayout() return fixed numeric records, but their binding cost is much larger than the Rust layout lookup. The optimization must preserve the existing signatures, complete object shape, errors, special numeric values, and detached-snapshot behavior on every supported runtime.

This PR does not add a public typed array, narrow getter, bulk API, or per-call typed-array allocation. Bulk reads remain a separate API decision in #41.

Implementation

The public methods still return the same value:

const layout = tree.getLayout(node);

layout.size.width;
layout.margin.left;

The binding now writes the 21 Layout numbers into one module-owned Float64Array, and TypeScript immediately reconstructs a fresh complete object. Rust validates the exact buffer length, writes every slot synchronously, and retains neither the view nor its pointer.

api/layout-codec.json is the only maintained Layout field inventory. The existing API codegen pipeline emits the public declaration, TypeScript decoder, Rust slot constants, and Rust writer from that model.

Every supported runtime shares one write path, but only if the scratch buffer is allocated over an explicit ArrayBuffer:

const layoutCodecBuffer = new Float64Array(new ArrayBuffer(layoutCodecByteLength));

JavaScriptCore materializes the backing buffer of a length-constructed typed array lazily, and Bun loses the first pointer write into any such buffer. A buffer built by length therefore reports a zero size for a laid-out 120x80 node on its first read and the true size afterwards, on Bun 1.2 and 1.3 alike. The allocation happens once at module load, so the per-call path is unchanged.

The Wasm target fills the same 21 slots through napi_set_element and never takes the pointer. Node-API promises that napi_get_typedarray_info yields a pointer into the typed array's own storage, and a Wasm module cannot be given one, because its linear memory cannot address the JavaScript heap. Both targets expose one method name and share one buffer, so the tree wrapper and the generated decoder are runtime-independent, and the Wasm loaders keep the transformation they already had.

Slot order

0 order
1-2 location.x/y
3-4 size.width/height
5-6 contentSize.width/height
7-8 scrollbarSize.width/height
9-12 border.left/right/top/bottom
13-16 padding.left/right/top/bottom
17-20 margin.left/right/top/bottom

Performance

Base: ee5c93396e89dbfba2bf36d81e53d93108b96ca0. Node v22.20.0, Linux x64. Medians cover 501 complete public calls over an already laid-out persistent tree, including NodeId validation, the binding call, Rust lookup and write, public object reconstruction, and caller field reads.

Public operation Before After Reduction
Native getUnroundedLayout() 1.359 ms 0.136 ms 90.0%
Native getLayout() 1.328 ms 0.133 ms 90.0%

A second machine reproduces the same shape on every supported runtime. macOS arm64, 500 public calls, base and head measured alternately:

Target and runtime Before After
Native, Node v24.19.0 0.960–0.999 ms 0.113–0.116 ms
Native, Bun 1.3.9 0.945 ms 0.115 ms
Wasm, Node v24.19.0 8.12–8.48 ms 8.03–8.43 ms

Wasm keeps its previous read time rather than gaining one: a Wasm read crosses the boundary 21 times instead of the 57 Node-API calls it replaces, and at that price the saved calls do not pay for the JavaScript-side reconstruction.

The unchanged read-sensitive Native scenarios also improved:

Scenario Before After
wide wrapping collection, 500 items 3.8104 ms 3.1295 ms (-17.9%)
coding-agent chat viewport resize 2.2987 ms 1.8798 ms (-18.2%)

The complete local benchmark suite passed cross-target output equivalence. No current TaffyJS scenario became faster than yoga-layout. benchmarks/results/published.json remains unchanged.

Validation

  • vp run check:codegen
  • vp run check
    • formatting, lint, release and platform checks, Rust fmt, and Clippy
    • 11 Rust tests
    • 48 native binding compatibility tests
    • 217 public @taffyjs/node tests
    • type tests and 557 Yoga facade tests
  • vp run check:wasm
    • the same 217 public tests on WASI
    • browser runtime and bundle, package contents, public types, packed npm/pnpm consumers, Yoga-WASI, and website build
  • Native and WASI smoke tests on Node, Bun 1.2.0, Bun 1.3.9, and Deno 2.2
  • vp run benchmark

Three new native binding tests pin the output buffer contract: a wrong length, a wrong element type, and a non-typed-array all fail, and a successful write leaves no stale slot.

The shared runtime smoke asserts the first read into the scratch buffer before the steady-state read, because a runtime that fails to expose the buffer's backing store loses exactly that write. It was red on Bun before the explicit ArrayBuffer allocation.

Ordinary CI gains a Wasm job that runs the whole check:wasm graph, because the binding compiles target-specific code for wasm32 that no other ordinary job builds or runs. .agents/docs/tooling-decisions.md records the revised platform boundary; macOS and the release targets stay at publication time.

No benchmark website data was updated.

Related to #41.

@hyfdev
hyfdev marked this pull request as ready for review August 22, 2026 06:02
Copilot AI lite review requested due to automatic review settings August 22, 2026 06:02

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hyfdev hyfdev changed the title perf(binding): compact Layout transport perf(binding): compact Layout transport and cover Wasm in ordinary CI Aug 22, 2026
@hyfdev
hyfdev force-pushed the codex/layout-output-transport-01a02590 branch from 84b0ba6 to e7b9f63 Compare August 22, 2026 14:12
@hyfdev hyfdev changed the title perf(binding): compact Layout transport and cover Wasm in ordinary CI perf(binding): compact Layout codec and cover Wasm in ordinary CI Aug 22, 2026
@hyfdev
hyfdev merged commit 8e2b99b into main Aug 22, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants