Skip to content

Commit e3aedac

Browse files
authored
Merge pull request #31 from hyfdev/agent/benchmark-design
Add Node and Wasm benchmarks
2 parents db6923f + 0bdcf43 commit e3aedac

33 files changed

Lines changed: 2072 additions & 4 deletions

‎.agents/docs/README.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,7 @@
22

33
- [Intent](intent.md) — the package family, its intended users, and the questions deliberately left open during bootstrap.
44
- [Public website](website.md) — the family-level product story, Guide and per-package documentation structure, and ownership of public content.
5+
- [Benchmark design](benchmark.md) — the Node-only public-API measurement boundary, scenario ownership, comparison groups, commands, retained results, and website presentation.
56
- [Architecture](architecture.md) — ownership, package boundaries, test placement and selection, and the complete-snapshot boundary.
67
- [Taffy-to-Node binding mapping](binding-mapping.md) — the repeatable rules for selecting Taffy's high-level Rust surface, mapping it through napi-rs, and preserving safety at the JavaScript boundary.
78
- [Selective query design notes](query-api-design.md) — the agreed high-level model for reading only selected Style, Layout, or DetailedLayoutInfo values without turning the API into a general query language.

‎.agents/docs/architecture.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@ This repository is a small Rust and JavaScript monorepo with one shared napi-rs
1010
- `packages/taffyjs-wasm` compiles the same authored source from `packages/taffyjs-node/src` against generated private Node and browser adapters that share one inline payload. It owns package metadata, generated artifacts, and build-time binding redirection; its package-specific typed generator lives under `tools/taffy-wasm`. It has no authored public wrapper source or separate initialization API.
1111
- `packages/taffyjs-yoga` owns the Node-only ESM Yoga 3.2.1 compatibility facade, Yoga declarations and visible state, input translation, and Yoga-shaped output projection. It depends only on the public `@taffyjs/node` package boundary and does not modify Taffy or the native binding.
1212
- `packages/taffyjs-yoga-wasm` builds the exact facade source owned by `packages/taffyjs-yoga/src` against the public `@taffyjs/wasm` backend. It owns only its package metadata and build-time backend redirection; it has no copied facade source or second compatibility policy.
13+
- `benchmarks` owns the Node-only end-to-end public JavaScript benchmark cases, their TypeScript coordinator, ignored local results, and the retained dataset consumed by the public website. Each scenario owns a direct child directory; the complete boundary is recorded in [Benchmark Design](benchmark.md).
1314
- `tests/taffyjs-node` is a private consumer package that tests `@taffyjs/node` through its package boundary.
1415
- `tests/taffyjs-wasm` is a private consumer package that checks the Wasm package through Node and browser package resolution, reuses the existing public type tests, mechanically inspects the published file set, and installs the packed result in fresh npm and pnpm consumers.
1516
- `tests/taffyjs-yoga` is the corresponding private consumer package for compatibility, differential, and end-to-end tests through the `@taffyjs/yoga` package boundary; it does not contain a redundant nested `tests/` directory.

‎.agents/docs/benchmark.md‎

Lines changed: 73 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,73 @@
1+
# Benchmark Design
2+
3+
[VOUCHED @hyfdev 2026-08-16]
4+
5+
## Purpose and comparison boundary
6+
7+
The public benchmark measures TaffyJS as a JavaScript product in Node.js. Every timed operation starts at a public JavaScript API and includes the complete implementation cost beneath that call. JavaScript validation and conversion, wrapper and compatibility work, Node-API, emnapi, WASI, Wasm memory transfer, callbacks, output construction, redundant computation, and unnecessary data movement all remain part of the result. The benchmark must not bypass a public wrapper, call a private binding, substitute an internal Rust or C++ timer, preconvert data into a private representation, or subtract work in order to make implementations look more alike.
8+
9+
The benchmark has two comparison groups:
10+
11+
- The Taffy API group runs the same public benchmark source against `@taffyjs/node` and `@taffyjs/wasm`.
12+
- The Yoga API group runs the same public Yoga-shaped benchmark source against `@taffyjs/yoga` and `yoga-layout`. A future `@taffyjs/yoga-wasm` target joins this group without changing the scenarios, runner, result format, or page structure.
13+
14+
Only implementations completing a semantically equivalent user transaction receive a speed comparison. Validation happens outside the timed interval and must never normalize the measured path. An unsupported implementation is reported as unsupported rather than receiving a transformed workload. Different API groups do not produce a shared ranking, and the benchmark never combines unrelated scenarios into an overall score or winner.
15+
16+
All official results use one fixed Node.js environment. Browser benchmark execution, browser result switching, and comparisons between Node and browser environments are outside this design.
17+
18+
## Scenarios
19+
20+
Every benchmark case is a scenario with a named question, input and scale, public API operation sequence, timed boundary, required output reads, applicable comparison groups, and completion check. A captured application snapshot, a modeled user transaction, a wide tree, a deep tree, and the same structure at different node counts are equal kinds of scenario on the benchmark page. Each explains what it simulates and what operation is measured; the project does not introduce a separate primary, synthetic, or Diagnostics result hierarchy.
21+
22+
Scenario timing follows the transaction being investigated. An initial-layout scenario may create the public tree, compute layout, and read the outputs the consumer needs. A persistent-tree scenario may mutate an existing tree, recompute it, and read the affected outputs without rebuilding the tree. A batch scenario may create, compute, read or serialize, and finish one complete item. Fixture parsing, result validation, and harness bookkeeping stay outside the timed interval unless they are themselves the public operation named by the scenario. A total is measured directly rather than reconstructed by adding separately aggregated phase results.
23+
24+
The exact initial scenario set remains open until it is researched and reviewed. It must be fixed before the official run whose result will be published; scenarios must not be selected or modified in response to which implementation wins.
25+
26+
## Repository ownership
27+
28+
Benchmark source and retained results belong in the top-level `benchmarks/` workspace rather than under `tests/`, `tools/`, or `apps/website`. Each benchmark case that cannot be combined with another owns a direct child directory. The benchmark workspace owns one private package, one Vite+ configuration, and one explicit TypeScript coordinator; individual scenarios do not receive package manifests or Vite configurations without a concrete execution requirement.
29+
30+
The intended shape is:
31+
32+
```text
33+
benchmarks/
34+
package.json
35+
vite.config.ts
36+
run.ts # serial coordinator and result writer
37+
worker.ts # one isolated target and scenario measurement
38+
suite.ts # explicit settings, scenario registry, and target list
39+
scenario.ts # shared scenario and result contract
40+
<scenario>/
41+
benchmark.ts
42+
fixtures.ts # only when the scenario needs maintained fixture data
43+
results/
44+
local/ # ignored local runs
45+
published.json # retained public result
46+
```
47+
48+
Maintained benchmark code follows the repository-wide TypeScript default. Shared code is introduced only for behavior that scenarios genuinely share; the benchmark must not grow a plugin system or a universal adapter that changes workloads merely to force unlike APIs into one implementation.
49+
50+
## Commands and execution
51+
52+
The root Vite+ task graph exposes two public commands:
53+
54+
- `vp run benchmark` runs local research benchmarks, may accept target or scenario filters, writes only ignored local output, and never changes tracked website data.
55+
- `vp run benchmark:update-website` builds and measures the complete official suite, validates every result, and replaces the retained website dataset only after the complete run succeeds. It does not deploy the website and does not accept a partial suite as a public update.
56+
57+
Both commands use the same TypeScript coordinator and benchmark implementations. Package metadata does not duplicate their orchestration as compound scripts. Timing runs serially because concurrent benchmark processes would interfere with one another. Official measurements use actual release artifacts, a fixed benchmark machine and Node.js version, repeated independent samples, and recorded repository and package revisions. The exact sampling and summary statistic remain open until the runner design is selected.
58+
59+
## Results and website
60+
61+
Local measurements belong under ignored `benchmarks/results/local/`. The website reads the tracked `benchmarks/results/published.json` directly at build time; it does not run benchmarks, copy a second canonical result into `apps/website`, or combine an old measurement with the current scenario source.
62+
63+
The published dataset is a self-contained snapshot of the facts needed to interpret its numbers, including the scenario name and parameters, measured transaction, targets, results, environment, and TaffyJS revision. It retains enough sample evidence to support the displayed comparison. Its concrete schema and file encoding remain implementation details. The dataset is committed because it is evidence for a public performance claim, not a distributable build artifact.
64+
65+
The benchmark page presents every scenario at the same level. Each applicable scenario shows a Taffy API comparison and, separately, a Yoga API comparison. It does not display browser results, a Diagnostics section, an overall score, an original-data download interface, or an empty column for a package that does not exist. Target columns come from the result's target list, so adding `@taffyjs/yoga-wasm` later is a target addition rather than a page redesign.
66+
67+
The page footer shows only the concise environment and source identity needed by a reader, including the Node.js version, operating system and benchmark machine, and TaffyJS commit.
68+
69+
## Delivery sequence
70+
71+
First select and document the initial scenarios and their exact transactions. Then implement the top-level benchmark workspace and the two comparison groups, prove completion and semantic equivalence outside timing, run the complete suite on the fixed Node.js benchmark host, update the retained dataset, and let the website render it. Before publication, review the result for comparison validity, measurement stability, and avoidable harness complexity.
72+
73+
The current delivery implements only the Taffy API comparison between `@taffyjs/node` and `@taffyjs/wasm`. The Yoga API comparison, captured Yoga application trees, and `@taffyjs/yoga-wasm` remain follow-up work after the relevant package pull requests merge. The current work may establish the shared result boundary, but it must not publish empty Yoga targets or manufacture a compatibility implementation in advance.

‎.agents/docs/taffyjs-node-decisions.md‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -174,6 +174,16 @@ New public state owners, compatibility layers, retained JavaScript values, callb
174174

175175
## Decided
176176

177+
### Native distribution targets
178+
179+
**Ruling:** `@taffyjs/node` must provide native packages for macOS arm64, Linux x64 GNU, and Windows x64 MSVC.
180+
181+
**Limits:** This adds `aarch64-apple-darwin` to the existing distribution model. It does not add macOS x64, Linux arm64, Linux musl, or another target. Each additional target still requires its own package metadata, artifact synchronization, public support documentation, and CI coverage.
182+
183+
**Why:** Yunfei required the Apple M3 Pro benchmark host to be supported as a real native target rather than bypassing the platform-package build step.
184+
185+
**Source:** Yunfei (`@hyfdev`), 2026-08-16; explicitly asked to support the current `darwin/arm64` host before continuing the benchmark website work.
186+
177187
### Generated numeric input shorthand
178188

179189
[VOUCHED @hyfdev 2026-08-15]

‎.agents/docs/technology-stack.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ napi-rs owns the Rust-to-Node boundary, private native declarations, native load
66

77
The generated ESM loader is a private build input. The authored wrapper imports it privately, and the bundled public entry does not re-export raw native operations.
88

9-
The root package has exact-version optional packages for the two supported targets: `@taffyjs/binding-linux-x64-gnu` and `@taffyjs/binding-win32-x64-msvc`. Publication is not configured.
9+
The root package has exact-version optional packages for the three supported targets: `@taffyjs/binding-darwin-arm64`, `@taffyjs/binding-linux-x64-gnu`, and `@taffyjs/binding-win32-x64-msvc`. Publication is not configured.
1010

1111
The bigint NodeId marker and its JavaScript validity registry require the authored wrapper, but not a custom loader or another package.
1212

@@ -30,7 +30,7 @@ For `packages/taffyjs-yoga-wasm`, Vite+ builds those same sibling source entries
3030

3131
`tools/api-codegen` owns source generation that must keep Rust and TypeScript API facts aligned. Its first maintained input is `api/numeric-families.json`; `vp run codegen` updates both language outputs. CI runs `vp run check:codegen`, which regenerates and rejects any resulting Git diff.
3232

33-
CI has five jobs. Ubuntu x64 and Windows x64 each build the native addon and run all Rust, JavaScript, and type tests with Node.js 22.18.0, then smoke-test the public `@taffyjs/node` and `@taffyjs/yoga` entries with Bun 1.2.0 and Deno 2.2.0; Ubuntu also rejects stale committed package JavaScript and declarations after the build. A separate Ubuntu WASIP job installs the Rust target and Playwright Chromium, builds `@taffyjs/wasm`, reruns the complete Node public API suite against it, runs type, package-content, packed-consumer, bundled-consumer, and browser-runtime checks against the generated package, and applies the same Bun 1.2.0 and Deno 2.2.0 smoke checks to the public entry. That job also builds `@taffyjs/yoga-wasm`, reruns the complete maintained Yoga behavior and declaration suites, inspects and installs the packed package with npm and pnpm, and exercises both public entries in bundled Chromium. Alternate-runtime smoke checks perform one fixed layout rather than copying the complete behavior suite. A Node-only job checks formatting, JavaScript and repository TypeScript including maintained tools through Vite+'s type-aware lint path, and generated-source drift. A Rust-only job checks formatting and Clippy. macOS and publication workflows are not configured.
33+
CI has six jobs. Ubuntu x64, Windows x64, and macOS arm64 each build the native addon and run all Rust, JavaScript, and type tests with Node.js 22.18.0; Ubuntu and Windows then smoke-test the public `@taffyjs/node` and `@taffyjs/yoga` entries with Bun 1.2.0 and Deno 2.2.0. Ubuntu also rejects stale committed package JavaScript and declarations after the build. A separate Ubuntu WASIP job installs the Rust target and Playwright Chromium, builds `@taffyjs/wasm`, reruns the complete Node public API suite against it, runs type, package-content, packed-consumer, bundled-consumer, and browser-runtime checks against the generated package, and applies the same Bun 1.2.0 and Deno 2.2.0 smoke checks to the public entry. That job also builds `@taffyjs/yoga-wasm`, reruns the complete maintained Yoga behavior and declaration suites, inspects and installs the packed package with npm and pnpm, and exercises both public entries in bundled Chromium. Alternate-runtime smoke checks perform one fixed layout rather than copying the complete behavior suite. A Node-only job checks formatting, JavaScript and repository TypeScript including maintained tools through Vite+'s type-aware lint path, and generated-source drift. A Rust-only job checks formatting and Clippy. Publication workflows are not configured.
3434

3535
## Public website
3636

‎.agents/docs/tooling-decisions.md‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -86,6 +86,18 @@ This ledger records only tooling judgments that Yunfei explicitly expressed for
8686

8787
**Source:** Yunfei (`@hyfdev`), 2026-08-09; explicitly selected the [vue-tui Vite+ configuration](https://github.com/vuejs-ai/vue-tui/blob/main/vite.config.ts) as the orchestration reference and requested `run.cache: false`.
8888

89+
### Benchmark ownership and command surface
90+
91+
[VOUCHED @hyfdev 2026-08-16]
92+
93+
**Ruling:** Maintained benchmark cases and results must belong to the top-level `benchmarks/` workspace, and the root Vite+ task graph must expose `vp run benchmark` for non-mutating local runs and `vp run benchmark:update-website` for a complete run that updates the retained website data.
94+
95+
**Limits:** This ruling fixes ownership and the two public commands, not the internal coordinator API, exact scenario set, sampling library, result schema, fixed benchmark machine, or Node.js version. Internal selection remains an argument to the local command rather than a larger public command family. Updating benchmark data does not deploy the website.
96+
97+
**Why:** Yunfei wanted benchmark examples to be first-class top-level cases rather than unrelated scripts, distinguished ordinary measurement from measurement that updates documentation data, and explicitly corrected the updating command to `benchmark:update-website`; no additional rationale was stated.
98+
99+
**Source:** Yunfei (`@hyfdev`), 2026-08-16; accepted and vouched the complete benchmark design and explicitly named the two public commands. See [Benchmark Design](benchmark.md).
100+
89101
### Typed tools with scope-revealing ownership
90102

91103
[VOUCHED @hyfdev 2026-08-16]

0 commit comments

Comments
 (0)