|
| 1 | +# Benchmark Design |
| 2 | + |
| 3 | +[VOUCHED @hyfdev 2026-08-16] |
| 4 | + |
| 5 | +## Purpose and comparison boundary |
| 6 | + |
| 7 | +The public benchmark measures TaffyJS as a JavaScript product in Node.js. Every timed operation starts at a public JavaScript API and includes the complete implementation cost beneath that call. JavaScript validation and conversion, wrapper and compatibility work, Node-API, emnapi, WASI, Wasm memory transfer, callbacks, output construction, redundant computation, and unnecessary data movement all remain part of the result. The benchmark must not bypass a public wrapper, call a private binding, substitute an internal Rust or C++ timer, preconvert data into a private representation, or subtract work in order to make implementations look more alike. |
| 8 | + |
| 9 | +The benchmark has two comparison groups: |
| 10 | + |
| 11 | +- The Taffy API group runs the same public benchmark source against `@taffyjs/node` and `@taffyjs/wasm`. |
| 12 | +- The Yoga API group runs the same public Yoga-shaped benchmark source against `@taffyjs/yoga` and `yoga-layout`. A future `@taffyjs/yoga-wasm` target joins this group without changing the scenarios, runner, result format, or page structure. |
| 13 | + |
| 14 | +Only implementations completing a semantically equivalent user transaction receive a speed comparison. Validation happens outside the timed interval and must never normalize the measured path. An unsupported implementation is reported as unsupported rather than receiving a transformed workload. Different API groups do not produce a shared ranking, and the benchmark never combines unrelated scenarios into an overall score or winner. |
| 15 | + |
| 16 | +All official results use one fixed Node.js environment. Browser benchmark execution, browser result switching, and comparisons between Node and browser environments are outside this design. |
| 17 | + |
| 18 | +## Scenarios |
| 19 | + |
| 20 | +Every benchmark case is a scenario with a named question, input and scale, public API operation sequence, timed boundary, required output reads, applicable comparison groups, and completion check. A captured application snapshot, a modeled user transaction, a wide tree, a deep tree, and the same structure at different node counts are equal kinds of scenario on the benchmark page. Each explains what it simulates and what operation is measured; the project does not introduce a separate primary, synthetic, or Diagnostics result hierarchy. |
| 21 | + |
| 22 | +Scenario timing follows the transaction being investigated. An initial-layout scenario may create the public tree, compute layout, and read the outputs the consumer needs. A persistent-tree scenario may mutate an existing tree, recompute it, and read the affected outputs without rebuilding the tree. A batch scenario may create, compute, read or serialize, and finish one complete item. Fixture parsing, result validation, and harness bookkeeping stay outside the timed interval unless they are themselves the public operation named by the scenario. A total is measured directly rather than reconstructed by adding separately aggregated phase results. |
| 23 | + |
| 24 | +The exact initial scenario set remains open until it is researched and reviewed. It must be fixed before the official run whose result will be published; scenarios must not be selected or modified in response to which implementation wins. |
| 25 | + |
| 26 | +## Repository ownership |
| 27 | + |
| 28 | +Benchmark source and retained results belong in the top-level `benchmarks/` workspace rather than under `tests/`, `tools/`, or `apps/website`. Each benchmark case that cannot be combined with another owns a direct child directory. The benchmark workspace owns one private package, one Vite+ configuration, and one explicit TypeScript coordinator; individual scenarios do not receive package manifests or Vite configurations without a concrete execution requirement. |
| 29 | + |
| 30 | +The intended shape is: |
| 31 | + |
| 32 | +```text |
| 33 | +benchmarks/ |
| 34 | + package.json |
| 35 | + vite.config.ts |
| 36 | + run.ts # serial coordinator and result writer |
| 37 | + worker.ts # one isolated target and scenario measurement |
| 38 | + suite.ts # explicit settings, scenario registry, and target list |
| 39 | + scenario.ts # shared scenario and result contract |
| 40 | + <scenario>/ |
| 41 | + benchmark.ts |
| 42 | + fixtures.ts # only when the scenario needs maintained fixture data |
| 43 | + results/ |
| 44 | + local/ # ignored local runs |
| 45 | + published.json # retained public result |
| 46 | +``` |
| 47 | + |
| 48 | +Maintained benchmark code follows the repository-wide TypeScript default. Shared code is introduced only for behavior that scenarios genuinely share; the benchmark must not grow a plugin system or a universal adapter that changes workloads merely to force unlike APIs into one implementation. |
| 49 | + |
| 50 | +## Commands and execution |
| 51 | + |
| 52 | +The root Vite+ task graph exposes two public commands: |
| 53 | + |
| 54 | +- `vp run benchmark` runs local research benchmarks, may accept target or scenario filters, writes only ignored local output, and never changes tracked website data. |
| 55 | +- `vp run benchmark:update-website` builds and measures the complete official suite, validates every result, and replaces the retained website dataset only after the complete run succeeds. It does not deploy the website and does not accept a partial suite as a public update. |
| 56 | + |
| 57 | +Both commands use the same TypeScript coordinator and benchmark implementations. Package metadata does not duplicate their orchestration as compound scripts. Timing runs serially because concurrent benchmark processes would interfere with one another. Official measurements use actual release artifacts, a fixed benchmark machine and Node.js version, repeated independent samples, and recorded repository and package revisions. The exact sampling and summary statistic remain open until the runner design is selected. |
| 58 | + |
| 59 | +## Results and website |
| 60 | + |
| 61 | +Local measurements belong under ignored `benchmarks/results/local/`. The website reads the tracked `benchmarks/results/published.json` directly at build time; it does not run benchmarks, copy a second canonical result into `apps/website`, or combine an old measurement with the current scenario source. |
| 62 | + |
| 63 | +The published dataset is a self-contained snapshot of the facts needed to interpret its numbers, including the scenario name and parameters, measured transaction, targets, results, environment, and TaffyJS revision. It retains enough sample evidence to support the displayed comparison. Its concrete schema and file encoding remain implementation details. The dataset is committed because it is evidence for a public performance claim, not a distributable build artifact. |
| 64 | + |
| 65 | +The benchmark page presents every scenario at the same level. Each applicable scenario shows a Taffy API comparison and, separately, a Yoga API comparison. It does not display browser results, a Diagnostics section, an overall score, an original-data download interface, or an empty column for a package that does not exist. Target columns come from the result's target list, so adding `@taffyjs/yoga-wasm` later is a target addition rather than a page redesign. |
| 66 | + |
| 67 | +The page footer shows only the concise environment and source identity needed by a reader, including the Node.js version, operating system and benchmark machine, and TaffyJS commit. |
| 68 | + |
| 69 | +## Delivery sequence |
| 70 | + |
| 71 | +First select and document the initial scenarios and their exact transactions. Then implement the top-level benchmark workspace and the two comparison groups, prove completion and semantic equivalence outside timing, run the complete suite on the fixed Node.js benchmark host, update the retained dataset, and let the website render it. Before publication, review the result for comparison validity, measurement stability, and avoidable harness complexity. |
| 72 | + |
| 73 | +The current delivery implements only the Taffy API comparison between `@taffyjs/node` and `@taffyjs/wasm`. The Yoga API comparison, captured Yoga application trees, and `@taffyjs/yoga-wasm` remain follow-up work after the relevant package pull requests merge. The current work may establish the shared result boundary, but it must not publish empty Yoga targets or manufacture a compatibility implementation in advance. |
0 commit comments