Skip to content

Latest commit

 

History

History
1120 lines (989 loc) · 58.4 KB

File metadata and controls

1120 lines (989 loc) · 58.4 KB

TinyTPU Tinyspec Coverage TODO

This TODO estimates the remaining work to support the tinyspec surface in tinygrad/spec/tinyspec.tex. It uses the current repo iteration style: one narrow, tested, committed behavior per iteration.

Overall Estimate

  • Broad functional tinyspec coverage: 250-400 iterations
  • Robust hardware-backed and well-tested coverage: 500-800 iterations

Current coverage includes hardware-backed TinyTPU execution for: GEMM (multi-WMMA, batched, deep-K, wide-N, with hardware fused bias+ReLU epilogue); full int32/bool VPU elementwise (ADD/MUL/SUB/MAX/MIN/DIV/MOD/ ABS, CMP{LT,EQ,NE}, AND/OR/XOR/NOT, SHL/SHR, WHERE, RELU, clip, fused add+relu, hardsigmoid, elu, mish); scalar-const variants and scalar broadcasting; all shapes with numel>16 through multi-tile elementwise loops; scalar, row-wise, and column-wise reductions (SUM/MAX/MIN/PROD) for any NxM through SXU_PROGRAM emitting VPU_{SUM,MAX,MIN,MUL}_REDUCE{,_COL,_TILE}; movement ops (reshape, contiguous slice/shrink, scalar expand, unrolled row expand); XLU transpose reachable via SXU_DISPATCH_XLU_TRANSPOSE at runtime level; TASM bundle assembler/disassembler; runtime bundle roundtrip + end-to-end sim tests.

Current Progress

  • Track tinyspec source in tinygrad/spec/tinyspec.tex
  • Runtime co-simulation bundle format for MXU programs
  • Runtime VMEM preload and VMEM result output path
  • Tinygrad GEMM lowering for supported tiled int32 cases through 4x4 MXU
  • Multi-row, batched, deep-K, wide-N GEMM test coverage
  • Tinygrad int32 single-tile VPU binary lowering
    • ADD
    • MUL
    • MAX
    • SUB
  • Tinygrad int32 single-tile VPU unary lowering
    • RELU
  • Tinygrad 4-element int32 reduction lowering
    • SUM via VPU_SUM_REDUCE
  • Full 16-lane VMEM tile coverage
    • ADD
    • MUL
    • MAX
    • RELU
  • Signed VPU test coverage
  • Runtime output validation for MXU and VMEM result lines
  • Remove generated bdpi/tinytpu_io.o from git tracking
  • TASM bundle assembler and disassembler (scripts/tasm.py, doc/tinytpu_asm.md)
  • Full-tile and multi-tile abs, IDIV, MOD coverage
  • Scalar broadcast (size-1 tensor) for MUL, SUB, MAX, multi-tile ADD/MUL
  • int32→bool and bool→int32 cast coverage
  • Clip (MIN+MAX program) full-tile and multi-tile
  • Fused add+relu full-tile and multi-tile
  • Tensor-tensor IDIV and MOD
  • Row-wise sum/max/min for NxM tensors — hardware-backed via VPU_{SUM,MAX,MIN}_REDUCE through SXU_PROGRAM for all N,M (legacy VPU_ROWSUM/HOST_ROWREDUCE removed)
  • Column-wise sum/max/min for NxM tensors — hardware-backed via VPU_*_REDUCE_COL for all N,M (legacy HOST_COLREDUCE removed)
  • Fix stale WAIT_MXU opcode in VPU-only test bundles
  • 2D tensor ops: all VPU binary/unary ops for arbitrary 2D shapes
  • Grouped scalar-const lowering for 2D/large tensors (NEG, x*c, x+c)
  • _tasm helper functions + bundle builders rewritten in readable assembly style
  • TASM bundle roundtrip tests for all bundle builders

Legacy Descriptor Removal Milestones

  • Milestone 1: remove legacy GEMM4x4
    • GEMM fallback (MULACC or scalar MUL+RANGE, no WMMA UOp) now emits SXU_PROGRAM with the same data_plan/instructions as the WMMA SXU path
    • _exec_gemm4x4 executor deleted
  • Milestone 2: remove legacy VPU_BINARY
    • Scalar-const IDIV, tensor-tensor IDIV, and bool→int32 cast now lower to SXU_PROGRAM
    • _exec_vpu_binary executor and VPU_BINARY emitter deleted
  • Milestone 3: remove legacy VPU_PROGRAM
    • Scalar-const MOD, tensor-tensor MOD, CLIP, ABS, and fused-add-relu already flow through SXU renderers
    • _exec_vpu_program executor deleted; the analyzer no longer emits VPU_PROGRAM
  • Milestone 4: remove legacy HOST_BINARY
    • Dead code: no emitter ever produced HOST_BINARY; executor and supported-op entry deleted
  • Milestone 5: remove legacy HOST_UNARY
    • RECIPROCAL lowered via FRECIP SXU program; TRUNC lowered via F2I+I2F SXU program
    • _exec_host_unary executor and HOST_UNARY emitter deleted
  • Milestone 6: simplify runtime after descriptor removal
    • _SUPPORTED_OPS = {"SXU_PROGRAM"}
    • _render_legacy_descriptor now returns only SXU descriptors (the GEMM fallback builds an SXU_PROGRAM)
    • analyze_tinytpu_uops deleted; ONNX trace reporting now reads renderer descriptors directly

InstSel Migration (UOp-Walking Renderer)

Replacing the ~47-recognizer _render_*_sxu_program waterfall with a single UOp-walking instruction-selection pass in tinygrad/tinygrad/runtime/support/tinytpu_lowering.py. See doc/plan-tinytpu-instsel.md.

  • Iteration 1: InstSel module — TpuInst/TpuKernel/encoder and the walker for int32/bool elementwise (ALU, WHERE, scalar broadcast, multi-tile). render() routes these kernels to the walker via the positive can_lower predicate.
  • float elementwise through the walker (VECTORIZE/GEP see-through, float CMPNE/CMPEQ, transparent bool→int cast); _render_elementwise_sxu_program and orphaned _find_alu_const deleted
  • transcendentals EXP2/LOG2/SIN routed through the walker (single VPU opcodes 51/52/53)
  • RECIPROCAL (FRECIP opcode 23) and SQRT (InstSel graph-rewrite to exp2(0.5·log2(x))) routed through the walker
  • deleted 34 transcendental/activation/elementwise _render_* recognizers and 3 orphaned pattern helpers (3062 lines); ops_tinytpu.py 6984 → 3551 lines
  • linear-scan VREG allocation with reuse (deep DAGs no longer burn a VREG per node)
  • divmod/trunc through the walker (MOD InstSel rewrite a-(a//b)*b, TRUNC F2I+I2F micro-pair); deleted _render_trunc and _render_scalar_const_divmod. Negative ///% now match tinygrad/numpy floor semantics — 4 tests that encoded the legacy recognizer's truncating result were corrected.

InstSel migration slice complete. Every elementwise / unary / transcendental / activation / divmod / trunc kernel is lowered by the UOp-walking InstSel pass in tinytpu_lowering.py. ops_tinytpu.py shrank from 6984 to ~3400 lines. The structural recognizers (reductions, broadcasts, pad, transpose, cast, copy) remain a deliberate later slice.

InstSel Migration — Structural Slice

  • Step 1 (ITER29): package split. tinytpu_lowering.py split into tinytpu_lowering/ package: common.py (shared types, opcode tables, graph helpers, InstSel pass), elementwise.py (walker), __init__.py (re-exports can_lower/lower_kernel). Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). See doc/plan-tinytpu-instsel-structural.md §6.
  • Step 2 (ITER30): classifier. classify.py introduces KernelClass enum (ELEMENTWISE, GEMM, UNSUPPORTED) and classify(uops) function. render() dispatches on KernelClass.ELEMENTWISE → lower_kernel; all other classes fall through to the structural recognizers unchanged. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions).
  • Step 3 (ITER31): reduction lowerer. reduction.py adds is_reduction(uops) (positive predicate) and lower_reduction(uops), a single classify-then-emit lowerer for scalar/row/column SUM/MAX/MIN/PROD reductions. KernelClass.REDUCTION is checked most-specific-first (after GEMM, before ELEMENTWISE); render() dispatches it to lower_reduction. The three legacy recognizers (_render_reduction/_render_rowreduce/ _render_colreduce, ~565 lines) and orphaned helpers _is_float_min_negation / _detect_reduce_op are deleted. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). See doc/plan-tinytpu-instsel-structural.md §7.
  • Step 4 (ITER32): broadcast lowerer. broadcast.py adds is_broadcast(uops) (positive predicate) and lower_broadcast(uops), a single classify-then-emit lowerer for row / column / column-where broadcasts. The axis is classified from the smaller operand's index relationship to the output ranges; colbc_where emits the column broadcast then the WHERE/select. KernelClass.BROADCAST is checked most-specific-first (after GEMM and REDUCTION, before ELEMENTWISE); render() dispatches it to lower_broadcast. The three legacy recognizers (_render_rowbc/_render_colbc/_render_colbc_where, ~300 lines) and the orphaned helper _classify_structured_broadcast_axis are deleted. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). See doc/plan-tinytpu-instsel-structural.md §8.
  • Step 5 (ITER33): trivial kernels. cast / copy / const-fill folded into the elementwise walker as degenerate elementwise maps. can_lower/ lower_kernel now accept a bare-LOAD DAG (copy, with a constant source offset for contiguous reshape / slice / shrink), a bare-CONST DAG (single-tile const-fill), and CAST value converts (I2F/F2I; bool→int stays transparent via _canon). _render_cast_sxu_program and _render_const_fill_sxu_program are deleted in full along with the orphaned _ALU_OP_NAMES. _render_copy_sxu_program (186 lines) is reduced to _render_rowbc_copy_sxu_program — the one non-degenerate case it carried, a single-input row broadcast (Tensor([[..]]).expand(N,M)), which is a structured broadcast, not a per-element map. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). See doc/plan-tinytpu-instsel-structural.md §6.
  • Step 6 (ITER34): GEMM relocate. Behavior-neutral relocation, two parts. (A) All bundle-instruction encoders (_vmem/_wmem/_amem, _load/_store/_vpu/_vpu_bg/_vpu_exp2/_vpu_log2/_vpu_sin, _select, _broadcast*, _mxu*, _psum*, _loop_*, _vzero/_vfill/ _vmov/_vneg/_vabs, _set_pred_*/_skip_*, _halt/_output_*/ _end/_bundle, …) moved verbatim into common.py; ops_tinytpu.py re-exports them so existing tests//scripts/ imports keep working. reduction.py/broadcast.py local encoder copies deleted and imported from common.py. (B) New gemm.py with lower_gemm/lower_gemm_fallback: the WMMA branch of _render_sxu_program, _render_gemm_fallback_sxu_program, and the GEMM-only helpers (_generate_gemm_sxu_instructions, _extract_wmma_epilogue, _apply_gemm_epilogue, _infer_tiling, _tiling_failure_note) moved verbatim. _find_unique_param_arg moved to common.py (shared with the structural recognizers). render() dispatches KernelClass.GEMM to lower_gemm; the non-WMMA matmul fallback runs via lower_gemm_fallback. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions); cosim passes. See doc/plan-tinytpu-instsel-structural.md §9.
  • Step 7-8 (ITER35): movement lowerer (Branch B). The Task 7 spike chose Branch B (renderer-side lowering): rangeify already dissolves the movement op itself, but the recognizers recover a non-affine tile access pattern the SXU model cannot express generically — they do real SXU-specific instruction selection, so they are relocated, not deleted. New movement.py adds is_movement(uops) (positive predicate) and lower_movement(uops), a classify-then-emit lowerer that ports the three legacy recognizers verbatim: pad / flip / non-affine permute (PAD_FILL scatter data plan), the 4×4 permute(1,0) (SXU_DISPATCH_XLU_TRANSPOSE), and single-input row-broadcast copy (SXU_BROADCAST_ROW). is_movement carries the legacy fallback-ordering gate (can_lower kernels — e.g. a scalar expand lowered as BROADCAST_SCALAR — were never reached by the recognizers and are excluded here). KernelClass.MOVEMENT is checked most-specific-first (after GEMM/REDUCTION/BROADCAST, before ELEMENTWISE); render() dispatches it to lower_movement. The three legacy recognizers (_render_pad/_render_transpose/_render_rowbc_copy_sxu_program, ~285 lines) and orphaned helpers _uop_contains/_has_load_src/_data_alu_ops/ _ALU_OPS are deleted. ops_tinytpu.py now has zero _render_*_sxu_program recognizers; _render_sxu_program is a trivial return None shell (Task 9 deletes it). Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). See doc/plan-tinytpu-instsel-structural.md §10.
  • Step 9 (ITER36): final cleanup. The empty _render_sxu_program shell and its call site in render() are deleted — render() is now a clean classify(uops) → per-class lowerer dispatch, with the non-WMMA matmul fallback (lower_gemm_fallback) preserved on the path before the UNSUPPORTED descriptor. Duplicated graph helpers _has_load_src / _data_alu_ops (byte-identical copies in movement.py/reduction.py/ broadcast.py) are hoisted into one canonical copy in common.py. The id()-based traversal closures in movement.py (_has_range, _consts_in, load-index distinctness set) are replaced with UOp.toposort()-based equivalents. The provably-dead WMMA reject entry is removed from the movement recognizers (classify routes WMMA → GEMM before is_movement); the MULACC guard is kept (a non-WMMA matmul can reach is_movement, so it is not provably dead). Dead _apply_gemm_epilogue (zero callers) is deleted from gemm.py. A deferred TODO is added at broadcast.py:_classify_broadcast_axis. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions); cosim passes.

Structural slice complete. All 12 structural recognizers migrated; ops_tinytpu.py has zero _render_* recognizers and is 1260 lines (from 6984); all kernel lowering is in the tinytpu_lowering package behind classify().

Unmasked hardware bug: the walker faithfully lowers tinygrad's decompositions, which exposed that the BSV EXP2/LOG2/SIN units are broken (EXP2 returns ~0 for positive inputs; LOG2 is very imprecise). The old activation recognizers used hardware-dodging formulations that hid this. ~40 transcendental/activation test failures, plus test_elu_positive_branch and test_log_compound_input, trace to this. Fix belongs in src/ BSV, not the backend — see doc/plan-primitive-ops-handoff.md.

Epilogue Gaps (surfaced by tests/test_e2e_epilogues.py)

End-to-end tests for the five CODA epilogue primitive classes (elementwise/ pairwise maps, vector loads/stores, tile loads/stores, tile reductions, stateful transforms) live in tests/test_e2e_epilogues.py. 18/19 pass; the following gaps are explicit:

  • Column-vector (M,1) broadcast over (M,N) mis-lowered as elementwise multiply. _classify_colbc picked the VPU op by the first non-zero op_count, so a single address-arithmetic MUL shadowed the real ADD. Fix: match the compute op by finding a binary UOp whose LOAD sources are the two input params (tinygrad/tinygrad/renderer/tinytpu/broadcast.py). Verified for ADD/MUL/SUB/rSUB/MAX with (M,1) operand.
  • Inline reduce inside an add chain silently produced garbage. Root cause: lower_gemm_fallback admitted any 3-param kernel with a single MUL + RANGE + STORE. An address-arithmetic MUL (RANGE * stride) was enough, so the unsupported fused-reduce-add kernel got phantom-lowered as a 1x1x1 GEMM. Fix: require the MUL to multiply two LOADed values (the actual acc += a*b signature). The fused kernel now raises NotImplementedError. The test tests/test_e2e_epilogues.py::TestStatefulEpilogues::test_inline_reduce_in_add_raises pins this behavior. A proper fused lowerer for (M,1) + reduce(tile, axis=1) is still open — callers lift the reduce to a named intermediate today.

Coverage Estimate by Area

Current VPU/Runtime Architecture: 20-40 iterations

  • Make VPU-only programs complete without dummy MXU dispatch
  • Add first-class runtime bundle builder for VMEM/VPU programs
  • Support multiple VMEM output tiles (col-reduce/row-reduce/GEMM emit multi-tile outputs through SXU_PROGRAM)
  • Support multiple VPU instructions in one tinygrad lowered program (SXU_PROGRAM multi-step paths: abs, clip, MOD, WHERE, row/col reductions)
  • Add runtime tests for VMEM preload/output protocol
  • Improve trace output for VPU-only programs
  • Document bundle records for VMEM input/output
  • Fuse GEMM epilogues in model runs
    • Post-run review on scripts/models/cnn_4_8_8_4.py showed the model runs end to end without host fallback or UNSUPPORTED, but each layer still lowers as MXU -> LOAD bias -> VPU ADD -> VPU RELU -> STORE rather than a first-class fused matmul + bias + relu epilogue path.
    • Add a direct runtime/compiler path so model kernels stop depending on separate bias-load and VPU epilogue instructions after every MXU tile.

General Elementwise Scalar/Tile Support: 30-60 iterations

  • Single-tile int32 ADD
  • Single-tile int32 MUL
  • Single-tile int32 MAX
  • Single-tile int32 RELU
  • Constants in VPU programs, e.g. x + 1, x * 2
    • x + scalar
    • x * scalar
    • maximum(x, scalar)
    • minimum(x, scalar)
    • x < scalar
    • x != scalar
  • Scalar broadcasting
    • Add explicit SXU/XLU broadcast opcodes (scalar/row/col)
    • Migrate scalar broadcast binary ops to SXU_PROGRAM via BROADCAST_SCALAR
  • Size-1 axis broadcasting (scalar-expand via BROADCAST_SCALAR; row-expand via BROADCAST_ROW for unrolled shapes)
  • Column-broadcast compare/select lowering for tinygrad workloads
    • (4,4) < (4,1) now lowers through SXU_PROGRAM with the BROADCAST_COL primitive instead of falling through to UNSUPPORTED.
    • The post-run review model path mask.where(y, y * 2) now closes through a dedicated single-tile BROADCAST_COL_SELECT SXU program when the compare and select stay fused in one kernel.
    • Follow-up: a separately realized bool mask still exposes a different mixed bool/int32 arithmetic gap in the generic VPU_BINARY path; keep that as a distinct issue rather than regressing the fused review-model path.
  • Arbitrary shapes with numel <= 16 (1D/2D/nD elementwise + reshape/slice/expand covered by copy + elementwise renderers)
    • Shape-preserving 2x2 elementwise coverage for supported VPU ops
  • Multi-tile elementwise loops for numel > 16
    • Multi-tile AND/OR/XOR/NOT for bool tensors
  • Mixed VPU op chains without host round trips (via VPU_PROGRAM path)
  • Output shape preservation for scalar, vector, and small matrix cases (1D/2D/nD coverage via copy + elementwise + reduction renderers)

More Elementwise Ops: 35-70 iterations

  • SUB as ADD + NEG, lowered to VPU_SUB
  • NEG as multiply by -1
  • CMPLT
  • CMPNE
  • CMPEQ
  • WHERE
  • AND
  • OR
  • XOR
  • NOT
  • SHL
  • SHR
  • MOD (via DIV+MUL+SUB bundle)
  • IDIV (via VPU_DIV, truncation semantics)
  • RECIP
  • TRUNC
  • Basic CAST (int32↔bool; int32↔float32 via VPU_I2F/F2I; fused same-dtype round-trips via COPY)
  • Basic BITCAST
  • Basic COPY (_render_copy_sxu_program handles same-dtype identity-index kernels via LOAD/STORE pairs)

Reductions: 30-60 iterations

  • 4-element int32 sum to scalar
  • Full-tile int32 sum to scalar (via VPU_SUM_REDUCE_TILE)
  • Multi-tile int32 sum to scalar (VPU_SUM_REDUCE_TILE per tile + VPU_ADD combine)
  • Row-wise sum/max/min over NxM tensor (SXU_PROGRAM row-reduce renderer, all N,M)
  • Column-wise sum/max/min over NxM tensor (SXU_PROGRAM col-reduce renderer, all N,M)
  • Full-tile sum (via VPU_SUM_REDUCE_TILE)
  • MAX reduction (4-elem, full-tile, multi-tile via VPU_MAX_REDUCE_TILE)
  • MIN reduction (4-elem, full-tile, multi-tile via VPU_MIN_REDUCE_TILE)
  • VPU col-reduce primitives (VPU_SUM/MAX/MIN_REDUCE_COL opcodes 29/30/31)
  • VPU tile-reduce primitives (VPU_SUM/MAX/MIN_REDUCE_TILE opcodes 32/33/34)
  • MUL reduction (VPU_MUL_REDUCE{,_COL,_TILE} hardware + tinygrad lowering for scalar/row/col prod)
  • keepdim behavior (rowsum/rowmax/rowmin keepdim tests pass)
  • Multi-tile reductions (scalar SUM/MAX/MIN with VPU_*_REDUCE_TILE per tile + VPU_ADD/MAX/MIN combine)
  • Reduction axis shape validation
  • Reduction result layout in VMEM

Movement Ops: 40-80 iterations

  • RESHAPE (copy SXU_PROGRAM with identity LOAD/STORE index mapping)
  • PERMUTE
  • TRANSPOSE via XLU — hardware: SXU_DISPATCH_XLU_TRANSPOSE opcode (12) + runtime test done; tinygrad permute(1,0) lowering still open
  • EXPAND (scalar and unrolled row-broadcast via BROADCAST_SCALAR / BROADCAST_ROW; RANGE-loop variant still open)
  • SHRINK (contiguous slice via copy renderer with affine offset)
  • PAD
  • FLIP
  • CAT
  • INDEX with simple affine patterns (contiguous slice with constant offset through copy renderer)
  • Gather-like indexing
  • XLU-backed broadcast primitives (scalar/row/col)
  • XLU-backed permutation paths
  • Multi-tile movement kernels

GEMM and Matmul Expansion: 25-50 iterations

  • 4x4 GEMM
  • Multi-row GEMM
  • Batched GEMM
  • Deep-K tiled GEMM
  • Wide-N tiled GEMM
  • Multi-WMMA lowering (multi-tile GEMM through WMMA path)
  • Bias epilogue (row-broadcast and full-tensor, hardware-backed for single-K-tile)
  • ReLU epilogue (hardware-backed for single-K-tile)
  • Fused bias+ReLU epilogue (hardware-backed via SXU_LOAD_MXU_RESULT → VPU)
  • M/N/K tail handling
  • Better unsupported shape diagnostics
  • Add/mul epilogue
  • Non-int8 operand policy
  • Accumulation overflow tests
  • Multi-output tile scheduling cleanup
  • Hardware epilogue for multi-K-tile GEMM — landed via PSUM bucket bank. Multi-K-tile GEMMs now run SXU_PSUM_CLEAR → N × DISPATCH_MXU(psum_acc) → SXU_PSUM_READ_ROW → bias/relu → STORE entirely in hardware. No numpy fallback path remains.

Dtypes: 25-50 iterations

  • int32 VMEM values for VPU paths
  • int8 operands for MXU paths
  • bool comparison outputs
  • int8 elementwise
  • uint8 elementwise
  • int16 elementwise
  • uint16 elementwise
  • uint32 elementwise
  • float32 policy (FADD/FSUB/FMUL/FMAX/FCMPLT dispatched via VPU_F* variants; scalar const and multi-tile covered; FRECIP+scalar fdiv work; tensor-tensor fdiv requires RECIPROCAL UOp detection, tracked)
  • cast saturation/wrapping behavior
  • comparison output dtype behavior
  • dtype range diagnostics

Call/Tuple/Ordering Semantics: 20-40 iterations

  • CALL
  • TUPLE
  • GETTUPLE
  • AFTER
  • assign/store dependency ordering
  • multiple outputs
  • multiple stores in one kernel
  • reusable captured graph fragments

Multi-Kernel and Memory Planning: 20-50 iterations

  • Multiple TinyTPU programs per tensor expression
  • Intermediate VMEM allocation
  • Spill/fill to HBM model
  • Host/device copy scheduling
  • Cross-kernel dependency tracking
  • Runtime buffer lifetime tests
  • Larger tensors split across VMEM tiles

Tinygrad Backend Integration Quality: 40-80 iterations

Shrink ops_tinytpu.py (~2335 → target <1200 lines)

Current bloat sources:

  • Structural lowerers in tinygrad/renderer/tinytpu/* still contain legacy recognizer ports. Convert repeated descriptor/data-plan construction into reusable builders.
  • Bundle builders (~400 lines): _build_vpu_binary_bundle, _build_vpu_where_bundle, _build_full_gemm_bundle, etc. Unify into a single generic bundle builder with a template pattern.
  • _exec_* methods (~400 lines): every op has the same chunk loop (for chunk_start in range(0, num_elems, _TILE_ELEMS)). Extract a shared _run_tiled_vpu helper.
  • Duplicate output parsing: _parse_vmem_output, _parse_multi_vmem_output, _parse_sim_output share the same structure.

Cleanup plan — eliminate analyze_tinytpu_uops via SXU_PROGRAM migration:

  • Add VPU_NOT hardware opcode (BSV + testbench + Python table)
  • Migrate scalar-const binary ops (x+c, x*c, NEG, NOT) to SXU_PROGRAM
  • Migrate bool-typed ops (AND/OR/XOR/NOT on bool tensors) to SXU_PROGRAM
  • Migrate WHERE (ternary select) to SXU_PROGRAM (now via first-class SXU_DISPATCH_SELECT using VPU_SELECT)
  • Migrate multi-step VPU_PROGRAM patterns (abs, clip, MOD, CMPEQ) to SXU_PROGRAM
  • Migrate scalar reductions (SUM/MAX/MIN to scalar) to SXU_PROGRAM
  • Migrate row-wise reduce to SXU_PROGRAM (done: _render_rowreduce_sxu_program, legacy VPU_ROWSUM removed)
  • Migrate row-broadcast binary (VPU_ROWBC_BINARY) to SXU_PROGRAM
  • Emit unsupported diagnostics directly from renderer descriptors
  • Delete dead analyze blocks, _exec_vpu_where/unary, _build_vpu_where, _render_wmma_descriptor
  • Delete remaining analyze_tinytpu_uops
  • Run selected upstream tinygrad tests on TINYTPU (2 tests pass via scripts/run_tinytpu_upstream_subset.py — expand list as coverage grows)
  • Add skipped/xfail manifest for unsupported tinyspec areas
  • Track coverage by tinyspec op category

Profiler and Debuggability: 20-40 iterations

  • Trace parser validation
  • Perfetto emission for existing traces
  • Full VPU-only trace coverage
  • VMEM bundle dump utility
  • Lowering decision dump (TINYTPU_DUMP_LOWERING env var emits renderer JSON)
  • Runtime bundle round-trip tests for all record types
  • Per-op cycle reports
  • MXU/VPU/VMEM utilization by lowered tinygrad op

Optional/Advanced Tinyspec Surface: 50-150 iterations

  • Multi-device COPY
  • REPLICATED
  • collectives: broadcast, scatter, gather, reduce, allgather
  • reduce_scatter
  • allreduce
  • optimizer axis semantics
  • BARRIER
  • SPECIAL
  • IF / ENDIF
  • WMMA mapping beyond current MXU path
  • CUSTOM
  • ATOMICADD
  • CUSTOMFUNCTION
  • PROGRAM / SOURCE / BINARY metadata handling

Major Hardware Capability Gaps vs tinyspec

These are the remaining hardware-level gaps that cannot be solved by software lowering alone. Ordered roughly by impact on real workloads.

Data path & reductions

  • Float reducer tile-scope primitives VPU_FSUM_REDUCE_TILE (opcode 38), VPU_FMAX_REDUCE_TILE (39), VPU_FMIN_REDUCE_TILE (40). Float sum/max scalar reductions lower end-to-end; float min lowers for single-tile kernels via a negation-around-max pattern rewrite.
  • Multi-tile float min combine — VPU_FMIN ALU opcode (41) added; renderer emits per-tile FMIN_REDUCE_TILE plus VPU_FMIN across tiles. Works up to 32 elements; larger sizes hit the pre-existing RANGE-loop reduction gap shared with other scalar reductions.
  • Float row-reduce primitives (VPU_FSUM_REDUCE 42, VPU_FMAX_REDUCE 43, VPU_FMIN_REDUCE 44). Per-sublane float reducer ops added; row-sum and row-max are wired through the tinygrad row renderer. Row-min still needs negation-decomp detection in the row path.
  • Float col-reduce primitives (VPU_FSUM_REDUCE_COL 45, VPU_FMAX_REDUCE_COL 46, VPU_FMIN_REDUCE_COL 47). Landed via the "hoist once outside the case block" pattern (same as the integer col reducers) using shared lane_f* helpers. Col-sum and col-max wired through tinygrad. Col-min pending negation rewrite.
  • Shared multi-cycle FP tile reducer — mkFpReducer sub-module with one FP adder + one comparator driven by a small FSM. Tile-scope FP reducers now dispatch through it; SXU stalls on vpu.isDone until the walk completes. Runtime build went from 1:48 to 1:29 after replacing the 3 combinational tile FP trees, and adding 3 col reducer opcodes on top cost only +7s (1:36), meeting the <10s/opcode acceptance bar.
  • Float prod reducer — no opcode; tinyspec Reduce(Mul, axes) on float would need a VPU_FMUL_REDUCE* family.

Build-time / bsc elaboration

  • Shared multi-cycle FP reducer unit — src/FpReducer.bsv landed. Tile-scope FP reducers now route through it; acceptance criterion met (+7s for 3 new col reducer opcodes). Row/col reducers still use the "hoist once" combinational pattern with shared lane_f* helpers, which is cheap at the current 4×4 lane count. If we scale lanes or add more reducer kinds (e.g. float prod), consider routing row and col through FpReducer too.
  • Narrow-dtype VPU lanes: int8, uint8, int16, uint16, uint32 elementwise. Current VPU is int32-only on the data path (except MXU which is int8). Blocks quantized-int8 inference kernels outside the MXU, and uint dtypes entirely.
  • Mixed-precision MXU: MXU is int8-only today. No bf16/fp16 MAC unit, no fp32 accumulator path for floating GEMM. Float32 GEMM currently can't lower at all.

Transcendental / math

  • Exp2/Log2/Sin/Cos hardware — Remez + range reduction. All four primitives live in src/TranscUnit.bsv: - Opcodes: VPU_EXP2 (51), VPU_LOG2 (52), VPU_SIN (53), VPU_COS (54). All dispatch through the multi-cycle walker; VPU isDone gates SXU collect. - Coefficients: Remez minimax for all four (EXP2 4× peak-error reduction, SIN 40×, COS 27×, LOG2 8× vs Taylor). - EXP2 range reduction: x = n + f via tr_trunc/tr_fp_to_int/ tr_pow2_int helpers; poly(f) on [-1, 1] scaled by 2^n via exponent-bit construction. Integer x gives exact 2^n results. - SIN range reduction: mod-2π + quadrant fold + sign-aware round-to-nearest bias. sin(π), sin(2π) etc. zero out; sin(5) accurate to <0.3% rel. - Tinygrad code_for_op declares Ops.EXP2/LOG2/SIN/SQRT hardware-supported. Renderers land Tensor-level: _render_scaled_exp2 (Tensor.exp), _render_sigmoid (FMUL+EXP2+FADD+FRECIP), _render_tanh, _render_scaled_sin (Tensor.cos), _render_scaled_log2 (Tensor.log), plus plain _render_{exp2,log2,sin,sqrt}_sxu_program. - Guards in elementwise/reciprocal/multistep-scalar-divide renderers reject kernels with EXP2/LOG2/SIN/SQRT so missing composite renderers fail loud. - Standalone TbTranscUnit (make test-transcunit) locks the coefficient contract at the unit layer. - End-to-end backend tests cover wide ranges: tanh |x|≤2.5, exp [-3,3], sin/cos [-3π,3π], sqrt [0.25,16].

Movement

  • [~] PAD primitive — small unrolled pads lower through the _render_pad_sxu_program renderer (PAD_FILL VMEM preload). Multi-tile + RANGE-loop pads still open.
  • [~] FLIP primitive — unrolled flip / non-affine permutes lower through the pad_fill renderer too. Multi-tile FLIP and RANGE-loop FLIP still open.
  • CAT primitive — concatenate along an axis. Could be a two-source SXU copy program driven by a range split, but no primitive exists today.
  • General Permute — only permute(1,0) (transpose) is wired through SXU_DISPATCH_XLU_TRANSPOSE. Non-transpose axis reorders for ≥3D tensors have no hardware path.

GEMM & matmul

  • Multi-K-tile hardware epilogue — landed via PSUM bucket bank (Item #2 below). Multi-K-tile GEMM now zeros a PSUM bucket via SXU_PSUM_CLEAR, accumulates every K-tile into row 0 via DISPATCH_MXU with psum_acc mode, then extracts the row with SXU_PSUM_READ_ROW and runs the existing bias+ReLU epilogue on the PSUM result. No numpy fallback path remains for multi-K-tile bias/relu.

Control flow & ordering

  • BARRIER / IF / ENDIF / AFTER SXU ops — SXU today is a straight-line sequencer. No conditional execution, no explicit ordering barriers for RAW hazards across dispatches.
  • ATOMICADD unit on VMEM banks — no RMW lane on the VMEM write port. Required for scatter-add style kernels.

Multi-device & parallelism

  • Collective primitives (broadcast/scatter/gather/allreduce/ reduce_scatter) expressed as NoC-backed hardware ops. NoC exists at the ring level but is not exposed as a tinyspec-shaped primitive.
  • REPLICATED / multi-device COPY model — the compiler has no concept of cross-chip axis today.
  • THREEFRY PRNG datapath — no ARX unit. All random init currently falls back to host.

Memory planning

  • VMEM spill/fill to HBM — compiler-managed spill path for tensors larger than the VMEM working set. Today everything must fit in VMEM tiles.

Architectural Refactors (higher-impact, bigger than opcode additions)

These are the structural improvements that came out of reviewing the TensorCore datapath against the NeuronCore reference and thinking about what's actually bottlenecked today. Ordered by impact/effort ratio — the top ones would most change what TinyTPU can run.

Tier 1 — parallelism and data movement

  • Decouple SXU into per-engine command queues. The single microsequencer is the visible biggest bottleneck: every engine shares one issue slot. MXU already overlaps via DISPATCH_MXU+WAIT_MXU, but VPU / FpReducer / XLU all serialize even though they touch disjoint state. Split the front-end into an SXU fetch/decode stage feeding four small per-engine FIFOs, each with its own execute-side controller; add a scoreboard keyed on VRegFile destination index for RAW/WAW. Closes the biggest visible gap vs NeuronCore's per-engine SEQs.
  • Dedicated PSUM accumulator bank. Landed. src/PSUMBank.bsv is an 8-bucket × 4×4 Int32 SRAM with write/accumulate/ writeRow/accumulateRow/readReq/readResp/peekRow/ clear and row-granular access for MXU. Shared between Controller (per-dispatch row deposit via startPsum) and SXU (opcodes 15-19: PSUM_WRITE, PSUM_ACCUMULATE, PSUM_READ, PSUM_READ_ROW, PSUM_CLEAR). The tinygrad renderer uses the bank for multi-K-tile GEMM accumulation, which closes the "Multi-K-tile hardware epilogue" gap above.
  • Engine-to-engine forwarding (bypass VRegFile). Today every hand-off is engine → resultReg → VRegFile → engine, four cycles. Add a muxed forwarding network on engine inputs so vs/vs2 can come straight from the MXU result reg / VPU resultReg / XLU output when SXU emits a FWD hint instead of a vreg index. The current SXU_LOAD_MXU_RESULT special case is the first half of this feature — generalize it.
  • Dual-issue slot for VPU + XLU. Landed. XLU dispatch rules now advance pc and return to FETCH immediately instead of stalling in an XLU_COLLECT state; a background do_xlu_collect_bg rule fires the cycle after dispatch (XLU.result carries 1-cycle latency) and writes xlu_dst while the main FSM fetches/executes the next non-XLU op. Structural-hazard guard !xlu_busy on every XLU dispatch rule prevents back-to-back XLU ops from issuing before the prior one collects. No RAW hazard in practice because the collect's vrf.write happens one cycle before any dependent read could fire (dispatch advances pc through FETCH before any EXEC reads vreg). End-to-end chained XLU broadcast sim test covers the path.

Tier 2 — memory hierarchy and programmability

  • Double-buffered Weight/ActivSRAM + small DMA engine. src/WeightSRAMDB.bsv and src/ActivationSRAMDB.bsv are now the default SRAMs inside TensorCore. The Controller consumes the .plain sub-interface (writes → INACTIVE, reads → ACTIVE); TensorCore exposes loadWeightTile / loadActivationTile (writeActive) for "preload then dispatch immediately" flows and preloadWeightTile / preloadActivationTile + swapWeightBanks / swapActivationBanks for DMA-overlap flows. Covered by make test-tc-db (ping-pong dispatch) plus the existing WeightDMA / ActivationDMA standalone benches.
  • Transcendental / programmable SIMD unit. sqrt / log2 / exp2 / sin are rejected today; NeuronCore covers these with its GPSIMD engine. Three options in increasing generality: (a) polynomial-approximator opcodes with fixed hardware LUTs, (b) a dedicated transcendental unit next to the VPU, (c) a tiny programmable SIMD lane (RISC-V style). (a) is the natural first iteration.
  • Predicate register + SKIP_IF_ZERO (baby IF / BARRIER). Landed. src/ScalarUnit.bsv gains a 1-bit pred Reg plus two opcodes: SXU_SET_PRED_IF_ZERO (opcode 20, pred := vreg[0][0] == 0) and SXU_SKIP_IF_PRED (opcode 21, advances pc by 2 when pred is set and auto-resets pred). TASM opcode table and Python wire helpers (_set_pred_if_zero, _skip_if_pred) landed; runtime sim tests cover both skip-taken and skip-not-taken paths. No tinygrad renderer is emitting these yet — they're infrastructure for future BARRIER / IF / ENDIF work.
  • Output-stationary dataflow mode on MXU. Real OS dataflow is now the primary path: - PE.feedPair(w, a) MACs with both operands arriving from neighbors; passWeight mirrors passActivation. - SystolicArray.feedPair(w_top, a_left) drives one systolic wavefront per cycle; getMatrix() returns the full per-PE psum (rows x cols). Covered by make test-array (2x2 OS matmul). - Controller.startOS(wBase, aBase, kLen) runs the full staircase: LoadWeights → LoadWeightsRespOS → LoadActsOS (K cycles) → StreamOS (K + (rows-1) + (cols-1) cycles) → Drain. Both operand staircase skews are applied inside the Controller. resultsMatrix() getter exposes the full output. make test-ctrl-os covers 4x4 identity @ arange = arange. - SXU opcode SXU_DISPATCH_MXU_OS (25) routes through it; TASM MXU_OS WMEM[w], AMEM[a], k=N and Python _mxu_os(w,a,k) emit the wire form. The old misnomer MXU_OS opcode (23) was renamed MXU_ACCUMULATE — see the dataflow roadmap below. - SXU_LOAD_MXU_MATRIX_ROW (26) copies one row of the psum matrix into a vreg so consumers can walk the drain row by row. - SXU_DISPATCH_MXU_OS_ACCUMULATE (33) dispatches without the drain-time clear, letting multi-K-tile OS scale past K == rows. - Covered end-to-end: test_mxu_os_end_to_end_identity_matmul, test_mxu_os_end_to_end_nontrivial_matmul, test_mxu_os_accumulate_multi_tile. Lean theorem accumulate_compose formalizes the multi-tile composition.

MXU dataflow roadmap

The four canonical systolic dataflows, ranked by TinyTPU fit:

  • Weight-stationary (WS). Fully implemented. Weights loaded once, activations + partial sums flow. Good for weight-heavy layers; matches TPU / NVDLA.
  • Output-stationary (OS). Fully implemented as of iter 42. PE.feedPair + SystolicArray.feedPair + Controller.startOS run the staircase; drain returns the full (rows x cols) psum. Dispatched via SXU opcode 25 (MXU_OS). High-value for depthwise conv and attention kernels.
  • WS-accumulate (K-batching on WS). Controller.startAccumulate preserves the WS accumulator across dispatches; dispatched via SXU opcode 23 (MXU_ACCUMULATE). Not a distinct dataflow — a scheduling overlay on WS that skips the drain-time clear. Also covered at a different granularity by the PSUMBank.
  • Input-stationary / activation-stationary (IS). Deferred. ActivationSRAM stores (rows,)-shaped vectors per address, not (rows, cols) tiles, so IS would need either a pre-load buffer built from K separate reads or a new ActivationTileSRAM shape. The compute win is narrow (batch=1 inference where activations are large and weights stream cheaply from HBM), and the OS path now covers most of the same workload class. Revisit if a real workload needs it.
  • Row-stationary (RS) — Eyeriss-style. Overkill for a 4x4 pure-matmul engine; skip unless the array grows + conv kernels become the bottleneck.

Hybrids worth noting:

  • K-batching / multi-tile accumulation via Controller.startAccumulate and/or PSUMBank accumulate buckets.
  • No-local-reuse (NLR): intentionally skipped — no reason to implement as a first-class mode.

Tier 3 — chip-level scale-out

  • Shared on-chip L2 between TensorCores. NoC exists, but cores re-fetch overlapping weights from HBM for data-parallel kernels. An L2 cache or an NoC broadcast primitive amortizes weight traffic across cores.
  • [~] Generalize SXU_LOAD_MXU_RESULT into engine-to-engine read ports. Hardware and TASM support landed: SXU_LOAD_VPU_RESULT (13) and SXU_LOAD_XLU_RESULT (14) both copy the engine's linger result register into a vreg. Bundle tests cover both. Remaining: tinygrad renderer is not yet using these opcodes to elide VRegFile round-trips in chained kernels (e.g. XLU-then-VPU, VPU chains that store and re-read).
  • Per-engine writeback resolved via VRegFile Vector-of-Reg. Before: mkRegFileFull had one write port, so bsc forced every vrf writer mutex with do_xlu_collect_bg (G0010 urgency conflicts — the dual-issue background rule was silently serialized). After: VRegFile is Vector#(numRegs, Reg#(T)); each index has its own write port. G0010 conflicts collapse to G0036 (same-cycle ordering — only matters if two rules target the same index, which valid programs never do). 865 sim-backed backend tests clean.

Recent SXU extensions (push 3)

New single-cycle opcodes added for instruction density and perf work:

  • SXU_LOAD_MXU_MATRIX_ROW (26) — drain one row of OS psum matrix.
  • SXU_READ_CYCLE (27) — free-running cycle counter into lane 0. Backed by a new unconditional cycle Reg (was TRACE-only).
  • SXU_LOOP_BEGIN (28) / SXU_LOOP_END (29) — single-level counted loop (counter in mxuTLen UInt#(8)). Renderers can emit a tight body without unrolling.
  • SXU_VZERO (30) — single-cycle zero vreg. Used by the softsign renderer to replace a zero-tile VMEM preload.
  • SXU_VFILL (31) — broadcast signed 8-bit immediate to all lanes.
  • SXU_VMOV (32) — single-cycle vreg copy.
  • SXU_DISPATCH_MXU_OS_ACCUMULATE (33) — multi-K-tile OS.
  • SXU_VNEG (34) — single-cycle lane-wise negate.
  • SXU_VABS (35) — single-cycle lane-wise absolute value. Int32 abs renderer now emits LOAD + VABS + STORE (3 instrs, was 5).

Renderer adoption of the new opcodes is an ongoing task — most renderers still emit the old multi-instruction patterns.

Recent SXU + VPU extensions (push 4)

Opcode total: SXU 36 → 41; VPU 55 → 82.

SXU additions (36-40):

  • SXU_LOAD_LOOP_DEPTH (36) — read current loop-stack depth (for debug introspection of nested loops).
  • SXU_DISPATCH_XLU_ROTATE (37) — exposes XLU's cyclic-rotate hardware that had no SXU entry point.
  • SXU_PSUM_CLEAR_ALL (38) — multi-cycle walker inside SXU that zeros every bucket. 1 instruction replaces 8 PSUM_CLEARs for multi-K GEMM epilogues.
  • SXU_SET_PRED_NE_ZERO (39) / SXU_SKIP_IF_NOT_PRED (40) — completes the predicate pair (if-zero and if-nonzero, skip-if-pred and skip-if-not-pred, both arms of if/else expressible).

SXU LOOP_BEGIN/LOOP_END now use a depth-4 stack (was single-level in push 3) — nested counted loops work up to 4 levels.

VPU additions (55-81):

  • Packed-int8 arithmetic (55-67): PACKED_I8_ADD / _SUB / _MAX / _MIN (55-58) byte-wise saturation; _NEG / _RELU (59-60) unary; _CMPLT / _CMPEQ (61-62) byte-wise compare; _MUL_LOW / _MUL_HIGH (63-64) multiply with wrap / Q-format; _ABS / _SIGN (65/67) unary.
  • Int32/float sign: VPU_SIGN (66), VPU_FSIGN (68), VPU_FABS (79). Softsign + float-abs renderers adopted FABS — drops 2 instructions per tile from both.
  • Reductions: VPU_ARGMIN / VPU_ARGMAX (69/70) per-row index.
  • Bit-manip: VPU_CLZ (71), VPU_POPCOUNT (72), VPU_CTZ (73), VPU_BYTE_REVERSE (74), VPU_ROTL (80), VPU_ROTR (81).
  • Saturating arithmetic: VPU_SAT_ADD_I32 (75), VPU_SAT_SUB_I32 (76), VPU_ABS_DIFF_I32 (77), VPU_PACKED_I8_ABS_DIFF (78).

Renderer adoption (push 4)

  • softsign (iter 26) — adopts VPU_FABS: drops the VZERO + FSUB + FMAX 3-instruction sequence, keeps only LOAD(1.0)+FABS.
  • float-abs (iter 27) — adopts VPU_FABS: drops the zero-tile VMEM preload + FSUB + FMAX (5 instrs → 3 per tile).
  • int32-abs (push 3) — already uses SXU_VABS (5 instrs → 3).

Lean theorems added (push 4): zero_weight_accum_unchanged, sign_idempotent, sign_odd, abs_step_preserves_weight, abs_fold_preserves_weight, sign_bounded, sign_times_self_nonneg — 7 new theorems on top of the 3 from push 3 (10 total).

VPU dual-issue (push 5 iters 1-4, LANDED)

Dedicated SXU_DISPATCH_VPU_BG opcode (41) for background-collect dual-issue. Unlike push-4 iter 2's attempt (which converted the existing sync path and broke multi-cycle ops), this iter keeps sync DISPATCH_VPU behavior identical and adds the BG path alongside. The main FSM's do_fetch dispatches both opcodes through one rule (do_vpu) that forks on pc_state: sync transitions to EXEC_VPU_COLLECT and stalls; BG advances pc and returns to FETCH, overlapping subsequent non-vreg-reader ops (MXU dispatch, PSUM clear, loop control, skip-if-pred). A single unified do_vpu_collect rule retires via vrf[vpu_wb_dst] for both paths — keeping the VRegFile writer count at parity with push-4 HEAD is what avoids blowing past bsc's -steps-max-intervals 20 elaboration budget.

Iter 3 added conservative !vpu_wb_pending RAW stall to every vreg- reader rule so multi-cycle BG (LOG2/EXP2/SIN/COS) and single-cycle BG followed by a dependent reader both stay correct. The fix is pessimistic: the reader stalls until the writeback lands even if its source vreg doesn't collide with vpu_wb_dst. A more precise curInstr.vregSrc == vpu_wb_dst check would let non-dependent readers overlap, but both the in-guard form and a fetch-time precompute form blow past the unfolding budget — likely because accessing packed-struct fields (curInstr.vregSrc) re-decodes the SxuInstr representation at every use site. Resolving that needs either a flat-field decode stage (separate Regs for vregSrc/vregSrc2/vregDst) or a restructured VRegFile with fewer writer sites per module.

Iter 4 formalized the dispatch/collect equivalence in Lean: bg_dispatch_then_collect_matches_sync, bg_collect_idempotent_when_empty, bg_dispatch_preserves_reg (13 theorems total, was 10).

Remaining push-5 work

  • Precise per-vreg RAW scoreboard. Fine-grained alternative to the coarse !vpu_wb_pending stall. Requires flat scalar Regs for curInstr.vregSrc/vregSrc2/vregDst populated by do_fetch so the hazard check cur_raw_pending stays cheap to elaborate. Unlocks real dual-issue overlap on VPU-adjacent readers.
  • Multi-input VPU command queue. A small FIFO of pending VPU dispatches so the main FSM can keep fetching/decoding while back-to-back VPU_BG ops queue up. Useful when pairing a VPU_BG with XLU/MXU dispatch to keep the engines busy. Needs careful elaboration budget: a SizedFIFO + drain rule adds ~2 vrf.read call sites, which may push bsc over.
  • Renderer adoption. Tinygrad's FABS / softsign / float-abs tile chains can emit VPU_BG for the final VPU op before STORE once the precise RAW scoreboard lands — without it, the coarse stall burns the dual-issue window.
  • Per-engine command queues. Full Tier-1 decouple: SXU fetch stage feeds four per-engine FIFOs (VPU/XLU/MXU/FpReducer) with independent executors. Groundwork is in place (vpu_wb_pending tracks VPU retire, xlu_busy tracks XLU retire); the missing pieces are the FIFOs and a RAW scoreboard keyed on vreg destination index.

Build-time budget: push-5 iter 3 baseline is ~186s CPU for a clean build/mkTbTinyTPURuntime.bexe. Treat >2x regressions as stop-the- line; the usual fix is collapsing vrf.write call sites or moving cross-field equality checks out of rule guards.

Deferred (called out but probably not next)

  • Mixed-precision MXU (bf16/fp16 MACs) — big surgery, unclear ROI without training workloads.
  • Hardware collective primitives on the ring NoC — worth the software plumbing first; hardware once we have multi-core workloads.
  • ATOMICADD on VMEM banks — only matters for scatter-add kernels which aren't on our critical path.
  • TranscUnit TANH / SIGMOID direct opcodes. Considered but deferred: Padé forms need a divider (not in TranscUnit); degree-5 odd Remez + saturation is feasible but requires careful coefficient fitting. Lower ROI now that exp2/log2 compositions cover the usage.
  • Int8 packed VPU arithmetic. [largely landed in push 4] Coverage: ADD/SUB (saturating), MAX/MIN, NEG (saturating), RELU, CMPLT/CMPEQ, MUL_LOW (wrap)/MUL_HIGH (Q1.7), ABS, SIGN — 10 opcodes (55-67). Remaining: (a) byte-wise logical ops reuse VPU_AND/OR/XOR/NOT via the 32-bit path (no new opcode needed); (b) unpack/pack between packed-i8 and full int32 tiles — deferred until a real quantized workload needs it (shape math is tricky); (c) renderer adoption (none yet — no int8 tinygrad workload has landed).
  • Nested loops in SXU. [done in push 4 iter 1] LOOP_BEGIN/END now use a depth-4 stack (loopCounterStack + loopReturnPcStack + loopTop). LOAD_LOOP_DEPTH (36) exposes the current depth for debug.

Milestones

Milestone 1: Useful Single-Tile Tensor Programs

Estimate: ~50 total iterations

  • Single-tile binary VPU ops
  • Single-tile ReLU
  • One simple reduction
  • Constants and scalar broadcasting
  • WHERE and comparisons
  • Basic reshape/transpose within one 4x4 tile (reshape via copy renderer; transpose reachable via SXU_DISPATCH_XLU_TRANSPOSE, tinygrad lowering pending)
  • VPU-only runtime completion cleanup

Milestone 2: Multi-Tile Core Tensor Ops

Estimate: ~120 total iterations

  • Multi-tile elementwise (ADD/MUL/SUB/MAX/MIN/compare/logic + scalar-const all cover numel>16 via tile-loop)
  • Multi-tile reductions (scalar via VPU_*_REDUCE_TILE; row/col via SXU_PROGRAM tile iteration)
  • Multi-tile movement ops (multi-tile reshape/slice/expand work through the copy renderer's tile loop)
  • GEMM epilogues (hardware-backed bias add, ReLU, fused bias+ReLU via SXU_LOAD_MXU_RESULT)
  • More shape coverage
  • First selected upstream tinygrad test subset passing on TINYTPU (test_plus_int, test_plus_big)

Milestone 3: Broad Core Tinyspec Semantics

Estimate: ~250 total iterations

  • Most movement ops
  • Most integer elementwise ops
  • Add/max/mul reductions (SUM/MAX/MIN/PROD scalar, row, col via VPU_*_REDUCE{,_COL,_TILE})
  • More dtype handling
  • Multi-kernel scheduling
  • Structural lowering pass

Milestone 4: Robust Full-Spec Direction

Estimate: 400+ iterations

  • Broad tinyspec semantic coverage
  • Well-tested runtime and lowering failures
  • Robust memory planning
  • Multi-device decisions documented
  • Advanced codegen ops either implemented or explicitly out of scope

Milestone 5: Serious Full-Spec Coverage

Estimate: 500-800 iterations

  • Full core tinyspec behavior
  • Broad upstream tinygrad test coverage
  • Hardware/runtime-backed execution where appropriate
  • Clear software fallback policy where hardware is not appropriate
  • Stable performance/profiling story

Recommended Next Iterations (refreshed 2026-04-20, end of second 40-iter push)

Current state: 889 backend tests pass + 95 BSV unit tests. ~120 new Python tests added this push.

Second-push additional landings (iters 24-31):

  • clip / hardtanh renderer — nested WHERE+CMPLT with 2 float CONSTs now emits FMIN+FMAX per tile.
  • single-bound clamp renderer — clamp(min=c) / clamp(max=c) with c != 0 emit FMAX / FMIN per tile. clamp(min=0) still routes through RELU (semantically equivalent).
  • RELU pattern guard tightening — requires CMPLT's CONST operand to be 0.0, so clamp(max=c) for c != 0 no longer silently false-matches as RELU.
  • leaky_relu renderer — max(alpha*x, x) for 0 < alpha < 1 via FMUL + FMAX.
  • softplus min-const false-match guard — the float-min-const path now rejects kernels carrying transcendentals / reciprocals, so softplus stops silently rendering as minimum(x, 0).
  • x4 regression guard** — self-square tightens to require LOAD tails.
  • Broad float / int32 correctness sweeps — 17 float ops + 16 int ops compared against numpy reference to catch future regressions.

Recommended Next Iterations (refreshed 2026-04-20, mid second 40-iter push)

Current state: 883 backend tests pass + 90 BSV unit tests (VPU 48 + FpReducer 7 + PSUMBank 9 + SXU 6 + SxuPSUM 2 + CtrlPSUM 1 + CtrlOS 1 + CtrlDB 2 + CtrlDBDMA 2 + WSRAMDB 3 + ASRAMDB 3 + WeightDMA 3 + ActivationDMA 3 + TranscUnit 5).

Landings in this second push (iters 1-23):

  • Item #8 (OS MXU): PE accumulator-hold DONE. startOS skips array.clearAll on drain; start/startPsum now pre-clear so WS always runs fresh. clearArray() method + SXU_MXU_CLEAR opcode (24) + TASM MXU_CLEAR mnemonic. End-to-end sim tests cover OS back-to-back accumulate + MXU_CLEAR reset.
  • Item #4 (Dual-issue XLU): DONE. XLU dispatch advances pc immediately; do_xlu_collect_bg writes the result one cycle later (XLU.result has 1-cycle latency). Structural-hazard guard !xlu_busy on every XLU dispatch rule. Chained-XLU sim test covers the path.
  • Item #5 (DMA stubs): DONE stubs + integration. WeightDMA + ActivationDMA synthesize deterministic tile patterns into inactive DB bank; standalone TBs + ping-pong TbCtrlDBDMA where MXU dispatches tile A on active while DMA preloads tile B on inactive. TensorCore default-wiring to DB SRAMs deferred (breaks existing preload semantics; needs an API revamp).
  • Composite activations + renderers:
    • Self-cube x**3 renderer (MUL(x, MUL(x, x)) pattern).
    • Self-square renderer tightened — no longer false-matches x**4.
    • Swish / silu renderer (x * sigmoid(x)).
    • Softsign renderer (x / (1 + |x|)).
  • Reducer silent-bug round-up: Five pattern guards added to scalar / col / row reducers so the following compound kernels route to UNSUPPORTED instead of silently returning a wrong answer:
    • hardtanh / clip (relu false-match via multi-CMPLT).
    • softsign (abs false-match via abs-in-MUL-RECIPROCAL tree).
    • sum(abs(x)) (pre-reduce WHERE/CMPLT).
    • reciprocal().sum() / exp().sum() (pre-reduce transcendental).
    • sum(x*x) / sum(-x) / max(-x) (pre-reduce data-path MUL).
    • log(x+k), exp(x+k), sin(x+k) scaled renderers now require LOAD-terminated non-const factors, refusing compound shifts.
  • Correctness wins:
    • Float mean() now returns correct results (reducer post-op now remapped to FADD/FMUL for float reductions).
    • Float mean(axis=0) / mean(axis=1) on fused kernels (col/row reducers now detect post-reduction FMUL(const)).
    • Float softsign end-to-end correct.

Recommended Next Iterations (updated 2026-04-20, end of FIRST 40-iter push)

Current state: 1006 Python tests + 87 BSV unit tests (VPU 48 + FpReducer 7 + PSUMBank 9 + SXU 6 + SxuPSUM 2 + CtrlPSUM 1 + CtrlOS 1 + CtrlDB 2 + WSRAMDB 3 + ASRAMDB 3 + TranscUnit 5).

Landings in this push:

  • Item #6 (Transcendentals): DONE with Remez + range reduction for EXP2/LOG2/SIN/COS. Tensor.exp, tanh, sigmoid, cos, log, sqrt, rsqrt all accurate across wide ranges.
  • Item #8 (OS MXU): DISPATCH PATH LIVE via Controller.startOS() + SXU DISPATCH_MXU_OS opcode. PE accumulator-hold FSM pending.
  • Item #5 (DB SRAM): CONTROLLER INTEGRATION. .plain sub-interface lets Controller consume DB SRAMs; preload-parallel pattern proven in TbCtrlDB. DMA stub + TensorCore default-wiring pending.
  • Item #4 (Dual-issue): SCOREBOARD LIVE. xlu_busy / xlu_dst set on dispatch, cleared on collect. Parallel rule + RAW-stall pending.
  • Renderer additions: Tensor.cos, Tensor.log, Tensor.rsqrt, Tensor.square, multi-tile wide-input tanh/exp/sigmoid/cos/sin tests.

Highest-leverage follow-ups (ranked):

  1. Finish Item #8 — PE accumulator-hold mode in SystolicArray + Controller FSM variant swapping operand roles so DISPATCH_MXU_OS delivers a genuinely distinct dataflow. ~6-8 iters.
  2. Finish Item #4 — parallel XLU issue rule + RAW-hazard stall consuming the existing scoreboard. ~4-6 iters.
  3. Finish Item #5 — DMA stub issuing background writes, plus TensorCore wiring to use DB SRAMs by default. ~4-6 iters.
  4. Engine-to-engine forwarding in renderers — emit SXU_LOAD_VPU_RESULT / LOAD_XLU_RESULT to elide VRegFile round-trips in chained kernels. ~3 iters.
  5. Composite activations: softplus (log(1+exp(x))), gelu (uses tanh), x**3 (MUL(x, MUL(x, x))). ~2-3 iters each.
  6. Narrow-dtype int8 VPU — big scope. Blocks quantized INT8 inference kernels outside the MXU. ~8-10 iters.

Prior Recommended Next Iterations (2026-04-12)

Current state: 379 backend tests + 8 runtime bundle tests. Row/col/scalar reductions (SUM/MAX/MIN/PROD) fully hardware-backed. Movement ops: reshape, contiguous slice/shrink, scalar and unrolled row-expand working. SXU_DISPATCH_XLU_TRANSPOSE opcode exists in hardware but tinygrad permute/transpose lowering not yet wired.

Highest-value next work:

  1. Wire tinygrad permute(1,0) to SXU_DISPATCH_XLU_TRANSPOSE. Detect the 2D index-swap pattern in the UOp graph and emit LOAD/XLU_TRANSPOSE/STORE for single-tile cases; extend to multi-tile via a tile-of-tiles pass.
  2. RANGE-loop variant of row-broadcast expand (shape (N, M) with M<_COLS and outer RANGE over rows) — current renderer handles only unrolled.
  3. Hardware epilogue for multi-K-tile GEMM (currently still a numpy fallback).
  4. Replace remaining hand-written structural recognizer duplication with reusable TinyTPU descriptor builders.
  5. Dtype expansion beyond int32/bool: int8/uint8/int16 elementwise policy.
  6. Movement gaps: PAD (requires CMPLT bounds check), FLIP (LOAD index = const - STORE index), CAT (WHERE-based select between two buffers), true permute (non-transpose axis reorder).
  7. Multi-output tile scheduling cleanup and intermediate VMEM allocation.
  8. Fused kernel detection improvements for 2-op chains tinygrad emits (slice-then-sum currently fuses to wrong-offset sum).

Model-blocking gaps (from mnist_gan.py bring-up)

  • leaky_relu(alpha): tinygrad emits float mul for the negative slope. SW path: integer approximation via SHR + WHERE select.
  • tanh, log_softmax, float32 GEMM: require float datapath (major arch change). SW fallback via host numpy or replace with int/argmax where semantically acceptable for inference-only models.
  • backward / autograd: training is not supported; export int8 weights and run inference only.