This TODO estimates the remaining work to support the tinyspec surface in
tinygrad/spec/tinyspec.tex. It uses the current repo iteration style: one narrow,
tested, committed behavior per iteration.
- Broad functional tinyspec coverage: 250-400 iterations
- Robust hardware-backed and well-tested coverage: 500-800 iterations
Current coverage includes hardware-backed TinyTPU execution for: GEMM (multi-WMMA, batched, deep-K, wide-N, with hardware fused bias+ReLU epilogue); full int32/bool VPU elementwise (ADD/MUL/SUB/MAX/MIN/DIV/MOD/ ABS, CMP{LT,EQ,NE}, AND/OR/XOR/NOT, SHL/SHR, WHERE, RELU, clip, fused add+relu, hardsigmoid, elu, mish); scalar-const variants and scalar broadcasting; all shapes with numel>16 through multi-tile elementwise loops; scalar, row-wise, and column-wise reductions (SUM/MAX/MIN/PROD) for any NxM through SXU_PROGRAM emitting VPU_{SUM,MAX,MIN,MUL}_REDUCE{,_COL,_TILE}; movement ops (reshape, contiguous slice/shrink, scalar expand, unrolled row expand); XLU transpose reachable via SXU_DISPATCH_XLU_TRANSPOSE at runtime level; TASM bundle assembler/disassembler; runtime bundle roundtrip + end-to-end sim tests.
- Track tinyspec source in
tinygrad/spec/tinyspec.tex - Runtime co-simulation bundle format for MXU programs
- Runtime VMEM preload and VMEM result output path
- Tinygrad GEMM lowering for supported tiled int32 cases through 4x4 MXU
- Multi-row, batched, deep-K, wide-N GEMM test coverage
- Tinygrad int32 single-tile VPU binary lowering
-
ADD -
MUL -
MAX -
SUB
-
- Tinygrad int32 single-tile VPU unary lowering
-
RELU
-
- Tinygrad 4-element int32 reduction lowering
-
SUMviaVPU_SUM_REDUCE
-
- Full 16-lane VMEM tile coverage
-
ADD -
MUL -
MAX -
RELU
-
- Signed VPU test coverage
- Runtime output validation for MXU and VMEM result lines
- Remove generated
bdpi/tinytpu_io.ofrom git tracking - TASM bundle assembler and disassembler (
scripts/tasm.py,doc/tinytpu_asm.md) - Full-tile and multi-tile abs, IDIV, MOD coverage
- Scalar broadcast (size-1 tensor) for MUL, SUB, MAX, multi-tile ADD/MUL
- int32→bool and bool→int32 cast coverage
- Clip (MIN+MAX program) full-tile and multi-tile
- Fused add+relu full-tile and multi-tile
- Tensor-tensor IDIV and MOD
- Row-wise sum/max/min for NxM tensors — hardware-backed via VPU_{SUM,MAX,MIN}_REDUCE through SXU_PROGRAM for all N,M (legacy VPU_ROWSUM/HOST_ROWREDUCE removed)
- Column-wise sum/max/min for NxM tensors — hardware-backed via VPU_*_REDUCE_COL for all N,M (legacy HOST_COLREDUCE removed)
- Fix stale WAIT_MXU opcode in VPU-only test bundles
- 2D tensor ops: all VPU binary/unary ops for arbitrary 2D shapes
- Grouped scalar-const lowering for 2D/large tensors (NEG, x*c, x+c)
- _tasm helper functions + bundle builders rewritten in readable assembly style
- TASM bundle roundtrip tests for all bundle builders
- Milestone 1: remove legacy
GEMM4x4- GEMM fallback (MULACC or scalar MUL+RANGE, no WMMA UOp) now emits
SXU_PROGRAMwith the same data_plan/instructions as the WMMA SXU path _exec_gemm4x4executor deleted
- GEMM fallback (MULACC or scalar MUL+RANGE, no WMMA UOp) now emits
- Milestone 2: remove legacy
VPU_BINARY- Scalar-const IDIV, tensor-tensor IDIV, and bool→int32 cast now lower to
SXU_PROGRAM _exec_vpu_binaryexecutor andVPU_BINARYemitter deleted
- Scalar-const IDIV, tensor-tensor IDIV, and bool→int32 cast now lower to
- Milestone 3: remove legacy
VPU_PROGRAM- Scalar-const MOD, tensor-tensor MOD, CLIP, ABS, and fused-add-relu already flow through SXU renderers
_exec_vpu_programexecutor deleted; the analyzer no longer emitsVPU_PROGRAM
- Milestone 4: remove legacy
HOST_BINARY- Dead code: no emitter ever produced
HOST_BINARY; executor and supported-op entry deleted
- Dead code: no emitter ever produced
- Milestone 5: remove legacy
HOST_UNARYRECIPROCALlowered viaFRECIPSXU program;TRUNClowered viaF2I+I2FSXU program_exec_host_unaryexecutor andHOST_UNARYemitter deleted
- Milestone 6: simplify runtime after descriptor removal
_SUPPORTED_OPS = {"SXU_PROGRAM"}_render_legacy_descriptornow returns only SXU descriptors (the GEMM fallback builds an SXU_PROGRAM)analyze_tinytpu_uopsdeleted; ONNX trace reporting now reads renderer descriptors directly
Replacing the ~47-recognizer _render_*_sxu_program waterfall with a single
UOp-walking instruction-selection pass in
tinygrad/tinygrad/runtime/support/tinytpu_lowering.py. See
doc/plan-tinytpu-instsel.md.
- Iteration 1: InstSel module —
TpuInst/TpuKernel/encoder and the walker for int32/bool elementwise (ALU, WHERE, scalar broadcast, multi-tile).render()routes these kernels to the walker via the positivecan_lowerpredicate. - float elementwise through the walker (VECTORIZE/GEP see-through, float
CMPNE/CMPEQ, transparent bool→int cast);
_render_elementwise_sxu_programand orphaned_find_alu_constdeleted - transcendentals EXP2/LOG2/SIN routed through the walker (single VPU opcodes 51/52/53)
- RECIPROCAL (FRECIP opcode 23) and SQRT (InstSel graph-rewrite to exp2(0.5·log2(x))) routed through the walker
- deleted 34 transcendental/activation/elementwise
_render_*recognizers and 3 orphaned pattern helpers (3062 lines);ops_tinytpu.py6984 → 3551 lines - linear-scan VREG allocation with reuse (deep DAGs no longer burn a VREG per node)
- divmod/trunc through the walker (MOD InstSel rewrite
a-(a//b)*b, TRUNC F2I+I2F micro-pair); deleted_render_truncand_render_scalar_const_divmod. Negative///%now match tinygrad/numpy floor semantics — 4 tests that encoded the legacy recognizer's truncating result were corrected.
InstSel migration slice complete. Every elementwise / unary / transcendental
/ activation / divmod / trunc kernel is lowered by the UOp-walking InstSel pass
in tinytpu_lowering.py. ops_tinytpu.py shrank from 6984 to ~3400 lines. The
structural recognizers (reductions, broadcasts, pad, transpose, cast, copy)
remain a deliberate later slice.
- Step 1 (ITER29): package split.
tinytpu_lowering.pysplit intotinytpu_lowering/package:common.py(shared types, opcode tables, graph helpers, InstSel pass),elementwise.py(walker),__init__.py(re-exportscan_lower/lower_kernel). Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). Seedoc/plan-tinytpu-instsel-structural.md§6. - Step 2 (ITER30): classifier.
classify.pyintroducesKernelClassenum (ELEMENTWISE,GEMM,UNSUPPORTED) andclassify(uops)function.render()dispatches onKernelClass.ELEMENTWISE→lower_kernel; all other classes fall through to the structural recognizers unchanged. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). - Step 3 (ITER31): reduction lowerer.
reduction.pyaddsis_reduction(uops)(positive predicate) andlower_reduction(uops), a single classify-then-emit lowerer for scalar/row/column SUM/MAX/MIN/PROD reductions.KernelClass.REDUCTIONis checked most-specific-first (after GEMM, before ELEMENTWISE);render()dispatches it tolower_reduction. The three legacy recognizers (_render_reduction/_render_rowreduce/_render_colreduce, ~565 lines) and orphaned helpers_is_float_min_negation/_detect_reduce_opare deleted. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). Seedoc/plan-tinytpu-instsel-structural.md§7. - Step 4 (ITER32): broadcast lowerer.
broadcast.pyaddsis_broadcast(uops)(positive predicate) andlower_broadcast(uops), a single classify-then-emit lowerer for row / column / column-where broadcasts. The axis is classified from the smaller operand's index relationship to the output ranges;colbc_whereemits the column broadcast then the WHERE/select.KernelClass.BROADCASTis checked most-specific-first (after GEMM and REDUCTION, before ELEMENTWISE);render()dispatches it tolower_broadcast. The three legacy recognizers (_render_rowbc/_render_colbc/_render_colbc_where, ~300 lines) and the orphaned helper_classify_structured_broadcast_axisare deleted. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). Seedoc/plan-tinytpu-instsel-structural.md§8. - Step 5 (ITER33): trivial kernels. cast / copy / const-fill folded
into the elementwise walker as degenerate elementwise maps.
can_lower/lower_kernelnow accept a bare-LOADDAG (copy, with a constant source offset for contiguous reshape / slice / shrink), a bare-CONSTDAG (single-tile const-fill), andCASTvalue converts (I2F/F2I; bool→int stays transparent via_canon)._render_cast_sxu_programand_render_const_fill_sxu_programare deleted in full along with the orphaned_ALU_OP_NAMES._render_copy_sxu_program(186 lines) is reduced to_render_rowbc_copy_sxu_program— the one non-degenerate case it carried, a single-input row broadcast (Tensor([[..]]).expand(N,M)), which is a structured broadcast, not a per-element map. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). Seedoc/plan-tinytpu-instsel-structural.md§6. - Step 6 (ITER34): GEMM relocate. Behavior-neutral relocation, two
parts. (A) All bundle-instruction encoders (
_vmem/_wmem/_amem,_load/_store/_vpu/_vpu_bg/_vpu_exp2/_vpu_log2/_vpu_sin,_select,_broadcast*,_mxu*,_psum*,_loop_*,_vzero/_vfill/_vmov/_vneg/_vabs,_set_pred_*/_skip_*,_halt/_output_*/_end/_bundle, …) moved verbatim intocommon.py;ops_tinytpu.pyre-exports them so existingtests//scripts/imports keep working.reduction.py/broadcast.pylocal encoder copies deleted and imported fromcommon.py. (B) Newgemm.pywithlower_gemm/lower_gemm_fallback: the WMMA branch of_render_sxu_program,_render_gemm_fallback_sxu_program, and the GEMM-only helpers (_generate_gemm_sxu_instructions,_extract_wmma_epilogue,_apply_gemm_epilogue,_infer_tiling,_tiling_failure_note) moved verbatim._find_unique_param_argmoved tocommon.py(shared with the structural recognizers).render()dispatchesKernelClass.GEMMtolower_gemm; the non-WMMA matmul fallback runs vialower_gemm_fallback. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions); cosim passes. Seedoc/plan-tinytpu-instsel-structural.md§9. - Step 7-8 (ITER35): movement lowerer (Branch B). The Task 7 spike
chose Branch B (renderer-side lowering): rangeify already dissolves the
movement op itself, but the recognizers recover a non-affine tile access
pattern the SXU model cannot express generically — they do real
SXU-specific instruction selection, so they are relocated, not deleted.
New
movement.pyaddsis_movement(uops)(positive predicate) andlower_movement(uops), a classify-then-emit lowerer that ports the three legacy recognizers verbatim: pad / flip / non-affine permute (PAD_FILL scatter data plan), the 4×4permute(1,0)(SXU_DISPATCH_XLU_TRANSPOSE), and single-input row-broadcast copy (SXU_BROADCAST_ROW).is_movementcarries the legacy fallback-ordering gate (can_lowerkernels — e.g. a scalarexpandlowered asBROADCAST_SCALAR— were never reached by the recognizers and are excluded here).KernelClass.MOVEMENTis checked most-specific-first (after GEMM/REDUCTION/BROADCAST, before ELEMENTWISE);render()dispatches it tolower_movement. The three legacy recognizers (_render_pad/_render_transpose/_render_rowbc_copy_sxu_program, ~285 lines) and orphaned helpers_uop_contains/_has_load_src/_data_alu_ops/_ALU_OPSare deleted.ops_tinytpu.pynow has zero_render_*_sxu_programrecognizers;_render_sxu_programis a trivialreturn Noneshell (Task 9 deletes it). Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions). Seedoc/plan-tinytpu-instsel-structural.md§10. - Step 9 (ITER36): final cleanup. The empty
_render_sxu_programshell and its call site inrender()are deleted —render()is now a cleanclassify(uops)→ per-class lowerer dispatch, with the non-WMMA matmul fallback (lower_gemm_fallback) preserved on the path before theUNSUPPORTEDdescriptor. Duplicated graph helpers_has_load_src/_data_alu_ops(byte-identical copies inmovement.py/reduction.py/broadcast.py) are hoisted into one canonical copy incommon.py. Theid()-based traversal closures inmovement.py(_has_range,_consts_in, load-index distinctness set) are replaced withUOp.toposort()-based equivalents. The provably-deadWMMAreject entry is removed from the movement recognizers (classifyroutes WMMA → GEMM beforeis_movement); theMULACCguard is kept (a non-WMMA matmul can reachis_movement, so it is not provably dead). Dead_apply_gemm_epilogue(zero callers) is deleted fromgemm.py. A deferred TODO is added atbroadcast.py:_classify_broadcast_axis. Behavior-neutral; suite unchanged (810 pass / 88 fail, 0 regressions); cosim passes.
Structural slice complete. All 12 structural recognizers migrated;
ops_tinytpu.py has zero _render_* recognizers and is 1260 lines (from
6984); all kernel lowering is in the tinytpu_lowering package behind
classify().
Unmasked hardware bug: the walker faithfully lowers tinygrad's
decompositions, which exposed that the BSV EXP2/LOG2/SIN units are broken
(EXP2 returns ~0 for positive inputs; LOG2 is very imprecise). The old
activation recognizers used hardware-dodging formulations that hid this. ~40
transcendental/activation test failures, plus test_elu_positive_branch and
test_log_compound_input, trace to this. Fix belongs in src/ BSV, not the
backend — see doc/plan-primitive-ops-handoff.md.
End-to-end tests for the five CODA epilogue primitive classes (elementwise/
pairwise maps, vector loads/stores, tile loads/stores, tile reductions,
stateful transforms) live in tests/test_e2e_epilogues.py. 18/19 pass; the
following gaps are explicit:
- Column-vector
(M,1)broadcast over(M,N)mis-lowered as elementwise multiply._classify_colbcpicked the VPU op by the first non-zero op_count, so a single address-arithmeticMULshadowed the realADD. Fix: match the compute op by finding a binary UOp whose LOAD sources are the two input params (tinygrad/tinygrad/renderer/tinytpu/broadcast.py). Verified for ADD/MUL/SUB/rSUB/MAX with(M,1)operand. - Inline reduce inside an add chain silently produced garbage. Root
cause:
lower_gemm_fallbackadmitted any 3-param kernel with a single MUL + RANGE + STORE. An address-arithmetic MUL (RANGE * stride) was enough, so the unsupported fused-reduce-add kernel got phantom-lowered as a 1x1x1 GEMM. Fix: require the MUL to multiply two LOADed values (the actualacc += a*bsignature). The fused kernel now raises NotImplementedError. The testtests/test_e2e_epilogues.py::TestStatefulEpilogues::test_inline_reduce_in_add_raisespins this behavior. A proper fused lowerer for(M,1) + reduce(tile, axis=1)is still open — callers lift the reduce to a named intermediate today.
- Make VPU-only programs complete without dummy MXU dispatch
- Add first-class runtime bundle builder for VMEM/VPU programs
- Support multiple VMEM output tiles (col-reduce/row-reduce/GEMM emit multi-tile outputs through SXU_PROGRAM)
- Support multiple VPU instructions in one tinygrad lowered program (SXU_PROGRAM multi-step paths: abs, clip, MOD, WHERE, row/col reductions)
- Add runtime tests for VMEM preload/output protocol
- Improve trace output for VPU-only programs
- Document bundle records for VMEM input/output
- Fuse GEMM epilogues in model runs
- Post-run review on
scripts/models/cnn_4_8_8_4.pyshowed the model runs end to end without host fallback orUNSUPPORTED, but each layer still lowers asMXU -> LOAD bias -> VPU ADD -> VPU RELU -> STORErather than a first-class fusedmatmul + bias + reluepilogue path. - Add a direct runtime/compiler path so model kernels stop depending on separate bias-load and VPU epilogue instructions after every MXU tile.
- Post-run review on
- Single-tile int32
ADD - Single-tile int32
MUL - Single-tile int32
MAX - Single-tile int32
RELU - Constants in VPU programs, e.g.
x + 1,x * 2-
x + scalar -
x * scalar -
maximum(x, scalar) -
minimum(x, scalar) -
x < scalar -
x != scalar
-
- Scalar broadcasting
- Add explicit SXU/XLU broadcast opcodes (scalar/row/col)
- Migrate scalar broadcast binary ops to SXU_PROGRAM via
BROADCAST_SCALAR
- Size-1 axis broadcasting (scalar-expand via BROADCAST_SCALAR; row-expand via BROADCAST_ROW for unrolled shapes)
- Column-broadcast compare/select lowering for tinygrad workloads
(4,4) < (4,1)now lowers throughSXU_PROGRAMwith theBROADCAST_COLprimitive instead of falling through toUNSUPPORTED.- The post-run review model path
mask.where(y, y * 2)now closes through a dedicated single-tileBROADCAST_COL_SELECTSXU program when the compare and select stay fused in one kernel. - Follow-up: a separately realized bool mask still exposes a different mixed bool/int32 arithmetic gap in the generic
VPU_BINARYpath; keep that as a distinct issue rather than regressing the fused review-model path.
- Arbitrary shapes with
numel <= 16(1D/2D/nD elementwise + reshape/slice/expand covered by copy + elementwise renderers)- Shape-preserving 2x2 elementwise coverage for supported VPU ops
- Multi-tile elementwise loops for
numel > 16- Multi-tile AND/OR/XOR/NOT for bool tensors
- Mixed VPU op chains without host round trips (via VPU_PROGRAM path)
- Output shape preservation for scalar, vector, and small matrix cases (1D/2D/nD coverage via copy + elementwise + reduction renderers)
-
SUBasADD + NEG, lowered toVPU_SUB -
NEGas multiply by-1 -
CMPLT -
CMPNE -
CMPEQ -
WHERE -
AND -
OR -
XOR -
NOT -
SHL -
SHR -
MOD(via DIV+MUL+SUB bundle) -
IDIV(via VPU_DIV, truncation semantics) -
RECIP -
TRUNC - Basic
CAST(int32↔bool; int32↔float32 via VPU_I2F/F2I; fused same-dtype round-trips via COPY) - Basic
BITCAST - Basic
COPY(_render_copy_sxu_program handles same-dtype identity-index kernels via LOAD/STORE pairs)
- 4-element int32 sum to scalar
- Full-tile int32 sum to scalar (via VPU_SUM_REDUCE_TILE)
- Multi-tile int32 sum to scalar (VPU_SUM_REDUCE_TILE per tile + VPU_ADD combine)
- Row-wise sum/max/min over NxM tensor (SXU_PROGRAM row-reduce renderer, all N,M)
- Column-wise sum/max/min over NxM tensor (SXU_PROGRAM col-reduce renderer, all N,M)
- Full-tile sum (via VPU_SUM_REDUCE_TILE)
-
MAXreduction (4-elem, full-tile, multi-tile via VPU_MAX_REDUCE_TILE) -
MINreduction (4-elem, full-tile, multi-tile via VPU_MIN_REDUCE_TILE) - VPU col-reduce primitives (VPU_SUM/MAX/MIN_REDUCE_COL opcodes 29/30/31)
- VPU tile-reduce primitives (VPU_SUM/MAX/MIN_REDUCE_TILE opcodes 32/33/34)
-
MULreduction (VPU_MUL_REDUCE{,_COL,_TILE} hardware + tinygrad lowering for scalar/row/col prod) -
keepdimbehavior (rowsum/rowmax/rowmin keepdim tests pass) - Multi-tile reductions (scalar SUM/MAX/MIN with VPU_*_REDUCE_TILE per tile + VPU_ADD/MAX/MIN combine)
- Reduction axis shape validation
- Reduction result layout in VMEM
-
RESHAPE(copy SXU_PROGRAM with identity LOAD/STORE index mapping) -
PERMUTE -
TRANSPOSEvia XLU — hardware: SXU_DISPATCH_XLU_TRANSPOSE opcode (12) + runtime test done; tinygrad permute(1,0) lowering still open -
EXPAND(scalar and unrolled row-broadcast via BROADCAST_SCALAR / BROADCAST_ROW; RANGE-loop variant still open) -
SHRINK(contiguous slice via copy renderer with affine offset) -
PAD -
FLIP -
CAT -
INDEXwith simple affine patterns (contiguous slice with constant offset through copy renderer) - Gather-like indexing
- XLU-backed broadcast primitives (scalar/row/col)
- XLU-backed permutation paths
- Multi-tile movement kernels
- 4x4 GEMM
- Multi-row GEMM
- Batched GEMM
- Deep-K tiled GEMM
- Wide-N tiled GEMM
- Multi-WMMA lowering (multi-tile GEMM through WMMA path)
- Bias epilogue (row-broadcast and full-tensor, hardware-backed for single-K-tile)
- ReLU epilogue (hardware-backed for single-K-tile)
- Fused bias+ReLU epilogue (hardware-backed via SXU_LOAD_MXU_RESULT → VPU)
- M/N/K tail handling
- Better unsupported shape diagnostics
- Add/mul epilogue
- Non-int8 operand policy
- Accumulation overflow tests
- Multi-output tile scheduling cleanup
- Hardware epilogue for multi-K-tile GEMM — landed via PSUM
bucket bank. Multi-K-tile GEMMs now run
SXU_PSUM_CLEAR→ N ×DISPATCH_MXU(psum_acc)→SXU_PSUM_READ_ROW→ bias/relu → STORE entirely in hardware. No numpy fallback path remains.
- int32 VMEM values for VPU paths
- int8 operands for MXU paths
- bool comparison outputs
- int8 elementwise
- uint8 elementwise
- int16 elementwise
- uint16 elementwise
- uint32 elementwise
- float32 policy (FADD/FSUB/FMUL/FMAX/FCMPLT dispatched via VPU_F* variants; scalar const and multi-tile covered; FRECIP+scalar fdiv work; tensor-tensor fdiv requires RECIPROCAL UOp detection, tracked)
- cast saturation/wrapping behavior
- comparison output dtype behavior
- dtype range diagnostics
-
CALL -
TUPLE -
GETTUPLE -
AFTER - assign/store dependency ordering
- multiple outputs
- multiple stores in one kernel
- reusable captured graph fragments
- Multiple TinyTPU programs per tensor expression
- Intermediate VMEM allocation
- Spill/fill to HBM model
- Host/device copy scheduling
- Cross-kernel dependency tracking
- Runtime buffer lifetime tests
- Larger tensors split across VMEM tiles
Current bloat sources:
- Structural lowerers in
tinygrad/renderer/tinytpu/*still contain legacy recognizer ports. Convert repeated descriptor/data-plan construction into reusable builders. - Bundle builders (~400 lines):
_build_vpu_binary_bundle,_build_vpu_where_bundle,_build_full_gemm_bundle, etc. Unify into a single generic bundle builder with a template pattern. _exec_*methods (~400 lines): every op has the same chunk loop (for chunk_start in range(0, num_elems, _TILE_ELEMS)). Extract a shared_run_tiled_vpuhelper.- Duplicate output parsing:
_parse_vmem_output,_parse_multi_vmem_output,_parse_sim_outputshare the same structure.
Cleanup plan — eliminate analyze_tinytpu_uops via SXU_PROGRAM migration:
- Add VPU_NOT hardware opcode (BSV + testbench + Python table)
- Migrate scalar-const binary ops (x+c, x*c, NEG, NOT) to SXU_PROGRAM
- Migrate bool-typed ops (AND/OR/XOR/NOT on bool tensors) to SXU_PROGRAM
- Migrate WHERE (ternary select) to SXU_PROGRAM (now via first-class SXU_DISPATCH_SELECT using VPU_SELECT)
- Migrate multi-step VPU_PROGRAM patterns (abs, clip, MOD, CMPEQ) to SXU_PROGRAM
- Migrate scalar reductions (SUM/MAX/MIN to scalar) to SXU_PROGRAM
- Migrate row-wise reduce to SXU_PROGRAM (done: _render_rowreduce_sxu_program, legacy VPU_ROWSUM removed)
- Migrate row-broadcast binary (VPU_ROWBC_BINARY) to SXU_PROGRAM
- Emit unsupported diagnostics directly from renderer descriptors
- Delete dead analyze blocks, _exec_vpu_where/unary, _build_vpu_where, _render_wmma_descriptor
- Delete remaining analyze_tinytpu_uops
- Run selected upstream tinygrad tests on
TINYTPU(2 tests pass via scripts/run_tinytpu_upstream_subset.py — expand list as coverage grows) - Add skipped/xfail manifest for unsupported tinyspec areas
- Track coverage by tinyspec op category
- Trace parser validation
- Perfetto emission for existing traces
- Full VPU-only trace coverage
- VMEM bundle dump utility
- Lowering decision dump (TINYTPU_DUMP_LOWERING env var emits renderer JSON)
- Runtime bundle round-trip tests for all record types
- Per-op cycle reports
- MXU/VPU/VMEM utilization by lowered tinygrad op
- Multi-device
COPY -
REPLICATED - collectives: broadcast, scatter, gather, reduce, allgather
- reduce_scatter
- allreduce
- optimizer axis semantics
-
BARRIER -
SPECIAL -
IF/ENDIF -
WMMAmapping beyond current MXU path -
CUSTOM -
ATOMICADD -
CUSTOMFUNCTION -
PROGRAM/SOURCE/BINARYmetadata handling
These are the remaining hardware-level gaps that cannot be solved by software lowering alone. Ordered roughly by impact on real workloads.
- Float reducer tile-scope primitives
VPU_FSUM_REDUCE_TILE(opcode 38),VPU_FMAX_REDUCE_TILE(39),VPU_FMIN_REDUCE_TILE(40). Float sum/max scalar reductions lower end-to-end; float min lowers for single-tile kernels via a negation-around-max pattern rewrite. - Multi-tile float min combine —
VPU_FMINALU opcode (41) added; renderer emits per-tileFMIN_REDUCE_TILEplusVPU_FMINacross tiles. Works up to 32 elements; larger sizes hit the pre-existing RANGE-loop reduction gap shared with other scalar reductions. - Float row-reduce primitives (
VPU_FSUM_REDUCE42,VPU_FMAX_REDUCE43,VPU_FMIN_REDUCE44). Per-sublane float reducer ops added; row-sum and row-max are wired through the tinygrad row renderer. Row-min still needs negation-decomp detection in the row path. - Float col-reduce primitives (
VPU_FSUM_REDUCE_COL45,VPU_FMAX_REDUCE_COL46,VPU_FMIN_REDUCE_COL47). Landed via the "hoist once outside the case block" pattern (same as the integer col reducers) using sharedlane_f*helpers. Col-sum and col-max wired through tinygrad. Col-min pending negation rewrite. - Shared multi-cycle FP tile reducer —
mkFpReducersub-module with one FP adder + one comparator driven by a small FSM. Tile-scope FP reducers now dispatch through it; SXU stalls onvpu.isDoneuntil the walk completes. Runtime build went from 1:48 to 1:29 after replacing the 3 combinational tile FP trees, and adding 3 col reducer opcodes on top cost only +7s (1:36), meeting the <10s/opcode acceptance bar. - Float prod reducer — no opcode; tinyspec
Reduce(Mul, axes)on float would need aVPU_FMUL_REDUCE*family.
- Shared multi-cycle FP reducer unit —
src/FpReducer.bsvlanded. Tile-scope FP reducers now route through it; acceptance criterion met (+7s for 3 new col reducer opcodes). Row/col reducers still use the "hoist once" combinational pattern with shared lane_f* helpers, which is cheap at the current 4×4 lane count. If we scale lanes or add more reducer kinds (e.g. float prod), consider routing row and col through FpReducer too. - Narrow-dtype VPU lanes: int8, uint8, int16, uint16, uint32 elementwise. Current VPU is int32-only on the data path (except MXU which is int8). Blocks quantized-int8 inference kernels outside the MXU, and uint dtypes entirely.
- Mixed-precision MXU: MXU is int8-only today. No bf16/fp16 MAC unit, no fp32 accumulator path for floating GEMM. Float32 GEMM currently can't lower at all.
-
Exp2/Log2/Sin/Coshardware — Remez + range reduction. All four primitives live insrc/TranscUnit.bsv: - Opcodes:VPU_EXP2(51),VPU_LOG2(52),VPU_SIN(53),VPU_COS(54). All dispatch through the multi-cycle walker; VPUisDonegates SXU collect. - Coefficients: Remez minimax for all four (EXP2 4× peak-error reduction, SIN 40×, COS 27×, LOG2 8× vs Taylor). - EXP2 range reduction: x = n + f viatr_trunc/tr_fp_to_int/tr_pow2_inthelpers; poly(f) on [-1, 1] scaled by 2^n via exponent-bit construction. Integer x gives exact 2^n results. - SIN range reduction: mod-2π + quadrant fold + sign-aware round-to-nearest bias. sin(π), sin(2π) etc. zero out; sin(5) accurate to <0.3% rel. - Tinygradcode_for_opdeclaresOps.EXP2/LOG2/SIN/SQRThardware-supported. Renderers land Tensor-level:_render_scaled_exp2(Tensor.exp),_render_sigmoid(FMUL+EXP2+FADD+FRECIP),_render_tanh,_render_scaled_sin(Tensor.cos),_render_scaled_log2(Tensor.log), plus plain_render_{exp2,log2,sin,sqrt}_sxu_program. - Guards in elementwise/reciprocal/multistep-scalar-divide renderers reject kernels with EXP2/LOG2/SIN/SQRT so missing composite renderers fail loud. - StandaloneTbTranscUnit(make test-transcunit) locks the coefficient contract at the unit layer. - End-to-end backend tests cover wide ranges: tanh |x|≤2.5, exp [-3,3], sin/cos [-3π,3π], sqrt [0.25,16].
- [~]
PADprimitive — small unrolled pads lower through the_render_pad_sxu_programrenderer (PAD_FILL VMEM preload). Multi-tile + RANGE-loop pads still open. - [~]
FLIPprimitive — unrolled flip / non-affine permutes lower through the pad_fill renderer too. Multi-tile FLIP and RANGE-loop FLIP still open. -
CATprimitive — concatenate along an axis. Could be a two-source SXU copy program driven by a range split, but no primitive exists today. - General
Permute— onlypermute(1,0)(transpose) is wired throughSXU_DISPATCH_XLU_TRANSPOSE. Non-transpose axis reorders for ≥3D tensors have no hardware path.
- Multi-K-tile hardware epilogue — landed via PSUM bucket
bank (Item #2 below). Multi-K-tile GEMM now zeros a PSUM bucket
via
SXU_PSUM_CLEAR, accumulates every K-tile into row 0 viaDISPATCH_MXUwithpsum_accmode, then extracts the row withSXU_PSUM_READ_ROWand runs the existing bias+ReLU epilogue on the PSUM result. No numpy fallback path remains for multi-K-tile bias/relu.
-
BARRIER/IF/ENDIF/AFTERSXU ops — SXU today is a straight-line sequencer. No conditional execution, no explicit ordering barriers for RAW hazards across dispatches. -
ATOMICADDunit on VMEM banks — no RMW lane on the VMEM write port. Required for scatter-add style kernels.
- Collective primitives (broadcast/scatter/gather/allreduce/ reduce_scatter) expressed as NoC-backed hardware ops. NoC exists at the ring level but is not exposed as a tinyspec-shaped primitive.
-
REPLICATED/ multi-deviceCOPYmodel — the compiler has no concept of cross-chip axis today. -
THREEFRYPRNG datapath — no ARX unit. All random init currently falls back to host.
- VMEM spill/fill to HBM — compiler-managed spill path for tensors larger than the VMEM working set. Today everything must fit in VMEM tiles.
These are the structural improvements that came out of reviewing the TensorCore datapath against the NeuronCore reference and thinking about what's actually bottlenecked today. Ordered by impact/effort ratio — the top ones would most change what TinyTPU can run.
- Decouple SXU into per-engine command queues. The single
microsequencer is the visible biggest bottleneck: every engine
shares one issue slot. MXU already overlaps via
DISPATCH_MXU+WAIT_MXU, but VPU / FpReducer / XLU all serialize even though they touch disjoint state. Split the front-end into an SXU fetch/decode stage feeding four small per-engine FIFOs, each with its own execute-side controller; add a scoreboard keyed on VRegFile destination index for RAW/WAW. Closes the biggest visible gap vs NeuronCore's per-engine SEQs. - Dedicated PSUM accumulator bank. Landed.
src/PSUMBank.bsvis an 8-bucket × 4×4 Int32 SRAM withwrite/accumulate/writeRow/accumulateRow/readReq/readResp/peekRow/clearand row-granular access for MXU. Shared between Controller (per-dispatch row deposit viastartPsum) and SXU (opcodes 15-19:PSUM_WRITE,PSUM_ACCUMULATE,PSUM_READ,PSUM_READ_ROW,PSUM_CLEAR). The tinygrad renderer uses the bank for multi-K-tile GEMM accumulation, which closes the "Multi-K-tile hardware epilogue" gap above. - Engine-to-engine forwarding (bypass VRegFile). Today every
hand-off is
engine → resultReg → VRegFile → engine, four cycles. Add a muxed forwarding network on engine inputs sovs/vs2can come straight from the MXU result reg / VPU resultReg / XLU output when SXU emits aFWDhint instead of a vreg index. The currentSXU_LOAD_MXU_RESULTspecial case is the first half of this feature — generalize it. - Dual-issue slot for VPU + XLU. Landed. XLU dispatch rules
now advance pc and return to FETCH immediately instead of
stalling in an XLU_COLLECT state; a background
do_xlu_collect_bgrule fires the cycle after dispatch (XLU.result carries 1-cycle latency) and writes xlu_dst while the main FSM fetches/executes the next non-XLU op. Structural-hazard guard!xlu_busyon every XLU dispatch rule prevents back-to-back XLU ops from issuing before the prior one collects. No RAW hazard in practice because the collect's vrf.write happens one cycle before any dependent read could fire (dispatch advances pc through FETCH before any EXEC reads vreg). End-to-end chained XLU broadcast sim test covers the path.
- Double-buffered Weight/ActivSRAM + small DMA engine.
src/WeightSRAMDB.bsvandsrc/ActivationSRAMDB.bsvare now the default SRAMs inside TensorCore. The Controller consumes the.plainsub-interface (writes → INACTIVE, reads → ACTIVE); TensorCore exposesloadWeightTile/loadActivationTile(writeActive) for "preload then dispatch immediately" flows andpreloadWeightTile/preloadActivationTile+swapWeightBanks/swapActivationBanksfor DMA-overlap flows. Covered bymake test-tc-db(ping-pong dispatch) plus the existing WeightDMA / ActivationDMA standalone benches. - Transcendental / programmable SIMD unit.
sqrt/log2/exp2/sinare rejected today; NeuronCore covers these with its GPSIMD engine. Three options in increasing generality: (a) polynomial-approximator opcodes with fixed hardware LUTs, (b) a dedicated transcendental unit next to the VPU, (c) a tiny programmable SIMD lane (RISC-V style). (a) is the natural first iteration. - Predicate register +
SKIP_IF_ZERO(babyIF/BARRIER). Landed.src/ScalarUnit.bsvgains a 1-bitpredReg plus two opcodes:SXU_SET_PRED_IF_ZERO(opcode 20, pred := vreg[0][0] == 0) andSXU_SKIP_IF_PRED(opcode 21, advances pc by 2 when pred is set and auto-resets pred). TASM opcode table and Python wire helpers (_set_pred_if_zero,_skip_if_pred) landed; runtime sim tests cover both skip-taken and skip-not-taken paths. No tinygrad renderer is emitting these yet — they're infrastructure for future BARRIER / IF / ENDIF work. - Output-stationary dataflow mode on MXU. Real OS dataflow is
now the primary path:
-
PE.feedPair(w, a)MACs with both operands arriving from neighbors;passWeightmirrorspassActivation. -SystolicArray.feedPair(w_top, a_left)drives one systolic wavefront per cycle;getMatrix()returns the full per-PE psum (rows x cols). Covered bymake test-array(2x2 OS matmul). -Controller.startOS(wBase, aBase, kLen)runs the full staircase:LoadWeights → LoadWeightsRespOS → LoadActsOS (K cycles) → StreamOS (K + (rows-1) + (cols-1) cycles) → Drain. Both operand staircase skews are applied inside the Controller.resultsMatrix()getter exposes the full output.make test-ctrl-oscovers 4x4 identity @ arange = arange. - SXU opcodeSXU_DISPATCH_MXU_OS(25) routes through it; TASMMXU_OS WMEM[w], AMEM[a], k=Nand Python_mxu_os(w,a,k)emit the wire form. The old misnomerMXU_OSopcode (23) was renamedMXU_ACCUMULATE— see the dataflow roadmap below. -SXU_LOAD_MXU_MATRIX_ROW(26) copies one row of the psum matrix into a vreg so consumers can walk the drain row by row. -SXU_DISPATCH_MXU_OS_ACCUMULATE(33) dispatches without the drain-time clear, letting multi-K-tile OS scale past K == rows. - Covered end-to-end:test_mxu_os_end_to_end_identity_matmul,test_mxu_os_end_to_end_nontrivial_matmul,test_mxu_os_accumulate_multi_tile. Lean theoremaccumulate_composeformalizes the multi-tile composition.
The four canonical systolic dataflows, ranked by TinyTPU fit:
- Weight-stationary (WS). Fully implemented. Weights loaded once, activations + partial sums flow. Good for weight-heavy layers; matches TPU / NVDLA.
- Output-stationary (OS). Fully implemented as of iter 42.
PE.feedPair+SystolicArray.feedPair+Controller.startOSrun the staircase; drain returns the full (rows x cols) psum. Dispatched via SXU opcode 25 (MXU_OS). High-value for depthwise conv and attention kernels. - WS-accumulate (K-batching on WS). Controller.startAccumulate
preserves the WS accumulator across dispatches; dispatched via
SXU opcode 23 (
MXU_ACCUMULATE). Not a distinct dataflow — a scheduling overlay on WS that skips the drain-time clear. Also covered at a different granularity by the PSUMBank. - Input-stationary / activation-stationary (IS). Deferred. ActivationSRAM stores (rows,)-shaped vectors per address, not (rows, cols) tiles, so IS would need either a pre-load buffer built from K separate reads or a new ActivationTileSRAM shape. The compute win is narrow (batch=1 inference where activations are large and weights stream cheaply from HBM), and the OS path now covers most of the same workload class. Revisit if a real workload needs it.
- Row-stationary (RS) — Eyeriss-style. Overkill for a 4x4 pure-matmul engine; skip unless the array grows + conv kernels become the bottleneck.
Hybrids worth noting:
- K-batching / multi-tile accumulation via
Controller.startAccumulateand/or PSUMBank accumulate buckets. - No-local-reuse (NLR): intentionally skipped — no reason to implement as a first-class mode.
- Shared on-chip L2 between TensorCores. NoC exists, but cores re-fetch overlapping weights from HBM for data-parallel kernels. An L2 cache or an NoC broadcast primitive amortizes weight traffic across cores.
- [~] Generalize
SXU_LOAD_MXU_RESULTinto engine-to-engine read ports. Hardware and TASM support landed:SXU_LOAD_VPU_RESULT(13) andSXU_LOAD_XLU_RESULT(14) both copy the engine's linger result register into a vreg. Bundle tests cover both. Remaining: tinygrad renderer is not yet using these opcodes to elide VRegFile round-trips in chained kernels (e.g. XLU-then-VPU, VPU chains that store and re-read). - Per-engine writeback resolved via VRegFile Vector-of-Reg.
Before:
mkRegFileFullhad one write port, so bsc forced every vrf writer mutex withdo_xlu_collect_bg(G0010 urgency conflicts — the dual-issue background rule was silently serialized). After:VRegFileisVector#(numRegs, Reg#(T)); each index has its own write port. G0010 conflicts collapse to G0036 (same-cycle ordering — only matters if two rules target the same index, which valid programs never do). 865 sim-backed backend tests clean.
New single-cycle opcodes added for instruction density and perf work:
SXU_LOAD_MXU_MATRIX_ROW(26) — drain one row of OS psum matrix.SXU_READ_CYCLE(27) — free-running cycle counter into lane 0. Backed by a new unconditionalcycleReg (was TRACE-only).SXU_LOOP_BEGIN(28) /SXU_LOOP_END(29) — single-level counted loop (counter in mxuTLen UInt#(8)). Renderers can emit a tight body without unrolling.SXU_VZERO(30) — single-cycle zero vreg. Used by the softsign renderer to replace a zero-tile VMEM preload.SXU_VFILL(31) — broadcast signed 8-bit immediate to all lanes.SXU_VMOV(32) — single-cycle vreg copy.SXU_DISPATCH_MXU_OS_ACCUMULATE(33) — multi-K-tile OS.SXU_VNEG(34) — single-cycle lane-wise negate.SXU_VABS(35) — single-cycle lane-wise absolute value. Int32 abs renderer now emits LOAD + VABS + STORE (3 instrs, was 5).
Renderer adoption of the new opcodes is an ongoing task — most renderers still emit the old multi-instruction patterns.
Opcode total: SXU 36 → 41; VPU 55 → 82.
SXU additions (36-40):
SXU_LOAD_LOOP_DEPTH(36) — read current loop-stack depth (for debug introspection of nested loops).SXU_DISPATCH_XLU_ROTATE(37) — exposes XLU's cyclic-rotate hardware that had no SXU entry point.SXU_PSUM_CLEAR_ALL(38) — multi-cycle walker inside SXU that zeros every bucket. 1 instruction replaces 8PSUM_CLEARs for multi-K GEMM epilogues.SXU_SET_PRED_NE_ZERO(39) /SXU_SKIP_IF_NOT_PRED(40) — completes the predicate pair (if-zero and if-nonzero, skip-if-pred and skip-if-not-pred, both arms ofif/elseexpressible).
SXU LOOP_BEGIN/LOOP_END now use a depth-4 stack (was single-level in push 3) — nested counted loops work up to 4 levels.
VPU additions (55-81):
- Packed-int8 arithmetic (55-67):
PACKED_I8_ADD/_SUB/_MAX/_MIN(55-58) byte-wise saturation;_NEG/_RELU(59-60) unary;_CMPLT/_CMPEQ(61-62) byte-wise compare;_MUL_LOW/_MUL_HIGH(63-64) multiply with wrap / Q-format;_ABS/_SIGN(65/67) unary. - Int32/float sign:
VPU_SIGN(66),VPU_FSIGN(68),VPU_FABS(79). Softsign + float-abs renderers adoptedFABS— drops 2 instructions per tile from both. - Reductions:
VPU_ARGMIN/VPU_ARGMAX(69/70) per-row index. - Bit-manip:
VPU_CLZ(71),VPU_POPCOUNT(72),VPU_CTZ(73),VPU_BYTE_REVERSE(74),VPU_ROTL(80),VPU_ROTR(81). - Saturating arithmetic:
VPU_SAT_ADD_I32(75),VPU_SAT_SUB_I32(76),VPU_ABS_DIFF_I32(77),VPU_PACKED_I8_ABS_DIFF(78).
- softsign (iter 26) — adopts
VPU_FABS: drops the VZERO + FSUB + FMAX 3-instruction sequence, keeps only LOAD(1.0)+FABS. - float-abs (iter 27) — adopts
VPU_FABS: drops the zero-tile VMEM preload + FSUB + FMAX (5 instrs → 3 per tile). - int32-abs (push 3) — already uses
SXU_VABS(5 instrs → 3).
Lean theorems added (push 4): zero_weight_accum_unchanged,
sign_idempotent, sign_odd, abs_step_preserves_weight,
abs_fold_preserves_weight, sign_bounded,
sign_times_self_nonneg — 7 new theorems on top of the 3 from
push 3 (10 total).
Dedicated SXU_DISPATCH_VPU_BG opcode (41) for background-collect
dual-issue. Unlike push-4 iter 2's attempt (which converted the
existing sync path and broke multi-cycle ops), this iter keeps sync
DISPATCH_VPU behavior identical and adds the BG path alongside.
The main FSM's do_fetch dispatches both opcodes through one rule
(do_vpu) that forks on pc_state: sync transitions to
EXEC_VPU_COLLECT and stalls; BG advances pc and returns to FETCH,
overlapping subsequent non-vreg-reader ops (MXU dispatch, PSUM
clear, loop control, skip-if-pred). A single unified do_vpu_collect
rule retires via vrf[vpu_wb_dst] for both paths — keeping the
VRegFile writer count at parity with push-4 HEAD is what avoids
blowing past bsc's -steps-max-intervals 20 elaboration budget.
Iter 3 added conservative !vpu_wb_pending RAW stall to every vreg-
reader rule so multi-cycle BG (LOG2/EXP2/SIN/COS) and single-cycle
BG followed by a dependent reader both stay correct. The fix is
pessimistic: the reader stalls until the writeback lands even if
its source vreg doesn't collide with vpu_wb_dst. A more precise
curInstr.vregSrc == vpu_wb_dst check would let non-dependent
readers overlap, but both the in-guard form and a fetch-time
precompute form blow past the unfolding budget — likely because
accessing packed-struct fields (curInstr.vregSrc) re-decodes
the SxuInstr representation at every use site. Resolving that
needs either a flat-field decode stage (separate Regs for
vregSrc/vregSrc2/vregDst) or a restructured VRegFile with fewer
writer sites per module.
Iter 4 formalized the dispatch/collect equivalence in Lean:
bg_dispatch_then_collect_matches_sync,
bg_collect_idempotent_when_empty,
bg_dispatch_preserves_reg (13 theorems total, was 10).
- Precise per-vreg RAW scoreboard. Fine-grained alternative
to the coarse !vpu_wb_pending stall. Requires flat scalar Regs
for curInstr.vregSrc/vregSrc2/vregDst populated by do_fetch so
the hazard check
cur_raw_pendingstays cheap to elaborate. Unlocks real dual-issue overlap on VPU-adjacent readers. - Multi-input VPU command queue. A small FIFO of pending VPU dispatches so the main FSM can keep fetching/decoding while back-to-back VPU_BG ops queue up. Useful when pairing a VPU_BG with XLU/MXU dispatch to keep the engines busy. Needs careful elaboration budget: a SizedFIFO + drain rule adds ~2 vrf.read call sites, which may push bsc over.
- Renderer adoption. Tinygrad's FABS / softsign / float-abs tile chains can emit VPU_BG for the final VPU op before STORE once the precise RAW scoreboard lands — without it, the coarse stall burns the dual-issue window.
- Per-engine command queues. Full Tier-1 decouple: SXU fetch stage feeds four per-engine FIFOs (VPU/XLU/MXU/FpReducer) with independent executors. Groundwork is in place (vpu_wb_pending tracks VPU retire, xlu_busy tracks XLU retire); the missing pieces are the FIFOs and a RAW scoreboard keyed on vreg destination index.
Build-time budget: push-5 iter 3 baseline is ~186s CPU for a clean
build/mkTbTinyTPURuntime.bexe. Treat >2x regressions as stop-the-
line; the usual fix is collapsing vrf.write call sites or moving
cross-field equality checks out of rule guards.
- Mixed-precision MXU (bf16/fp16 MACs) — big surgery, unclear ROI without training workloads.
- Hardware collective primitives on the ring NoC — worth the software plumbing first; hardware once we have multi-core workloads.
ATOMICADDon VMEM banks — only matters for scatter-add kernels which aren't on our critical path.- TranscUnit TANH / SIGMOID direct opcodes. Considered but deferred: Padé forms need a divider (not in TranscUnit); degree-5 odd Remez + saturation is feasible but requires careful coefficient fitting. Lower ROI now that exp2/log2 compositions cover the usage.
- Int8 packed VPU arithmetic. [largely landed in push 4] Coverage: ADD/SUB (saturating), MAX/MIN, NEG (saturating), RELU, CMPLT/CMPEQ, MUL_LOW (wrap)/MUL_HIGH (Q1.7), ABS, SIGN — 10 opcodes (55-67). Remaining: (a) byte-wise logical ops reuse VPU_AND/OR/XOR/NOT via the 32-bit path (no new opcode needed); (b) unpack/pack between packed-i8 and full int32 tiles — deferred until a real quantized workload needs it (shape math is tricky); (c) renderer adoption (none yet — no int8 tinygrad workload has landed).
- Nested loops in SXU. [done in push 4 iter 1] LOOP_BEGIN/END now use
a depth-4 stack (
loopCounterStack+loopReturnPcStack+loopTop).LOAD_LOOP_DEPTH(36) exposes the current depth for debug.
Estimate: ~50 total iterations
- Single-tile binary VPU ops
- Single-tile ReLU
- One simple reduction
- Constants and scalar broadcasting
-
WHEREand comparisons - Basic reshape/transpose within one 4x4 tile (reshape via copy renderer; transpose reachable via SXU_DISPATCH_XLU_TRANSPOSE, tinygrad lowering pending)
- VPU-only runtime completion cleanup
Estimate: ~120 total iterations
- Multi-tile elementwise (ADD/MUL/SUB/MAX/MIN/compare/logic + scalar-const all cover numel>16 via tile-loop)
- Multi-tile reductions (scalar via VPU_*_REDUCE_TILE; row/col via SXU_PROGRAM tile iteration)
- Multi-tile movement ops (multi-tile reshape/slice/expand work through the copy renderer's tile loop)
- GEMM epilogues (hardware-backed bias add, ReLU, fused bias+ReLU via SXU_LOAD_MXU_RESULT)
- More shape coverage
- First selected upstream tinygrad test subset passing on
TINYTPU(test_plus_int, test_plus_big)
Estimate: ~250 total iterations
- Most movement ops
- Most integer elementwise ops
- Add/max/mul reductions (SUM/MAX/MIN/PROD scalar, row, col via VPU_*_REDUCE{,_COL,_TILE})
- More dtype handling
- Multi-kernel scheduling
- Structural lowering pass
Estimate: 400+ iterations
- Broad tinyspec semantic coverage
- Well-tested runtime and lowering failures
- Robust memory planning
- Multi-device decisions documented
- Advanced codegen ops either implemented or explicitly out of scope
Estimate: 500-800 iterations
- Full core tinyspec behavior
- Broad upstream tinygrad test coverage
- Hardware/runtime-backed execution where appropriate
- Clear software fallback policy where hardware is not appropriate
- Stable performance/profiling story
Current state: 889 backend tests pass + 95 BSV unit tests. ~120 new Python tests added this push.
Second-push additional landings (iters 24-31):
- clip / hardtanh renderer — nested WHERE+CMPLT with 2 float CONSTs now emits FMIN+FMAX per tile.
- single-bound clamp renderer — clamp(min=c) / clamp(max=c) with c != 0 emit FMAX / FMIN per tile. clamp(min=0) still routes through RELU (semantically equivalent).
- RELU pattern guard tightening — requires CMPLT's CONST operand to be 0.0, so clamp(max=c) for c != 0 no longer silently false-matches as RELU.
- leaky_relu renderer — max(alpha*x, x) for 0 < alpha < 1 via FMUL + FMAX.
- softplus min-const false-match guard — the float-min-const path now rejects kernels carrying transcendentals / reciprocals, so softplus stops silently rendering as minimum(x, 0).
- x4 regression guard** — self-square tightens to require LOAD tails.
- Broad float / int32 correctness sweeps — 17 float ops + 16 int ops compared against numpy reference to catch future regressions.
Current state: 883 backend tests pass + 90 BSV unit tests (VPU 48 + FpReducer 7 + PSUMBank 9 + SXU 6 + SxuPSUM 2 + CtrlPSUM 1 + CtrlOS 1 + CtrlDB 2 + CtrlDBDMA 2 + WSRAMDB 3 + ASRAMDB 3 + WeightDMA 3 + ActivationDMA 3 + TranscUnit 5).
Landings in this second push (iters 1-23):
- Item #8 (OS MXU): PE accumulator-hold DONE. startOS skips
array.clearAll on drain; start/startPsum now pre-clear so WS always
runs fresh. clearArray() method + SXU_MXU_CLEAR opcode (24) + TASM
MXU_CLEARmnemonic. End-to-end sim tests cover OS back-to-back accumulate + MXU_CLEAR reset. - Item #4 (Dual-issue XLU): DONE. XLU dispatch advances pc
immediately; do_xlu_collect_bg writes the result one cycle later
(XLU.result has 1-cycle latency). Structural-hazard guard
!xlu_busyon every XLU dispatch rule. Chained-XLU sim test covers the path. - Item #5 (DMA stubs): DONE stubs + integration. WeightDMA + ActivationDMA synthesize deterministic tile patterns into inactive DB bank; standalone TBs + ping-pong TbCtrlDBDMA where MXU dispatches tile A on active while DMA preloads tile B on inactive. TensorCore default-wiring to DB SRAMs deferred (breaks existing preload semantics; needs an API revamp).
- Composite activations + renderers:
- Self-cube
x**3renderer (MUL(x, MUL(x, x)) pattern). - Self-square renderer tightened — no longer false-matches
x**4. - Swish / silu renderer (x * sigmoid(x)).
- Softsign renderer (x / (1 + |x|)).
- Self-cube
- Reducer silent-bug round-up: Five pattern guards added to
scalar / col / row reducers so the following compound kernels route
to UNSUPPORTED instead of silently returning a wrong answer:
- hardtanh / clip (relu false-match via multi-CMPLT).
- softsign (abs false-match via abs-in-MUL-RECIPROCAL tree).
sum(abs(x))(pre-reduce WHERE/CMPLT).reciprocal().sum()/exp().sum()(pre-reduce transcendental).sum(x*x)/sum(-x)/max(-x)(pre-reduce data-path MUL).log(x+k),exp(x+k),sin(x+k)scaled renderers now require LOAD-terminated non-const factors, refusing compound shifts.
- Correctness wins:
- Float
mean()now returns correct results (reducer post-op now remapped to FADD/FMUL for float reductions). - Float
mean(axis=0)/mean(axis=1)on fused kernels (col/row reducers now detect post-reduction FMUL(const)). - Float softsign end-to-end correct.
- Float
Current state: 1006 Python tests + 87 BSV unit tests (VPU 48 + FpReducer 7 + PSUMBank 9 + SXU 6 + SxuPSUM 2 + CtrlPSUM 1 + CtrlOS 1 + CtrlDB 2 + WSRAMDB 3 + ASRAMDB 3 + TranscUnit 5).
Landings in this push:
- Item #6 (Transcendentals): DONE with Remez + range reduction for EXP2/LOG2/SIN/COS. Tensor.exp, tanh, sigmoid, cos, log, sqrt, rsqrt all accurate across wide ranges.
- Item #8 (OS MXU): DISPATCH PATH LIVE via Controller.startOS() + SXU DISPATCH_MXU_OS opcode. PE accumulator-hold FSM pending.
- Item #5 (DB SRAM): CONTROLLER INTEGRATION.
.plainsub-interface lets Controller consume DB SRAMs; preload-parallel pattern proven in TbCtrlDB. DMA stub + TensorCore default-wiring pending. - Item #4 (Dual-issue): SCOREBOARD LIVE.
xlu_busy/xlu_dstset on dispatch, cleared on collect. Parallel rule + RAW-stall pending. - Renderer additions: Tensor.cos, Tensor.log, Tensor.rsqrt, Tensor.square, multi-tile wide-input tanh/exp/sigmoid/cos/sin tests.
Highest-leverage follow-ups (ranked):
- Finish Item #8 — PE accumulator-hold mode in SystolicArray + Controller FSM variant swapping operand roles so DISPATCH_MXU_OS delivers a genuinely distinct dataflow. ~6-8 iters.
- Finish Item #4 — parallel XLU issue rule + RAW-hazard stall consuming the existing scoreboard. ~4-6 iters.
- Finish Item #5 — DMA stub issuing background writes, plus TensorCore wiring to use DB SRAMs by default. ~4-6 iters.
- Engine-to-engine forwarding in renderers — emit SXU_LOAD_VPU_RESULT / LOAD_XLU_RESULT to elide VRegFile round-trips in chained kernels. ~3 iters.
- Composite activations: softplus (log(1+exp(x))), gelu (uses tanh), x**3 (MUL(x, MUL(x, x))). ~2-3 iters each.
- Narrow-dtype int8 VPU — big scope. Blocks quantized INT8 inference kernels outside the MXU. ~8-10 iters.
Current state: 379 backend tests + 8 runtime bundle tests. Row/col/scalar reductions (SUM/MAX/MIN/PROD) fully hardware-backed. Movement ops: reshape, contiguous slice/shrink, scalar and unrolled row-expand working. SXU_DISPATCH_XLU_TRANSPOSE opcode exists in hardware but tinygrad permute/transpose lowering not yet wired.
Highest-value next work:
- Wire tinygrad
permute(1,0)to SXU_DISPATCH_XLU_TRANSPOSE. Detect the 2D index-swap pattern in the UOp graph and emit LOAD/XLU_TRANSPOSE/STORE for single-tile cases; extend to multi-tile via a tile-of-tiles pass. - RANGE-loop variant of row-broadcast expand (shape (N, M) with M<_COLS and outer RANGE over rows) — current renderer handles only unrolled.
- Hardware epilogue for multi-K-tile GEMM (currently still a numpy fallback).
- Replace remaining hand-written structural recognizer duplication with reusable TinyTPU descriptor builders.
- Dtype expansion beyond int32/bool: int8/uint8/int16 elementwise policy.
- Movement gaps: PAD (requires CMPLT bounds check), FLIP (LOAD index = const - STORE index), CAT (WHERE-based select between two buffers), true permute (non-transpose axis reorder).
- Multi-output tile scheduling cleanup and intermediate VMEM allocation.
- Fused kernel detection improvements for 2-op chains tinygrad emits (slice-then-sum currently fuses to wrong-offset sum).
leaky_relu(alpha): tinygrad emits float mul for the negative slope. SW path: integer approximation via SHR + WHERE select.tanh,log_softmax, float32 GEMM: require float datapath (major arch change). SW fallback via host numpy or replace with int/argmax where semantically acceptable for inference-only models.backward/ autograd: training is not supported; export int8 weights and run inference only.