Expand TinyTPU implementation coverage autonomously.
The default mission is to make the system handle more real workloads through the smallest defensible changes across:
- BSV hardware in
src/ - BSV tests in
test/ - tinygrad TinyTPU backend in
tinygrad/tinygrad/runtime/ops_tinytpu.py - supporting scripts and tests in
scripts/andtests/
The agent should prefer changes that improve supported functionality, correctness, and debuggability without adding unnecessary complexity.
Before starting a new implementation push, read the repo context:
README.md- The directly relevant design/spec doc under
doc/ - The source and test files in the area being changed
tinygrad/tinygrad/runtime/ops_tinytpu.pywhen the task touches lowering or co-sim
Do not make blind edits. Build context first, then choose the smallest useful slice.
Loop continuously until interrupted:
- Check git state and confirm the current branch/commit.
- Choose one bounded implementation gap.
- Reproduce the gap with the smallest relevant test.
- Implement the minimum coherent fix.
- Run the narrowest useful verification first, then broader verification if warranted.
- Keep the change only if it advances functionality or simplifies the code without regression.
- Update
TODO.mdwhen the iteration changes tinyspec coverage, milestone progress, known gaps, or recommended next work. - Update
results.tsvwith one tab-separated progress row for the completed iteration or iteration batch. - Commit the improvement.
- Repeat from the new head.
The loop is intended to keep expanding real capability, not just churn code.
Preferred order:
- A failing or missing end-to-end TinyTPU capability
- A lowering/runtime gap in the tinygrad
TINYTPUbackend - A missing hardware instruction/data-path needed by software
- Missing tests for behavior the repo claims to support
- Simplifications that preserve behavior and reduce code or concepts
Good task shapes:
- Add one missing SXU/VPU/XLU/backend capability end to end
- Turn a current
NotImplementedErrorpath into working behavior - Add a missing testbench case and implement the minimal hardware/software support
- Fix a real mismatch between the software contract and the BSV implementation
Avoid:
- Large speculative rewrites without a failing test or concrete target
- Changing multiple subsystems at once unless the boundary requires it
- Complexity that is not justified by clear capability gain
For model bring-up and backend expansion, use this rule:
- Start from the failing ONNX model or tinygrad workload.
- Compile through tinygrad targeting
TINYTPU. - Inspect where execution fails in the TinyTPU renderer/runtime path.
- First assume it is a software/lowering problem.
- Escalate to a TinyTPU hardware issue only when the required behavior cannot be expressed with the current instruction set or data paths.
Expected output for each investigated workload:
- What tinygrad produced
- What the TinyTPU backend could and could not lower
- The exact missing instruction sequence, lowering rule, or hardware capability
- The chosen fix location: tinygrad software or TinyTPU hardware
After any successful model or workload run, explicitly check ISA support sufficiency:
- Did the workload run end to end on
TINYTPUwithout host fallback? - Did lowering use direct primitives or short regular SXU programs instead of large analyzer special-cases?
- Did any unsupported path,
UNSUPPORTEDdescriptor, or simulator-side software convention appear? - Did the generated bundle count and instruction mix look structurally reasonable for the workload?
If the answer to any of these is "no", record the exact missing primitive, lowering gap, or hardware contract in TODO.md. Do not treat "model ran" as equivalent to "ISA support is sufficient".
- Prefer test-driven expansion: add or tighten a failing test before the fix when practical.
- Keep edits local to the subsystem you are advancing.
- Preserve existing architectural boundaries unless there is a concrete reason to change them.
- Favor simple instruction sequences and explicit data movement over clever abstractions.
- If a tinygrad-side workaround can express the behavior cleanly with the current ISA, do that before inventing new hardware.
- If hardware is required, make the missing contract explicit in code and tests.
The target architecture for ops_tinytpu.py is a UOp-walking renderer that
emits SXU instructions directly from the tinygrad UOp graph — one SXU/VPU
instruction per UOp, like how CStyleLanguage emits one line of C per UOp.
When a UOp has no corresponding SXU/VPU instruction, add a hardware primitive rather than writing complex UOp-graph pattern matching in software. Prefer a new VPU opcode or SXU instruction over a multi-hundred- line graph walker that reverse-engineers tinygrad's decomposition.
Concrete priority order:
- Add a hardware primitive (new VPU opcode, new SXU instruction) that maps 1:1 to the tinygrad UOp. This is always preferred.
- Emit a short SXU microprogram (2–4 existing instructions) when the pattern is simple and stable (e.g. WHERE via COPY+SELECT).
- Use pattern matching as a last resort only when the tinygrad decomposition is too complex for a single hardware primitive and no short microprogram exists. Document why and plan the hardware fix.
Examples of hardware-first decisions:
VPU_SELECTreplaced a 4-instruction MUL/SUB/MUL/ADD WHERE sequenceVPU_NOTreplaced XOR-with-all-ones constant tile loadingVPU_MINis preferred over detecting XOR+MAX decomposition in software- Remaining decomposition-heavy ops should be resolved by adding direct hardware opcodes where appropriate, not by reintroducing whole-kernel graph counting.
Propose ISA additions in TODO.md and get user approval before implementing
(per the microarchitecture rule below). The goal is a 1:1 UOp→instruction
mapping that avoids whole-kernel pattern-matching layers.
When a workload hits an unsupported op, dtype, or shape:
- Record it in
TODO.mdunder the appropriate coverage area with a concrete description. - Decide whether the fix belongs in software (tinygrad backend lowering) or hardware (BSV):
- Software first: if the behavior can be expressed with existing SXU/VPU/MXU/XLU instructions, lower it in
ops_tinytpu.py. Examples: new elementwise op via existing VPU opcodes, host fallback for unsupported dtypes, shape remapping. - Hardware needed: if no existing instruction sequence can express the behavior, or if a software workaround would be unreasonably slow. Examples: new VPU opcode, new SXU instruction, wider data paths. In this case, note the required BSV change in
TODO.mdand ask the user before implementing — do not expand the microarchitecture (new opcodes, new functional units, wider data paths, new SRAMs) without explicit approval.
- Software first: if the behavior can be expressed with existing SXU/VPU/MXU/XLU instructions, lower it in
- If the unsupported feature blocks a real model (not just a synthetic test), prioritize it higher.
- Do not silently skip unsupported features — always emit a clear diagnostic via the
UNSUPPORTEDdescriptor path so the gap is visible.
Use the narrowest command that proves the change:
make test-<unit>for unit-level BSV workmake testwhen cross-cutting changes justify the costpython3 scripts/test_cosim.pyfor end-to-end tinygrad co-simpytest tests/...for Python-side tooling or profiler work
When fixing a bug:
- Reproduce it with a targeted test.
- Make the test pass.
- Run adjacent tests that could plausibly regress.
If a change cannot be verified locally, state exactly what remains unverified and why.
Keep a change if it does at least one of these:
- Enables a new supported behavior
- Fixes an incorrect result
- Removes code while preserving behavior
- Improves debuggability or diagnostics for a real failure mode
Discard or rework a change if it:
- Adds complexity with negligible functional gain
- Leaves behavior ambiguous or untested
- Fixes one path by hard-coding around the design
- Regresses a nearby unit or end-to-end flow
Commit only coherent advances. Each commit should describe one implementation step, for example:
xlu: add transpose path for runtime bundlestinytpu: lower elementwise add through vpu sequencetensorcore: fix mxu completion handshake
Do not mix unrelated cleanups into an implementation commit.
When updating TODO.md, keep the edit scoped to the completed iteration:
- Mark newly supported behavior as complete.
- Add newly discovered gaps or limitations.
- Adjust recommended next iterations if the priority changed.
- Do not rewrite unrelated estimates or checklist sections unless the iteration made them obsolete.
When updating results.tsv, append one row per completed iteration or batch. Keep the file tab-separated with this schema:
date commit iterations scope supported_delta tests_passed todo_delta remaining_gap runtime_build_s
Rules for results.tsv:
- Use ISO dates.
- Use the short commit hash after the commit is created.
- Keep fields concise and avoid tabs inside field values.
- Set
iterationsto the number of loop iterations represented by the row. - Use
supported_deltafor newly supported behavior ornonefor docs-only/tooling-only work. - Use
tests_passedfor the highest-signal verification command and result. - Use
todo_deltato summarize TODO progress changed by the iteration. - Use
remaining_gapfor the next concrete blocker surfaced by the work. - Use
runtime_build_sfor a clean-rebuild wall-clock timing ofbuild/mkTbTinyTPURuntime.bexe(seconds, from thetotalline oftime make build/mkTbTinyTPURuntime.bexeafter removing the relevant.bo/.bexeartifacts). If an iter didn't touch BSV or didn't need the runtime sim, reuse the last iter's number. If it jumps >2× vs. the prior iter, treat it as a regression and investigate before landing (usually: collapsing extra vrf.read / vrf.write call sites, or moving cross-field struct-equality checks out of rule guards and into fetch-time cached Bools).
For each completed loop iteration, leave the repo in a state where a human can see:
- What gap was targeted
- What test or workload reproduced it
- What was changed
- What verification passed
- What
TODO.mdprogress was updated, if applicable - What
results.tsvprogress row was appended - What limitations remain, if any