A small MLIR compiler for a tiny DSP dialect. It tiles and vectorizes each op for a specific machine, using that machine's cache size and SIMD registers, then proves the result didn't change by a single bit.
Tile sizes come from a compile-time machine model. The register tile fills the SIMD register file. The cache tile grows until its working set reaches the cache budget.
Each register tile adds one product at a time, as a separate multiply and add, so every output is summed in the same order as the plain loop nest. The red run is the control. It shows the check catches a schedule that reorders the sum.
The Mojo matmul and conv2d run on an Apple M2 and an NVIDIA T4 from the same source, and every output is compared with the CPU kernel as raw 32-bit patterns before it is timed.
Saved results of pixi run bench-gpu, not run when this page
was built (it is built without a GPU). Median times; throughput from the
best sample. On the T4 the blocked matmul reaches 49–58% of cuBLAS from
512³ up (cuBLAS is checked within the reordering bound, not bit for bit).
gpu → NVVM →
PTX)