MLIR · LLVM · hardware-aware compilation

nano-dsp-mlir

A small MLIR compiler for a tiny DSP dialect. It tiles and vectorizes each op for a specific machine, using that machine's cache size and SIMD registers, then proves the result didn't change by a single bit.

Pipeline

One matmul, from source to machine code


  
Hardware-aware

Same program, different machine, different code

Tile sizes come from a compile-time machine model. The register tile fills the SIMD register file. The cache tile grows until its working set reaches the cache budget.

Generated schedule (Transform dialect)

      
Inner loop, machine code

      
Correctness

Optimized output is bit-identical

Each register tile adds one product at a time, as a separate multiply and add, so every output is summed in the same order as the plain loop nest. The red run is the control. It shows the check catches a schedule that reorders the sum.

GPU

The same numbers on a GPU

The Mojo matmul and conv2d run on an Apple M2 and an NVIDIA T4 from the same source, and every output is compared with the CPU kernel as raw 32-bit patterns before it is timed.

Saved results of pixi run bench-gpu, not run when this page was built (it is built without a GPU). Median times; throughput from the best sample. On the T4 the blocked matmul reaches 49–58% of cuBLAS from 512³ up (cuBLAS is checked within the reordering bound, not bit for bit).

Roadmap