MLIR · developer tooling

Compiler Tooling Lab

Three pinned MLIR projects, read as one pipeline: source, diagnostics, MLIR, pass inspection, profiling.

This page follows three separate compiler projects through the stages of a developer-tooling pipeline. Each project is pinned to one commit. Every output quoted below was captured from the project's own tools at that commit, or is a file shown verbatim from its repository. Where something could not be captured, the page says so instead of filling the gap.

  1. Source. JSON Schema constraints as schema-dialect IR. json-schema-mlir
  2. Diagnostics. Verifier errors with file, line and column. json-schema-mlir
  3. MLIR. Custom dialects lowered to the LLVM dialect. json-schema-mlir, nano-dsp-mlir
  4. Pass inspection. IR snapshots and structural diffs, per pass. VizMLIR, nano-dsp-mlir
  5. Profiling. Compile time per pass, plus runtime numbers where they exist. nano-dsp-mlir

1. json-schema-mlir

JSON Schema constraints as an MLIR dialect, optimized as IR and lowered to native code.

Source · Diagnostics · MLIR
v0.1.0-12-g6f4be81 · Source at 6f4be81 · README.md · Architecture · Run it locally · Tests and benchmarks

json-schema-mlir represents JSON Schema validation rules as operations in a schema dialect. Constraints are optimized as IR (subsumption, fusion, removing redundant checks) before being lowered to arith/scf/math and the LLVM dialect.

The same constraint lattice also exists as a Mojo library (mojo/schema/), with the canonicalizer's subsumption and meet rules and the lowered validation semantics. Its tests check that merging two constraints never changes which values are accepted.

At the pinned commit, the entry point is schema-opt on .mlir input. The JSON-to-IR front end shown in the upstream pipeline diagram is not in the tree yet.

Follow one real test function from the dialect, through canonicalization and lowering, to the LLVM dialect, plus one verifier diagnostic.

Start from a real test input

Three validate_string checks on the same value, combined with arith.andi. This function comes verbatim from the repository's canonicalization tests.

Open the captured output @fuse_constraint_tree (from the canonicalization tests)

Diagnostics point at the source

A contradictory constraint (max_length below min_length) is rejected by the op verifier. The error carries file, line and column. This input file was written for the lab; the diagnostic is schema-opt's real output.

Open the captured output Contradictory constraint (input written for this lab) · Verifier diagnostic

Every pass in the pipeline

--schema-to-std-pipeline runs the passes below. Rows marked unchanged left the IR exactly as they found it.

Open the captured output pass table

Canonicalization fuses the constraint tree

The three checks become one validate_string carrying the meet of the constraints: min_length 5 subsumes min_length 2. Both andi ops disappear.

Open the captured output Input as parsed and printed by schema-opt · IR after SchemaCanonicalizerPass (schema-canonicalize)

Lowering guards on the runtime type tag

The validator becomes an scf.if on the JSON kind, so a length check never runs on a non-string. String length and regex matching go through the runtime ABI.

Open the captured output IR after SymbolDCEPass (symbol-dce)

Down to the LLVM dialect

--schema-to-llvm-pipeline continues through control flow and LLVM conversion.

Open the captured output LLVM dialect output

The same lattice in Mojo

The Mojo library repeats each canonicalization test case, then checks a grid of constraint pairs and values, including NaN and infinity: whenever two constraints merge, the merged one accepts exactly what both accepted.

Open the captured output Mojo constraint-lattice tests

2. VizMLIR

Browser-only MLIR graph viewer and structural diff, with a Rust parser compiled to WebAssembly.

Pass inspection
v0.3.0-14-gb4f2c91 · Source at b4f2c91 · README.md · Architecture · Run it locally · Tests and benchmarks

VizMLIR parses MLIR in the browser with a Rust parser compiled to WebAssembly, draws the operation graph on a canvas, and diffs two IR snapshots with SSA names normalized. No backend is involved.

In this lab it is the pass-inspection step: the IR captured from json-schema-mlir and nano-dsp-mlir is run through VizMLIR's own WASM parser and diff module at the pinned commit.

VizMLIR has a live app: joepothiboot.github.io/vizmlir/. It deploys from the project's main branch, so it can be newer than the pinned b4f2c91 described here.

Diff real compiler passes from the other projects with VizMLIR's own WASM parser, then open the app and try it yourself.

Diff json-schema-mlir's canonicalization

VizMLIR matches operations by kind and SSA-normalized label. The fused validate_string keeps its op name, so it matches. The two absorbed checks and both andi ops are reported as removed.

Open the captured output json-schema-mlir canonicalization, diffed by VizMLIR

A real finding: function arguments

In the graph summaries above, VizMLIR reports each use of %arg0 as an undefined SSA value. At this commit the parser does not appear to register function block arguments as definitions. This is captured output, shown as-is, and a candidate upstream fix.

Diff the lowering to standard dialects

The larger diff after --lower-schema-to-std shows the runtime ABI declarations and the scf.if guard being introduced.

Open the captured output json-schema-mlir lowering to std, diffed by VizMLIR

Diff nano-dsp-mlir's dsp → linalg conversion

The same diff applied to the other compiler: dsp ops become linalg.generic with affine maps.

Open the captured output nano-dsp-mlir dsp → linalg, diffed by VizMLIR

3. nano-dsp-mlir

A small image/math DSL dialect lowered to linalg and LLVM, with differential execution tests.

MLIR · Pass inspection · Profiling
v0.1.0-6-gccb09e1 · Source at ccb09e1 · README.md · Architecture · Run it locally · Tests and benchmarks

nano-dsp-mlir defines a dsp dialect (add, relu, matmul, conv2d) with shape verifiers, and lowers it to linalg.generic. From there the tests lower through stock MLIR passes to LLVM and execute with mlir-runner to check numeric results.

Each dsp op is also implemented as a SIMD Mojo kernel (mojo/nanodsp/) and as a plain C++ loop nest (reference/). All three implementations are tested against the same expected values, and the Mojo kernels are benchmarked. The upstream README also describes pieces not in the tree yet: transform-dialect schedules, -nanodsp-optimize, the Python DSL and sweep scripts.

Follow relu(conv2d(image) + bias) from the dsp dialect through 19 passes to the LLVM dialect, execute it, and see where compile time goes.

The program

An integration test: a 3×3 box blur over a 4×4 image, plus a bias, through relu. The CHECK lines state the expected output.

Open the captured output relu(conv2d(image) + bias) integration test

Verifier diagnostics

The invalid-op tests, run without -verify-diagnostics, so the raw errors and their locations are visible.

Open the captured output Verifier diagnostics for the invalid-op tests

dsp → linalg

The only pass this project owns. conv2d, add and relu become linalg.generic ops with explicit indexing maps.

Open the captured output IR after ConvertDSPToLinalg (convert-dsp-to-linalg)

Every pass to the LLVM dialect

The stock lowering chain the tests use. Unchanged rows are passes that found nothing to do.

Open the captured output pass table

Execute and check

The lowered module run with mlir-runner. The printed memref is compared with the test's CHECK lines.

Open the captured output Executed with mlir-runner

The same ops in Mojo and C++

The Mojo kernels and the C++ reference are checked against the same expected values as the MLIR tests above. The Mojo tests also compare against a naive loop nest on odd sizes, so the code after each SIMD chunk runs too.

Open the captured output Mojo kernel tests · C++ reference tests

Profiling

Compile-time wall clock per pass, from one run on the capture host. It shows relative cost, not a benchmark. The kernel numbers are the untiled Mojo kernels on one core, the baseline the planned tiling stage has to beat.

Open the captured output Compile-time pass timing · Mojo kernel throughput (untiled, one core, best of 3-5 runs)

Appendix: captured outputs and notes

The material linked from the article above, in order.

json-schema-mlir

Architecture

  • schema dialect for JSON Schema string and number constraints
  • Constraint-lattice canonicalization (subsumption, fusion)
  • Lowering to arith/scf/math/func and the LLVM dialect
  • Verifier diagnostics with MLIR source locations
  • Mojo library with the same constraint lattice, tested for soundness of meet
  1. schema dialect

    ODS definitions for validate_string, validate_number and struct, with verifiers.

    include/Schema/SchemaOps.td

  2. --schema-canonicalize

    Constraint-lattice subsumption and conjunction fusion.

    lib/Schema/SchemaCanonicalizerPass.cpp

  3. --lower-schema-to-std

    Type-guarded lowering to arith/scf/math plus runtime ABI calls.

    lib/Schema/LowerToStandard.cpp

  4. --schema-to-llvm-pipeline

    Registers the pipelines that continue to the LLVM dialect.

    tools/schema-opt/schema-opt.cpp

  5. Mojo lattice library

    StringConstraints and NumberConstraints with subsumes, meet and validate, following the same rules as the two passes above.

    mojo/schema/lattice.mojo

json-schema-mlir

Run it locally

git clone https://github.com/joepothiboot/json-schema-mlir && cd json-schema-mlir
git checkout 6f4be813566745002c9bebcb8b4b929bade347cd
MLIR_INSTALL="$(brew --prefix llvm)" ./build.sh   # configures, builds schema-opt, runs check-schema
./build/bin/schema-opt test/Dialect/Schema/schema-canonicalize.mlir --schema-canonicalize --split-input-file
pixi run test-mojo   # Mojo lattice tests; pixi installs the pinned Mojo

Needs LLVM/MLIR with FileCheck and lit (captured with the toolchain listed below). Clone into a path without spaces, because lit's %s substitution breaks on them.

json-schema-mlir

Tests and benchmarks

lit regression suite (check-schema)

3 passed, 0 failed, 0 errors, 0 skipped/deselected (lit)

Test log
PASS: JSON-SCHEMA-MLIR :: Dialect/Schema/schema-canonicalize.mlir (1 of 3)
PASS: JSON-SCHEMA-MLIR :: Dialect/Schema/ops.mlir (2 of 3)
PASS: JSON-SCHEMA-MLIR :: Lowering/lower-to-std.mlir (3 of 3)

Total Discovered Tests: 3
  Passed: 3 (100.00%)

Captured Output of lit -v build/test at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Mojo constraint-lattice tests

12 passed, 0 failed, 0 errors, 0 skipped/deselected (mojo)

Test log
schema: 12 tests passed
✨ Pixi task (test-mojo): mojo run -I mojo mojo/tests/test_lattice.mojo

Captured Output of pixi run test-mojo at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Compile-time pass timing

wall time per pass, one run; total 0.0012 s
PasssShare
Parser0.000218.1%
CanonicalizerPass0.00003.4%
'func.func' Pipeline0.00001.8%
SchemaCanonicalizerPass0.00001.7%
LowerSchemaToStandardPass0.000215.4%
ReconcileUnrealizedCastsPass0.00014.8%
CanonicalizerPass0.00017.0%
CSEPass0.00000.9%
(A) DominanceInfo0.00000.0%
SymbolDCEPass0.00014.5%
Output0.00016.1%
Rest0.000537.2%

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-std-pipeline --mlir-timing -o /dev/null at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

json-schema-mlir

Start from a real test input

@fuse_constraint_tree (from the canonicalization tests)

Stage: schema

From test/Dialect/Schema/schema-canonicalize.mlir:29

func.func @fuse_constraint_tree(%doc: !schema.value) -> i1 {
  %a = schema.validate_string %doc { min_length = 5 : i64 } : !schema.value
  %b = schema.validate_string %doc { max_length = 32 : i64 } : !schema.value
  %c = schema.validate_string %doc { min_length = 2 : i64, pattern = "^[a-z]+$" } : !schema.value
  %ab  = arith.andi %a, %b : i1
  %abc = arith.andi %ab, %c : i1
  return %abc : i1
}

Static file test/Dialect/Schema/schema-canonicalize.mlir#L29-L36 at json-schema-mlir@6f4be81, shown verbatim.

json-schema-mlir

Diagnostics point at the source

Contradictory constraint (input written for this lab)

From inputs/schema-mlir/contradictory-bounds.mlir:1

// Written for compiler-tooling-lab: a deliberately contradictory constraint,
// used to capture a real schema-opt verifier diagnostic.
func.func @bad(%doc: !schema.value) -> i1 {
  %v = schema.validate_string %doc { min_length = 8 : i64, max_length = 3 : i64 } : !schema.value
  return %v : i1
}

Static file compiler-tooling-lab/inputs/schema-mlir/contradictory-bounds.mlir, shown verbatim.

Verifier diagnostic

  • error inputs/schema-mlir/contradictory-bounds.mlir:4:8 'schema.validate_string' op 'max_length' (3) must be >= 'min_length' (8)
  • note inputs/schema-mlir/contradictory-bounds.mlir:4:8 see current operation: %0 = "schema.validate_string"(%arg0) <{max_length = 3 : i64, min_length = 8 : i64}> : (!schema.value) -> i1

schema-opt exit code 1

Captured Output of schema-opt 'lab:inputs/schema-mlir/contradictory-bounds.mlir' at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

json-schema-mlir

Every pass in the pipeline

Pass events, in execution order
#PassEffectIR lines after
1CanonicalizerPass (canonicalize)unchanged10
2SchemaCanonicalizerPass (schema-canonicalize)changed6
3LowerSchemaToStandardPass (lower-schema-to-std)changed29
4ReconcileUnrealizedCastsPass (reconcile-unrealized-casts)unchanged29
5CanonicalizerPass (canonicalize)changed29
6CSEPass (cse)unchanged29
7SymbolDCEPass (symbol-dce)changed26

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-std-pipeline --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

json-schema-mlir

Canonicalization fuses the constraint tree

Input as parsed and printed by schema-opt

Stage: schema

module {
  func.func @fuse_constraint_tree(%arg0: !schema.value) -> i1 {
    %0 = schema.validate_string %arg0 {min_length = 5 : i64} : !schema.value
    %1 = schema.validate_string %arg0 {max_length = 32 : i64} : !schema.value
    %2 = schema.validate_string %arg0 {min_length = 2 : i64, pattern = "^[a-z]+$"} : !schema.value
    %3 = arith.andi %0, %1 : i1
    %4 = arith.andi %3, %2 : i1
    return %4 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

IR after SchemaCanonicalizerPass (schema-canonicalize)

Stage: schema

module {
  func.func @fuse_constraint_tree(%arg0: !schema.value) -> i1 {
    %0 = schema.validate_string %arg0 {max_length = 32 : i64, min_length = 5 : i64, pattern = "^[a-z]+$"} : !schema.value
    return %0 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-std-pipeline --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

json-schema-mlir

Lowering guards on the runtime type tag

IR after SymbolDCEPass (symbol-dce)

Stage: llvm

module attributes {schema.string_pool = ["^[a-z]+$"]} {
  func.func private @__schema_rt_kind(i64) -> i32 attributes {llvm.readnone}
  func.func private @__schema_rt_str_len(i64) -> i64 attributes {llvm.readnone}
  func.func private @__schema_rt_str_matches(i64, i64) -> i1 attributes {llvm.readnone}
  func.func @fuse_constraint_tree(%arg0: i64) -> i1 {
    %false = arith.constant false
    %c0_i64 = arith.constant 0 : i64
    %c32_i64 = arith.constant 32 : i64
    %c5_i64 = arith.constant 5 : i64
    %c3_i32 = arith.constant 3 : i32
    %0 = call @__schema_rt_kind(%arg0) : (i64) -> i32
    %1 = arith.cmpi eq, %0, %c3_i32 : i32
    %2 = scf.if %1 -> (i1) {
      %3 = func.call @__schema_rt_str_len(%arg0) : (i64) -> i64
      %4 = arith.cmpi sge, %3, %c5_i64 : i64
      %5 = arith.cmpi sle, %3, %c32_i64 : i64
      %6 = arith.andi %4, %5 : i1
      %7 = func.call @__schema_rt_str_matches(%arg0, %c0_i64) : (i64, i64) -> i1
      %8 = arith.andi %6, %7 : i1
      scf.yield %8 : i1
    } else {
      scf.yield %false : i1
    }
    return %2 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-std-pipeline --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

json-schema-mlir

Down to the LLVM dialect

LLVM dialect output

Stage: llvm

module attributes {schema.string_pool = ["^[a-z]+$"]} {
  llvm.func @__schema_rt_kind(i64) -> i32 attributes {llvm.readnone, memory_effects = #llvm.memory_effects<other = none, argMem = none, inaccessibleMem = none, errnoMem = none, targetMem0 = none, targetMem1 = none>, sym_visibility = "private"}
  llvm.func @__schema_rt_str_len(i64) -> i64 attributes {llvm.readnone, memory_effects = #llvm.memory_effects<other = none, argMem = none, inaccessibleMem = none, errnoMem = none, targetMem0 = none, targetMem1 = none>, sym_visibility = "private"}
  llvm.func @__schema_rt_str_matches(i64, i64) -> i1 attributes {llvm.readnone, memory_effects = #llvm.memory_effects<other = none, argMem = none, inaccessibleMem = none, errnoMem = none, targetMem0 = none, targetMem1 = none>, sym_visibility = "private"}
  llvm.func @fuse_constraint_tree(%arg0: i64) -> i1 {
    %0 = llvm.mlir.constant(false) : i1
    %1 = llvm.mlir.constant(0 : i64) : i64
    %2 = llvm.mlir.constant(32 : i64) : i64
    %3 = llvm.mlir.constant(5 : i64) : i64
    %4 = llvm.mlir.constant(3 : i32) : i32
    %5 = llvm.call @__schema_rt_kind(%arg0) : (i64) -> i32
    %6 = llvm.icmp "eq" %5, %4 : i32
    llvm.cond_br %6, ^bb1, ^bb2
  ^bb1:  // pred: ^bb0
    %7 = llvm.call @__schema_rt_str_len(%arg0) : (i64) -> i64
    %8 = llvm.icmp "sge" %7, %3 : i64
    %9 = llvm.icmp "sle" %7, %2 : i64
    %10 = llvm.and %8, %9 : i1
    %11 = llvm.call @__schema_rt_str_matches(%arg0, %1) : (i64, i64) -> i1
    %12 = llvm.and %10, %11 : i1
    llvm.br ^bb3(%12 : i1)
  ^bb2:  // pred: ^bb0
    llvm.br ^bb3(%0 : i1)
  ^bb3(%13: i1):  // 2 preds: ^bb1, ^bb2
    llvm.br ^bb4
  ^bb4:  // pred: ^bb3
    llvm.return %13 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-llvm-pipeline at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

json-schema-mlir

The same lattice in Mojo

Mojo constraint-lattice tests

12 passed, 0 failed, 0 errors, 0 skipped/deselected (mojo)

Test log
schema: 12 tests passed
✨ Pixi task (test-mojo): mojo run -I mojo mojo/tests/test_lattice.mojo

Captured Output of pixi run test-mojo at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

VizMLIR

Architecture

  • Browser-only MLIR parser compiled from Rust to WebAssembly
  • Operation graph rendering on canvas
  • SSA-renumbering-aware structural diff between two IR snapshots
  • Unit tests for the trace parser, diff and WASM bridge (vitest)
  1. Lexer and parser (Rust)

    Tokenizes and parses MLIR into an arena-allocated AST with interned strings.

    wasm/src/parser/lexer.rs

  2. WASM ABI

    A fixed shared-memory header (ABI v3) exposing nodes, edges, layout and diagnostics.

    wasm/src/abi.rs

  3. JS bridge

    Reads the header through typed-array views; no copying or JSON.

    src/wasm/bridge.js

  4. Diff and renderer

    SSA-normalized structural diff; canvas graph renderer.

    src/diff.js

VizMLIR

Run it locally

git clone https://github.com/joepothiboot/vizmlir && cd vizmlir
git checkout b4f2c91a9b1734b03f8ea362c97e5597b0654080
rustup target add wasm32-unknown-unknown
npm ci && npm run dev

Needs Node 20+ and Rust with the wasm32-unknown-unknown target. wasm-opt (binaryen) is optional.

VizMLIR

Tests and benchmarks

Unit tests (vitest)

136 passed, 0 failed, 0 errors, 0 skipped/deselected (vitest)

Test log
✓ tests/unit/palette.test.js (5 tests) 4ms
 ✓ tests/unit/export.test.js (7 tests) 24ms
 ✓ tests/unit/abi-sync.test.js (9 tests) 23ms
 ✓ tests/unit/samples.test.js (8 tests) 83ms
 ✓ tests/unit/timing.test.js (19 tests) 57ms
 ✓ tests/unit/trace.test.js (44 tests) 29ms
 ✓ tests/unit/diff.test.js (18 tests) 748ms
   ✓ diffSnapshots > stays fast on large modules  688ms
 ✓ tests/unit/mlir-highlight.test.js (26 tests) 43ms
 Test Files  8 passed (8)
      Tests  136 passed (136)

Captured Output of npx vitest run at VizMLIR@b4f2c91 · darwin-arm64, LLVM 23.1.1, 2026-09-27

No benchmarks are published for this commit.

VizMLIR

Diff json-schema-mlir's canonicalization

json-schema-mlir canonicalization, diffed by VizMLIR

Pass SchemaCanonicalizerPass (schema-canonicalize) (#2) — changed the IR.

0 added, 4 removed, 0 changed (VizMLIR diffSnapshots)

  • − schema.validate_string
  • − schema.validate_string
  • − arith.andi
  • − arith.andi

Before: 9 nodes, 10 edges; VizMLIR diagnostics: undefined SSA value %arg0 ×4

After: 5 nodes, 4 edges; VizMLIR diagnostics: undefined SSA value %arg0 ×2

The IR VizMLIR parsed (paste into its baseline and current editors)

Input as parsed and printed by schema-opt

Stage: schema

module {
  func.func @fuse_constraint_tree(%arg0: !schema.value) -> i1 {
    %0 = schema.validate_string %arg0 {min_length = 5 : i64} : !schema.value
    %1 = schema.validate_string %arg0 {max_length = 32 : i64} : !schema.value
    %2 = schema.validate_string %arg0 {min_length = 2 : i64, pattern = "^[a-z]+$"} : !schema.value
    %3 = arith.andi %0, %1 : i1
    %4 = arith.andi %3, %2 : i1
    return %4 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

IR after SchemaCanonicalizerPass (schema-canonicalize)

Stage: schema

module {
  func.func @fuse_constraint_tree(%arg0: !schema.value) -> i1 {
    %0 = schema.validate_string %arg0 {max_length = 32 : i64, min_length = 5 : i64, pattern = "^[a-z]+$"} : !schema.value
    return %0 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-std-pipeline --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Captured Output of node scripts/capture.mjs (imports VizMLIR src/wasm/bridge.js, src/diff.js and runs public/mlir_core.wasm) at VizMLIR@b4f2c91 · darwin-arm64, LLVM 23.1.1, 2026-09-27

VizMLIR

Diff the lowering to standard dialects

json-schema-mlir lowering to std, diffed by VizMLIR

Pass LowerSchemaToStandardPass (lower-schema-to-std) (#3) — changed the IR.

23 added, 1 removed, 2 changed (VizMLIR diffSnapshots)

Show 26 diff rows
  • ~ module → [
  • ~ func.func @fuse_constraint_tree → func.func @__schema_rt_kind
  • + func.func @__schema_rt_as_f64
  • + func.func @__schema_rt_str_len
  • + func.func @__schema_rt_str_matches
  • + func.func @__schema_rt_str_format
  • + func.func @__schema_rt_has_field
  • + func.func @fuse_constraint_tree
  • + call @__schema_rt_kind
  • + arith.constant
  • + arith.cmpi
  • + scf.if
  • + func.call @__schema_rt_str_len
  • + arith.constant
  • + arith.cmpi
  • + arith.constant
  • + arith.cmpi
  • + arith.andi
  • + arith.constant
  • + func.call @__schema_rt_str_matches
  • + arith.andi
  • + scf.yield
  • + else
  • + arith.constant
  • + scf.yield
  • − schema.validate_string

Before: 5 nodes, 4 edges; VizMLIR diagnostics: undefined SSA value %arg0 ×2

After: 27 nodes, 31 edges; VizMLIR diagnostics: undefined SSA value %arg0 ×4

The IR VizMLIR parsed (paste into its baseline and current editors)

IR after SchemaCanonicalizerPass (schema-canonicalize)

Stage: schema

module {
  func.func @fuse_constraint_tree(%arg0: !schema.value) -> i1 {
    %0 = schema.validate_string %arg0 {max_length = 32 : i64, min_length = 5 : i64, pattern = "^[a-z]+$"} : !schema.value
    return %0 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-std-pipeline --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

IR after LowerSchemaToStandardPass (lower-schema-to-std)

Stage: llvm

module attributes {schema.string_pool = ["^[a-z]+$"]} {
  func.func private @__schema_rt_kind(i64) -> i32 attributes {llvm.readnone}
  func.func private @__schema_rt_as_f64(i64) -> f64 attributes {llvm.readnone}
  func.func private @__schema_rt_str_len(i64) -> i64 attributes {llvm.readnone}
  func.func private @__schema_rt_str_matches(i64, i64) -> i1 attributes {llvm.readnone}
  func.func private @__schema_rt_str_format(i64, i64) -> i1 attributes {llvm.readnone}
  func.func private @__schema_rt_has_field(i64, i64) -> i1 attributes {llvm.readnone}
  func.func @fuse_constraint_tree(%arg0: i64) -> i1 {
    %0 = call @__schema_rt_kind(%arg0) : (i64) -> i32
    %c3_i32 = arith.constant 3 : i32
    %1 = arith.cmpi eq, %0, %c3_i32 : i32
    %2 = scf.if %1 -> (i1) {
      %3 = func.call @__schema_rt_str_len(%arg0) : (i64) -> i64
      %c5_i64 = arith.constant 5 : i64
      %4 = arith.cmpi sge, %3, %c5_i64 : i64
      %c32_i64 = arith.constant 32 : i64
      %5 = arith.cmpi sle, %3, %c32_i64 : i64
      %6 = arith.andi %4, %5 : i1
      %c0_i64 = arith.constant 0 : i64
      %7 = func.call @__schema_rt_str_matches(%arg0, %c0_i64) : (i64, i64) -> i1
      %8 = arith.andi %6, %7 : i1
      scf.yield %8 : i1
    } else {
      %false = arith.constant false
      scf.yield %false : i1
    }
    return %2 : i1
  }
}

Captured Output of schema-opt fuse_constraint_tree.mlir --schema-to-std-pipeline --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at json-schema-mlir@6f4be81 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Captured Output of node scripts/capture.mjs (imports VizMLIR src/wasm/bridge.js, src/diff.js and runs public/mlir_core.wasm) at VizMLIR@b4f2c91 · darwin-arm64, LLVM 23.1.1, 2026-09-27

VizMLIR

Diff nano-dsp-mlir's dsp → linalg conversion

nano-dsp-mlir dsp → linalg, diffed by VizMLIR

Pass ConvertDSPToLinalg (convert-dsp-to-linalg) (#1) — changed the IR.

19 added, 0 removed, 6 changed (VizMLIR diffSnapshots)

Show 25 diff rows
  • ~ dsp.conv2d → tensor.empty
  • ~ dsp.add → arith.constant
  • ~ dsp.relu → linalg.fill
  • ~ tensor.cast → linalg.generic
  • + ^bb0
  • + arith.mulf
  • + arith.addf
  • + linalg.yield
  • ~ call @printMemrefF32 → ->
  • ~ return → tensor.empty
  • + linalg.generic
  • + ^bb0
  • + arith.addf
  • + linalg.yield
  • + ->
  • + tensor.empty
  • + arith.constant
  • + linalg.generic
  • + ^bb0
  • + arith.maximumf
  • + linalg.yield
  • + ->
  • + tensor.cast
  • + call @printMemrefF32
  • + return

Before: 13 nodes, 14 edges; no VizMLIR diagnostics

After: 32 nodes, 44 edges; no VizMLIR diagnostics

The IR VizMLIR parsed (paste into its baseline and current editors)

Input as parsed and printed by nanodsp-opt

Stage: dsp

module {
  func.func private @printMemrefF32(tensor<*xf32>)
  func.func @main() {
    %cst = arith.constant dense<[[[[1.000000e+00], [2.000000e+00], [3.000000e+00], [4.000000e+00]], [[5.000000e+00], [6.000000e+00], [7.000000e+00], [8.000000e+00]], [[9.000000e+00], [1.000000e+01], [1.100000e+01], [1.200000e+01]], [[1.300000e+01], [1.400000e+01], [1.500000e+01], [1.600000e+01]]]]> : tensor<1x4x4x1xf32>
    %cst_0 = arith.constant dense<1.000000e+00> : tensor<3x3x1x1xf32>
    %cst_1 = arith.constant dense<-6.000000e+01> : tensor<1x2x2x1xf32>
    %0 = dsp.conv2d %cst, %cst_0 : (tensor<1x4x4x1xf32>, tensor<3x3x1x1xf32>) -> tensor<1x2x2x1xf32>
    %1 = dsp.add %0, %cst_1 : tensor<1x2x2x1xf32>
    %2 = dsp.relu %1 : tensor<1x2x2x1xf32>
    %cast = tensor.cast %2 : tensor<1x2x2x1xf32> to tensor<*xf32>
    call @printMemrefF32(%cast) : (tensor<*xf32>) -> ()
    return
  }
}

Captured Output of nanodsp-opt test/Integration/DSPToLinalg/pipeline.mlir at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

IR after ConvertDSPToLinalg (convert-dsp-to-linalg)

Stage: linalg

#map = affine_map<(d0, d1, d2, d3, d4, d5, d6) -> (d0, d1 + d4, d2 + d5, d6)>
#map1 = affine_map<(d0, d1, d2, d3, d4, d5, d6) -> (d4, d5, d6, d3)>
#map2 = affine_map<(d0, d1, d2, d3, d4, d5, d6) -> (d0, d1, d2, d3)>
#map3 = affine_map<(d0, d1, d2, d3) -> (d0, d1, d2, d3)>
module {
  func.func private @printMemrefF32(tensor<*xf32>)
  func.func @main() {
    %cst = arith.constant dense<[[[[1.000000e+00], [2.000000e+00], [3.000000e+00], [4.000000e+00]], [[5.000000e+00], [6.000000e+00], [7.000000e+00], [8.000000e+00]], [[9.000000e+00], [1.000000e+01], [1.100000e+01], [1.200000e+01]], [[1.300000e+01], [1.400000e+01], [1.500000e+01], [1.600000e+01]]]]> : tensor<1x4x4x1xf32>
    %cst_0 = arith.constant dense<1.000000e+00> : tensor<3x3x1x1xf32>
    %cst_1 = arith.constant dense<-6.000000e+01> : tensor<1x2x2x1xf32>
    %0 = tensor.empty() : tensor<1x2x2x1xf32>
    %cst_2 = arith.constant 0.000000e+00 : f32
    %1 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<1x2x2x1xf32>) -> tensor<1x2x2x1xf32>
    %2 = linalg.generic {indexing_maps = [#map, #map1, #map2], iterator_types = ["parallel", "parallel", "parallel", "parallel", "reduction", "reduction", "reduction"]} ins(%cst, %cst_0 : tensor<1x4x4x1xf32>, tensor<3x3x1x1xf32>) outs(%1 : tensor<1x2x2x1xf32>) {
    ^bb0(%in: f32, %in_4: f32, %out: f32):
      %7 = arith.mulf %in, %in_4 : f32
      %8 = arith.addf %out, %7 : f32
      linalg.yield %8 : f32
    } -> tensor<1x2x2x1xf32>
    %3 = tensor.empty() : tensor<1x2x2x1xf32>
    %4 = linalg.generic {indexing_maps = [#map3, #map3, #map3], iterator_types = ["parallel", "parallel", "parallel", "parallel"]} ins(%2, %cst_1 : tensor<1x2x2x1xf32>, tensor<1x2x2x1xf32>) outs(%3 : tensor<1x2x2x1xf32>) {
    ^bb0(%in: f32, %in_4: f32, %out: f32):
      %7 = arith.addf %in, %in_4 : f32
      linalg.yield %7 : f32
    } -> tensor<1x2x2x1xf32>
    %5 = tensor.empty() : tensor<1x2x2x1xf32>
    %cst_3 = arith.constant 0.000000e+00 : f32
    %6 = linalg.generic {indexing_maps = [#map3, #map3], iterator_types = ["parallel", "parallel", "parallel", "parallel"]} ins(%4 : tensor<1x2x2x1xf32>) outs(%5 : tensor<1x2x2x1xf32>) {
    ^bb0(%in: f32, %out: f32):
      %7 = arith.maximumf %in, %cst_3 : f32
      linalg.yield %7 : f32
    } -> tensor<1x2x2x1xf32>
    %cast = tensor.cast %6 : tensor<1x2x2x1xf32> to tensor<*xf32>
    call @printMemrefF32(%cast) : (tensor<*xf32>) -> ()
    return
  }
}

Captured Output of nanodsp-opt test/Integration/DSPToLinalg/pipeline.mlir -convert-dsp-to-linalg -one-shot-bufferize=bufferize-function-boundaries -buffer-deallocation-pipeline -convert-linalg-to-loops -convert-scf-to-cf -expand-strided-metadata -lower-affine -convert-arith-to-llvm -finalize-memref-to-llvm -convert-func-to-llvm -convert-cf-to-llvm -reconcile-unrealized-casts --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Captured Output of node scripts/capture.mjs (imports VizMLIR src/wasm/bridge.js, src/diff.js and runs public/mlir_core.wasm) at VizMLIR@b4f2c91 · darwin-arm64, LLVM 23.1.1, 2026-09-27

nano-dsp-mlir

Architecture

  • dsp dialect: add, relu, matmul, conv2d with shape verifiers
  • dsp → linalg.generic conversion
  • Differential execution tests via mlir-runner
  • Compile-time pass timing
  • SIMD Mojo kernels and a C++ reference, tested against the same values; Mojo kernel benchmark
  1. dsp dialect

    Value-semantics ops on static f32 tensors, with verifiers.

    include/nanodsp/Dialect/DSP/IR/DSPOps.td

  2. -convert-dsp-to-linalg

    Rewrites each op to linalg.generic in destination-passing style.

    lib/Conversion/DSPToLinalg/DSPToLinalg.cpp

  3. Stock lowering (test oracle)

    Upstream bufferization and LLVM conversion, used only so the tests can execute.

    test/lit.cfg.py

  4. nanodsp-opt

    mlir-opt-style driver with the dialect and passes registered.

    tools/nanodsp-opt/nanodsp-opt.cpp

  5. Mojo kernels

    Generic SIMD add, relu, matmul and conv2d with the dialect's semantics; the width comes from the target at compile time.

    mojo/nanodsp/kernels.mojo

  6. C++ reference

    Scalar loop nests used as an independent check on both of the above.

    reference/nanodsp_ref.h

nano-dsp-mlir

Run it locally

git clone https://github.com/joepothiboot/nano-dsp-mlir && cd nano-dsp-mlir
git checkout ccb09e1b01eb6ae5dd0fe9e6108eeb90d2a08ae5
./test.sh   # finds Homebrew LLVM/MLIR or $MLIR_DIR, builds, runs check-nanodsp
pixi run test-mojo && pixi run test-reference && pixi run bench

Needs LLVM/MLIR with mlir-runner, FileCheck and lit. Clone into a path without spaces. The Mojo and C++ steps need only pixi, which installs the pinned Mojo.

nano-dsp-mlir

Tests and benchmarks

lit regression suite (check-nanodsp)

15 passed, 0 failed, 0 errors, 0 skipped/deselected (lit)

Test log
PASS: NANO-DSP-MLIR :: Dialect/DSP/invalid.mlir (1 of 15)
PASS: NANO-DSP-MLIR :: Dialect/DSP/canonicalize.mlir (2 of 15)
PASS: NANO-DSP-MLIR :: Dialect/DSP/ops.mlir (3 of 15)
PASS: NANO-DSP-MLIR :: Conversion/DSPToLinalg/relu.mlir (4 of 15)
PASS: NANO-DSP-MLIR :: Conversion/DSPToLinalg/matmul.mlir (5 of 15)
PASS: NANO-DSP-MLIR :: Conversion/DSPToLinalg/add.mlir (6 of 15)
PASS: NANO-DSP-MLIR :: Conversion/DSPToLinalg/conv2d.mlir (7 of 15)
PASS: NANO-DSP-MLIR :: Integration/DSPToLinalg/conv2d-multichannel.mlir (8 of 15)
PASS: NANO-DSP-MLIR :: Integration/DSPToLinalg/conv2d-strided.mlir (9 of 15)
PASS: NANO-DSP-MLIR :: Integration/end-to-end.mlir (10 of 15)
PASS: NANO-DSP-MLIR :: Integration/DSPToLinalg/relu.mlir (11 of 15)
PASS: NANO-DSP-MLIR :: Integration/DSPToLinalg/matmul.mlir (12 of 15)
PASS: NANO-DSP-MLIR :: Integration/DSPToLinalg/conv2d.mlir (13 of 15)
PASS: NANO-DSP-MLIR :: Integration/DSPToLinalg/pipeline.mlir (14 of 15)
PASS: NANO-DSP-MLIR :: Integration/DSPToLinalg/add.mlir (15 of 15)

Total Discovered Tests: 15
  Passed: 15 (100.00%)

Captured Output of lit -v build/test at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Mojo kernel tests

12 passed, 0 failed, 0 errors, 0 skipped/deselected (mojo)

Test log
nanodsp: 12 tests passed
✨ Pixi task (test-mojo): mojo run -I mojo mojo/tests/test_kernels.mojo

Captured Output of pixi run test-mojo at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

C++ reference tests

5 passed, 0 failed, 0 errors, 0 skipped/deselected (c++)

Test log
PASS add
PASS relu
PASS relu propagates NaN
PASS matmul
PASS conv2d

Captured Output of pixi run test-reference at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Compile-time pass timing

wall time per pass, one run; total 0.0071 s
PasssShare
Parser0.001014.3%
ConvertDSPToLinalg0.00033.9%
OneShotBufferizePass0.00034.3%
ExpandReallocPass0.00011.2%
CanonicalizerPass0.00022.2%
OwnershipBasedBufferDeallocationPass0.00011.9%
CanonicalizerPass0.00012.0%
BufferDeallocationSimplificationPass0.00011.6%
LowerDeallocationsPass0.00011.4%
CSEPass0.00000.6%
(A) DominanceInfo0.00000.0%
CanonicalizerPass0.00011.9%
ConvertLinalgToLoopsPass0.00045.7%
SCFToControlFlowPass0.00022.6%
ExpandStridedMetadataPass0.00022.8%
LowerAffinePass0.00011.9%
ArithToLLVMConversionPass0.00045.8%
FinalizeMemRefToLLVMConversionPass0.00068.0%
(A) DataLayoutAnalysis0.00000.1%
ConvertFuncToLLVMPass0.00034.4%
(A) DataLayoutAnalysis0.00000.1%
ConvertControlFlowToLLVMPass0.00045.1%
ReconcileUnrealizedCastsPass0.00022.5%
Output0.00045.3%
Rest0.001520.7%

Captured Output of nanodsp-opt test/Integration/DSPToLinalg/pipeline.mlir -convert-dsp-to-linalg -one-shot-bufferize=bufferize-function-boundaries -buffer-deallocation-pipeline -convert-linalg-to-loops -convert-scf-to-cf -expand-strided-metadata -lower-affine -convert-arith-to-llvm -finalize-memref-to-llvm -convert-func-to-llvm -convert-cf-to-llvm -reconcile-unrealized-casts --mlir-timing -o /dev/null at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Mojo kernel throughput (untiled, one core, best of 3-5 runs)

Exit code 0

matmul   64 x 64 x 64    0.031 ms   16.91251612903226 GFLOP/s
matmul   128 x 128 x 128    0.199 ms   21.076904522613066 GFLOP/s
matmul   256 x 256 x 256    1.443 ms   23.25324462924463 GFLOP/s
matmul   512 x 512 x 512    13.844 ms   19.39002138110373 GFLOP/s
conv2d   56 x 56 x 64 -> 64    13.663 ms   15.735259313474346 GFLOP/s
conv2d   28 x 28 x 128 -> 128    9.178 ms   21.72156373937677 GFLOP/s

Captured Output of pixi run bench at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

nano-dsp-mlir

The program

relu(conv2d(image) + bias) integration test

Stage: dsp

From test/Integration/DSPToLinalg/pipeline.mlir:1

// RUN: nanodsp-opt %s -convert-dsp-to-linalg \
// RUN: | mlir-opt %stock_lower_to_llvm \
// RUN: | mlir-runner -e main --entry-point-result=void \
// RUN:     --shared-libs=%mlir_runner_utils --shared-libs=%mlir_c_runner_utils \
// RUN: | FileCheck %s

func.func private @printMemrefF32(%ptr : tensor<*xf32>)

// relu(blur(image) + bias)
func.func @main() {
  %in = arith.constant dense<[[[[ 1.0],[ 2.0],[ 3.0],[ 4.0]],
                               [[ 5.0],[ 6.0],[ 7.0],[ 8.0]],
                               [[ 9.0],[10.0],[11.0],[12.0]],
                               [[13.0],[14.0],[15.0],[16.0]]]]> : tensor<1x4x4x1xf32>
  %k = arith.constant dense<1.0> : tensor<3x3x1x1xf32>

  // conv result is [[54, 63],[90, 99]]; bias drives two lanes negative.
  %bias = arith.constant dense<[[[[-60.0],[-60.0]],
                                 [[-60.0],[-60.0]]]]> : tensor<1x2x2x1xf32>

  %c = dsp.conv2d %in, %k : (tensor<1x4x4x1xf32>, tensor<3x3x1x1xf32>) -> tensor<1x2x2x1xf32>
  %s = dsp.add %c, %bias : tensor<1x2x2x1xf32>
  %r = dsp.relu %s : tensor<1x2x2x1xf32>

  // 54-60 = -6 -> 0 ; 63-60 =  3 ->  3
  // 90-60 = 30 -> 30 ; 99-60 = 39 -> 39
  %u = tensor.cast %r : tensor<1x2x2x1xf32> to tensor<*xf32>
  call @printMemrefF32(%u) : (tensor<*xf32>) -> ()
  return
}

// CHECK: rank = 4 offset = 0 sizes = [1, 2, 2, 1]
// CHECK: 0
// CHECK: 3
// CHECK: 30
// CHECK: 39

Static file test/Integration/DSPToLinalg/pipeline.mlir at nano-dsp-mlir@ccb09e1, shown verbatim.

nano-dsp-mlir

Verifier diagnostics

Verifier diagnostics for the invalid-op tests

nanodsp-opt exit code 1

Captured Output of nanodsp-opt test/Dialect/DSP/invalid.mlir -split-input-file -o /dev/null at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

nano-dsp-mlir

dsp → linalg

IR after ConvertDSPToLinalg (convert-dsp-to-linalg)

Stage: linalg

#map = affine_map<(d0, d1, d2, d3, d4, d5, d6) -> (d0, d1 + d4, d2 + d5, d6)>
#map1 = affine_map<(d0, d1, d2, d3, d4, d5, d6) -> (d4, d5, d6, d3)>
#map2 = affine_map<(d0, d1, d2, d3, d4, d5, d6) -> (d0, d1, d2, d3)>
#map3 = affine_map<(d0, d1, d2, d3) -> (d0, d1, d2, d3)>
module {
  func.func private @printMemrefF32(tensor<*xf32>)
  func.func @main() {
    %cst = arith.constant dense<[[[[1.000000e+00], [2.000000e+00], [3.000000e+00], [4.000000e+00]], [[5.000000e+00], [6.000000e+00], [7.000000e+00], [8.000000e+00]], [[9.000000e+00], [1.000000e+01], [1.100000e+01], [1.200000e+01]], [[1.300000e+01], [1.400000e+01], [1.500000e+01], [1.600000e+01]]]]> : tensor<1x4x4x1xf32>
    %cst_0 = arith.constant dense<1.000000e+00> : tensor<3x3x1x1xf32>
    %cst_1 = arith.constant dense<-6.000000e+01> : tensor<1x2x2x1xf32>
    %0 = tensor.empty() : tensor<1x2x2x1xf32>
    %cst_2 = arith.constant 0.000000e+00 : f32
    %1 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<1x2x2x1xf32>) -> tensor<1x2x2x1xf32>
    %2 = linalg.generic {indexing_maps = [#map, #map1, #map2], iterator_types = ["parallel", "parallel", "parallel", "parallel", "reduction", "reduction", "reduction"]} ins(%cst, %cst_0 : tensor<1x4x4x1xf32>, tensor<3x3x1x1xf32>) outs(%1 : tensor<1x2x2x1xf32>) {
    ^bb0(%in: f32, %in_4: f32, %out: f32):
      %7 = arith.mulf %in, %in_4 : f32
      %8 = arith.addf %out, %7 : f32
      linalg.yield %8 : f32
    } -> tensor<1x2x2x1xf32>
    %3 = tensor.empty() : tensor<1x2x2x1xf32>
    %4 = linalg.generic {indexing_maps = [#map3, #map3, #map3], iterator_types = ["parallel", "parallel", "parallel", "parallel"]} ins(%2, %cst_1 : tensor<1x2x2x1xf32>, tensor<1x2x2x1xf32>) outs(%3 : tensor<1x2x2x1xf32>) {
    ^bb0(%in: f32, %in_4: f32, %out: f32):
      %7 = arith.addf %in, %in_4 : f32
      linalg.yield %7 : f32
    } -> tensor<1x2x2x1xf32>
    %5 = tensor.empty() : tensor<1x2x2x1xf32>
    %cst_3 = arith.constant 0.000000e+00 : f32
    %6 = linalg.generic {indexing_maps = [#map3, #map3], iterator_types = ["parallel", "parallel", "parallel", "parallel"]} ins(%4 : tensor<1x2x2x1xf32>) outs(%5 : tensor<1x2x2x1xf32>) {
    ^bb0(%in: f32, %out: f32):
      %7 = arith.maximumf %in, %cst_3 : f32
      linalg.yield %7 : f32
    } -> tensor<1x2x2x1xf32>
    %cast = tensor.cast %6 : tensor<1x2x2x1xf32> to tensor<*xf32>
    call @printMemrefF32(%cast) : (tensor<*xf32>) -> ()
    return
  }
}

Captured Output of nanodsp-opt test/Integration/DSPToLinalg/pipeline.mlir -convert-dsp-to-linalg -one-shot-bufferize=bufferize-function-boundaries -buffer-deallocation-pipeline -convert-linalg-to-loops -convert-scf-to-cf -expand-strided-metadata -lower-affine -convert-arith-to-llvm -finalize-memref-to-llvm -convert-func-to-llvm -convert-cf-to-llvm -reconcile-unrealized-casts --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

nano-dsp-mlir

Every pass to the LLVM dialect

Pass events, in execution order
#PassEffectIR lines after
1ConvertDSPToLinalg (convert-dsp-to-linalg)changed37
2OneShotBufferizePass (one-shot-bufferize)changed40
3ExpandReallocPass (expand-realloc)unchanged40
4CanonicalizerPass (canonicalize)changed39
5OwnershipBasedBufferDeallocationPass (ownership-based-buffer-deallocation)changed50
6CanonicalizerPass (canonicalize)changed41
7BufferDeallocationSimplificationPass (buffer-deallocation-simplification)changed43
8LowerDeallocationsPass (bufferization-lower-deallocations)changed49
9CSEPass (cse)unchanged49
10CanonicalizerPass (canonicalize)changed42
11ConvertLinalgToLoopsPass (convert-linalg-to-loops)changed76
12SCFToControlFlowPass (convert-scf-to-cf)changed190
13ExpandStridedMetadataPass (expand-strided-metadata)unchanged190
14LowerAffinePass (lower-affine)changed189
15ArithToLLVMConversionPass (convert-arith-to-llvm)changed230
16FinalizeMemRefToLLVMConversionPass (finalize-memref-to-llvm)changed453
17ConvertFuncToLLVMPass (convert-func-to-llvm)changed455
18ConvertControlFlowToLLVMPass (convert-cf-to-llvm)changed474
19ReconcileUnrealizedCastsPass (reconcile-unrealized-casts)changed396

Captured Output of nanodsp-opt test/Integration/DSPToLinalg/pipeline.mlir -convert-dsp-to-linalg -one-shot-bufferize=bufferize-function-boundaries -buffer-deallocation-pipeline -convert-linalg-to-loops -convert-scf-to-cf -expand-strided-metadata -lower-affine -convert-arith-to-llvm -finalize-memref-to-llvm -convert-func-to-llvm -convert-cf-to-llvm -reconcile-unrealized-casts --mlir-print-ir-after-all --mlir-print-ir-module-scope --mlir-disable-threading -o /dev/null at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

nano-dsp-mlir

Execute and check

Executed with mlir-runner

Exit code 0

Unranked Memref base@ = 0x7c52d4d540 rank = 4 offset = 0 sizes = [1, 2, 2, 1] strides = [4, 2, 1, 1] data = 
[[[[0], 
   [3]], 
  [[30], 
   [39]]]]

All 5 expected lines found in order (checked by the lab build against the test's CHECK lines):

  • rank = 4 offset = 0 sizes = [1, 2, 2, 1]
  • 0
  • 3
  • 30
  • 39

Captured Output of nanodsp-opt test/Integration/DSPToLinalg/pipeline.mlir -convert-dsp-to-linalg -one-shot-bufferize=bufferize-function-boundaries -buffer-deallocation-pipeline -convert-linalg-to-loops -convert-scf-to-cf -expand-strided-metadata -lower-affine -convert-arith-to-llvm -finalize-memref-to-llvm -convert-func-to-llvm -convert-cf-to-llvm -reconcile-unrealized-casts | mlir-runner -e main --entry-point-result=void --shared-libs=$LLVM/lib/libmlir_runner_utils.dylib --shared-libs=$LLVM/lib/libmlir_c_runner_utils.dylib at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

nano-dsp-mlir

The same ops in Mojo and C++

Mojo kernel tests

12 passed, 0 failed, 0 errors, 0 skipped/deselected (mojo)

Test log
nanodsp: 12 tests passed
✨ Pixi task (test-mojo): mojo run -I mojo mojo/tests/test_kernels.mojo

Captured Output of pixi run test-mojo at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

C++ reference tests

5 passed, 0 failed, 0 errors, 0 skipped/deselected (c++)

Test log
PASS add
PASS relu
PASS relu propagates NaN
PASS matmul
PASS conv2d

Captured Output of pixi run test-reference at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

nano-dsp-mlir

Profiling

Compile-time pass timing

wall time per pass, one run; total 0.0071 s
PasssShare
Parser0.001014.3%
ConvertDSPToLinalg0.00033.9%
OneShotBufferizePass0.00034.3%
ExpandReallocPass0.00011.2%
CanonicalizerPass0.00022.2%
OwnershipBasedBufferDeallocationPass0.00011.9%
CanonicalizerPass0.00012.0%
BufferDeallocationSimplificationPass0.00011.6%
LowerDeallocationsPass0.00011.4%
CSEPass0.00000.6%
(A) DominanceInfo0.00000.0%
CanonicalizerPass0.00011.9%
ConvertLinalgToLoopsPass0.00045.7%
SCFToControlFlowPass0.00022.6%
ExpandStridedMetadataPass0.00022.8%
LowerAffinePass0.00011.9%
ArithToLLVMConversionPass0.00045.8%
FinalizeMemRefToLLVMConversionPass0.00068.0%
(A) DataLayoutAnalysis0.00000.1%
ConvertFuncToLLVMPass0.00034.4%
(A) DataLayoutAnalysis0.00000.1%
ConvertControlFlowToLLVMPass0.00045.1%
ReconcileUnrealizedCastsPass0.00022.5%
Output0.00045.3%
Rest0.001520.7%

Captured Output of nanodsp-opt test/Integration/DSPToLinalg/pipeline.mlir -convert-dsp-to-linalg -one-shot-bufferize=bufferize-function-boundaries -buffer-deallocation-pipeline -convert-linalg-to-loops -convert-scf-to-cf -expand-strided-metadata -lower-affine -convert-arith-to-llvm -finalize-memref-to-llvm -convert-func-to-llvm -convert-cf-to-llvm -reconcile-unrealized-casts --mlir-timing -o /dev/null at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27

Mojo kernel throughput (untiled, one core, best of 3-5 runs)

Exit code 0

matmul   64 x 64 x 64    0.031 ms   16.91251612903226 GFLOP/s
matmul   128 x 128 x 128    0.199 ms   21.076904522613066 GFLOP/s
matmul   256 x 256 x 256    1.443 ms   23.25324462924463 GFLOP/s
matmul   512 x 512 x 512    13.844 ms   19.39002138110373 GFLOP/s
conv2d   56 x 56 x 64 -> 64    13.663 ms   15.735259313474346 GFLOP/s
conv2d   28 x 28 x 128 -> 128    9.178 ms   21.72156373937677 GFLOP/s

Captured Output of pixi run bench at nano-dsp-mlir@ccb09e1 · darwin-arm64, LLVM 23.1.1, 2026-09-27