Running an ML Engine on a GPU Its Own Toolkit Had Never Heard Of

TL;DR: The PTX technique that got ONNX Runtime and TensorRT onto Grace Blackwell ARM64 months before the toolchain supported it — now an MIT-licensed repo, in case it saves you a weekend.

In November 2025 I got a photo-management app’s machine-learning engine running at ~95% GPU utilization on a chip that, as far as its own CUDA toolkit was concerned, did not exist.

The app was Immich. The GPU was an NVIDIA Grace Blackwell GB10 — the chip in the DGX Spark server I call Sparky. And the problem, in one sentence, was this: the GPU was newer than the software built to talk to it.

The technique that solved it is the useful part, it generalizes well past this particular chip, and the whole thing is on GitHub under an MIT license. Here’s how it works.

The wall

Immich’s ML container handles face detection and CLIP-based smart search, running on ONNX Runtime with optional TensorRT acceleration. In November 2025 there were no pre-built ARM64 ONNX Runtime wheels with TensorRT support, so it had to be compiled from source. Fine. I’ve built worse.

Then the hardware bit back. The GB10’s compute capability is sm_121, and here’s the mismatch:

  • The host driver on Sparky was CUDA 13.0, which knows about GB10.
  • The container’s CUDA toolkit was 12.2.2, which does not.

Which left two options, both bad. Build for a generic GPU architecture and ONNX Runtime compiles fine — then dies the instant it touches the GPU with cudaErrorNoKernelImageForDevice, CUDA’s polite way of saying “I have no machine code that runs on this card.” Or target sm_121 directly and it won’t even compile: nvcc fatal: Unsupported gpu architecture 'compute_121'. The toolkit had simply never heard of the chip.

The build history, abbreviated:

Builds What I tried What happened
early Chasing dependencies — cuDNN, CMake, Eigen, pip resolution Failures at every stage
middle Generic CUDA architectures Built, then cudaErrorNoKernelImageForDevice at runtime
v14 Target sm_121 explicitly Unsupported gpu architecture 'compute_121' — wouldn’t build
v15 PTX for a virtual architecture It worked.

The trick

The thing that unlocked it is a feature of NVIDIA’s toolchain I should have reached for sooner.

You don’t have to compile all the way down to a specific GPU. You can compile to PTX — NVIDIA’s architecture-independent intermediate representation, essentially bytecode for GPUs. PTX gets JIT-compiled into native machine code by the driver, at runtime. And the driver on my box — CUDA 13.0 — absolutely knew what a GB10 was.

So instead of fighting the toolkit, I told it to emit PTX for the newest virtual architecture it did understand:

--cmake_extra_defines CMAKE_CUDA_ARCHITECTURES="89-virtual"

89-virtual produces PTX for compute_89 — the highest virtual arch CUDA 12.2 knows. No sm_121 ever appears in the build. At runtime the CUDA 13.0 driver JIT-compiled that PTX straight into native GB10 code.

The error vanished. ONNX Runtime came up with all three execution providers — TensorRT, CUDA, and CPU. The GPU pinned at ~95% during inference, 5–10× faster than CPU-only.

The part worth keeping

When your container’s CUDA toolkit is older than your GPU, build for the newest virtual (PTX) architecture the toolkit supports and let the newer host driver JIT it the rest of the way.

It costs a few seconds of JIT on first run and nothing after that. And it will keep being useful, because the underlying situation recurs on a schedule: toolchains trail the newest silicon by a few months, every single generation. PTX-plus-driver-JIT is the bridge across that gap, and it was designed to be exactly that.

As far as I can document, this was the first working ONNX Runtime + TensorRT build on Blackwell-generation ARM64. I say “as far as I can document” on purpose — I looked, and I didn’t find a prior example. That’s the honest version of the claim.

Why I stopped here, and why that’s the good ending

I ran it, I proved the technique, and then I made a call to set Immich aside. The GPU was only part of the story. My library is built around Canon CR3 RAW files, and I never got Immich on this machine to read them. I built a custom server image with the RAW libraries installed, but it never produced a working CR3 thumbnail — when I retested it months later, a CR3 still came back as “not a known file format.” Indexing a 200,000-image library was also slower than I could live with day to day. A photo manager that can’t show you most of your photos isn’t one you keep.

On top of that, every Immich release would have meant rebuilding a 5–7 hour custom ML container with no path to upstream the fix, and on ARM64 that means keeping a parallel build infrastructure alive indefinitely. That’s a part-time job, and it wasn’t the job I wanted.

So I wrote down what I’d learned and let the ecosystem catch up. It did, quickly:

  • CUDA 12.9 added native sm_121 support and 13.x supports it fully, so for Blackwell specifically you now target arch 121 directly and skip the gymnastics entirely.
  • In March 2026 a developer (@volschin) published a working DGX-Spark Dockerfile for Immich on a CUDA 13 base doing exactly that — the clean path, four months after I did it the hard way.

I count that as a win rather than a loss. The specific workaround aged out in four months. The general technique didn’t, and won’t.

It’s yours

The Dockerfile, the GPU monitoring scripts, and a curated build history — both dead ends and the attempt that worked — are in a small MIT-licensed repo: onnxruntime-arm64-blackwell-ptx. Archived rather than maintained; the technique is the deliverable.

It built, it ran, and now it’s yours. If it saves you a weekend on whatever absurdly new GPU you just got your hands on, I’m happy I could help.

The projects, experience and opinions here are mine. AI helped me turn my notes and build records into this piece and polished it for Cairoglyphics.ai.