Model Bring-up Engineer / ML Compiler Engineer

General Compute

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

General Compute in San Francisco seeks a senior IC to bring up new models on our ASIC. You will own the end-to-end process from reference weights to first correct tokens, collaborating with the hardware partner to ensure correctness and performance.

You will build the agentic bringup loop, verify via diffs and parity checks, and optimize after correctness with fusion, quantization, and memory layout for production-ready models.

Qualifications

  • 5+ years in systems or ML systems, with real depth in at least one of: ML compilers, model porting/bringup, or high-performance kernels.
  • Strong on the internals of modern LLM inference: transformers, attention, KV cache, MoE routing, quantization, batching.
  • Fluent with agentic tooling and self-directed work style.
  • Comfortable inside a compiler stack — MLIR/LLVM, XLA, or a vendor graph compiler.

Responsibilities

  • Own model bringup end-to-end. Take a new architecture from reference weights to first correct tokens running on our ASIC, in days, not quarters.
  • Build the agentic bringup loop harness that compiles, runs, diffs against reference, and locates failing ops.
  • Live in the compiler space: graph capture, IR lowering, op coverage, kernel selection when a model won’t compile or numbers are off.
  • Own correctness before speed; build verification harness with layer-by-layer diffs, logit parity, end-to-end evals.
  • Then optimize: operator fusion, quantization, memory layout, batching and KV-cache behavior for our hardware.
  • Collaborate shoulder-to-shoulder with hardware partner’s compiler and runtime team to ship production-grade models.

Skills

ML systems
Model bringup
High-performance kernels
Transformers internals
KV cache
MoE routing
Quantization
Agentic tooling
Self-directed
Compiler stacks

Tools

MLIR/LLVM
XLA
CUDA
Triton
TensorRT-LLM
TVM
vLLM
TGI

Job description

About Us

We are the first AI inference neocloud, using ASIC compute to generate tokens 5–7× faster than existing GPU-based competitors. We just closed an oversubscribed seed round and quadrupled our compute allocation to $97M. Multiple lender conversations are live on a $200M asset-backed equipment facility.

About the role

You’ll take a new model and get it running — correctly — on our ASIC in record time. When a frontier model drops, the only question that matters is how fast we can land it on our silicon and start serving it. You own that loop: from reference weights, through the compiler, to first correct tokens. The low-level runtime is co-owned with our hardware partner today; your job is everything it takes to get a brand-new architecture compiled, verified, and fast on top of it.

The bet of this role is that bring-up should be an agentic loop, not a hand-port. You’ll build the harness of agents that compiles, runs, diffs against reference, and localizes failures — so the marginal model comes up faster than the last one did. Correctness first, optimization second: get it right, prove it’s right, then make it cheap. This is a senior IC role on a small team. You’ll own the bring-up pipeline, not tickets.

What You’ll Do:

  • Own model bringup end-to-end. Take a new architecture — a frontier LLM, an MoE, a multimodal model — from reference weights to first correct tokens running on our ASIC, in days, not quarters.

  • Build the agentic bringup loop. The differentiator isn’t hand-porting one model — it’s the harness of agents that compiles, runs, diffs against reference, localizes the failing op, and iterates without you in the inner loop. Each model you land should make the loop better at landing the next one.

  • Live in the compiler. Graph capture, IR lowering, op coverage, kernel selection — when a model won’t compile or produces wrong numbers, the fix is yours, whether it’s a missing lowering, a fused-kernel bug, or a numerics mismatch.

  • Own correctness before speed. Build the verification harness — layer-by-layer activation diffs, logit parity, end-to-end evals — that proves a freshly brought-up model matches reference before anyone trusts a token of it.

  • Then optimize. Once it’s correct, make it fast: operator fusion, quantization, memory layout, batching and KV-cache behavior on our hardware. Bringup gets it running; this is where it earns its cost-per-token.

  • Work shoulder-to-shoulder with our hardware partner’s compiler and runtime team. You’re the person who turns ‘the chip can technically run this’ into ’this model is live and correct in production.

What we need from you:

  • 5+ years in systems or ML systems, with real depth in at least one of: ML compilers, model porting/bringup, or high-performance kernels.

  • You’ve taken a model architecture you didn’t design and made it run — and run correctly — on a target it wasn’t written for. Numerics debugging doesn’t scare you.

  • Strong on the internals of modern LLM inference: transformers, attention, KV cache, MoE routing, quantization, batching. You can read a new model’s reference implementation and know what will be hard to lower.

  • Comfortable inside a compiler stack — MLIR/LLVM, XLA, or a vendor graph compiler — at the level of IR, lowering, and op coverage, not just calling into one.

  • Fluent with agentic tooling. You’d rather build the agent that runs the tedious bringup loop than run it by hand — and you have the taste to know where the loop still needs a human.

  • Self-directed. We don’t assign tickets — you’ll see the next model coming and have it half brought-up before anyone asks.

Nice-to-haves:

  • Have worked on a non-NVIDIA accelerator — TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar — at the compiler or model-bringup layer.

  • Kernel-level experience in CUDA, Triton, or a vendor kernel language. You know why a fused attention kernel beats three unfused ops.

  • Have built eval and numerics-verification harnesses (logit parity, activation diffing) for models in production.

  • Contributed to a graph compiler or serving runtime — XLA, TVM, MLIR, vLLM, TGI, TensorRT-LLM, or SGLang.

  • Have built agent loops or LLM-driven tooling that did real engineering work, not demos.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model Bringup Engineer / ML Compiler Engineer
Model Bringup Engineer / ML Compiler Engineer

General Compute Inc. • New York (NY)

On-site
USD 180,000 - 300,000
Head of Infrastructure
Head of Infrastructure

General Compute • San Francisco (CA)

On-site
USD 190,000 - 280,000
Member of Technical Staff, AI-Driven Compilation
Member of Technical Staff, AI-Driven Compilation

SF Tensor • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Equity and benefits
Office in San Francisco
Software Engineer, Inference Platform
Software Engineer, Inference Platform

General Compute Inc. • New York (NY)

On-site
USD 130,000 - 160,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

General Compute Inc. • San Francisco (CA)

On-site
USD 200,000 - 260,000
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems • Denver (CO)

On-site
USD 150,000 - 210,000
Health benefits
Paid time off
401(k) match
+1
Head of Infrastructure
Head of Infrastructure

General Compute Inc. • New York (NY)

On-site
USD 120,000 - 150,000
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000