Model Bringup Engineer / ML Compiler Engineer

General Compute Inc.

New York (NY)

On-site

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

General Compute Inc. seeks a senior IC to own bringup and validation of frontier LLM architectures on our ASIC. You’ll build and own the end-to-end bringup loop from reference weights to first correct tokens, working with a hardware partner.

Focus is correctness first, then optimization, with responsibilities spanning compiler internals, graph lowering, and high-performance kernels. You’ll ship a production-ready path that makes models run fast and correctly.

Qualifications

  • 5+ years in systems or ML systems.
  • Deep depth in at least one: ML compilers, model porting/bringup, or high-performance kernels.
  • Strong knowledge of modern LLM inference: transformers, attention, KV cache, MoE routing, quantization, batching.
  • Comfortable inside a compiler stack (MLIR/LLVM, XLA, or vendor graph compiler).
  • Fluent with agentic tooling and proactive bringup.
  • Self-directed; not assigned tickets.

Responsibilities

  • Own model bringup end-to-end, from reference weights to first correct tokens on an ASIC.
  • Build the agentic bringup loop: harness of agents that compiles, runs, diffs, and localizes failures.
  • Work in the compiler space: graph capture, IR lowering, op coverage, kernel selection; fix lowering or fused-kernel issues.
  • Own correctness before speed: verification harness and end-to-end evals to match reference.
  • Then optimize: operator fusion, quantization, memory layout, batching, KV-cache behavior.
  • Collaborate with hardware partner's compiler and runtime team to productionize models.

Skills

ML systems
Model bringup
Compiler stacks
LLM inference internals
Agentic tooling

Job description

Take a new model and get it running — correctly — on our ASIC in record time. When a frontier model drops, the only question that matters is how fast we can land it on our silicon and start serving it. You own that loop: from reference weights, through the compiler, to first correct tokens. The low-level runtime is co-owned with our hardware partner today; your job is everything it takes to get a brand-new architecture compiled, verified, and fast on top of it.

The bet of this role is that bringup should be an agentic loop, not a hand-port. You'll build the harness of agents that compiles, runs, diffs against reference, and localizes failures — so the marginal model comes up faster than the last one did. Correctness first, optimization second: get it right, prove it's right, then make it cheap. This is a senior IC role on a small team. You'll own the bringup pipeline, not tickets.

Responsibilities
  • Own model bringup end-to-end. Take a new architecture — a frontier LLM, an MoE, a multimodal model — from reference weights to first correct tokens running on our ASIC, in days, not quarters.
  • Build the agentic bringup loop. The differentiator isn't hand-porting one model — it's the harness of agents that compiles, runs, diffs against reference, localizes the failing op, and iterates without you in the inner loop. Each model you land should make the loop better at landing the next one.
  • Live in the compiler. Graph capture, IR lowering, op coverage, kernel selection — when a model won't compile or produces wrong numbers, the fix is yours, whether it's a missing lowering, a fused-kernel bug, or a numerics mismatch.
  • Own correctness before speed. Build the verification harness — layer-by-layer activation diffs, logit parity, end-to-end evals — that proves a freshly brought-up model matches reference before anyone trusts a token of it.
  • Then optimize. Once it's correct, make it fast: operator fusion, quantization, memory layout, batching and KV-cache behavior on our hardware. Bringup gets it running; this is where it earns its cost-per-token.
  • Work shoulder-to-shoulder with our hardware partner's compiler and runtime team. You're the person who turns 'the chip can technically run this' into 'this model is live and correct in production.'
What we're looking for
  • 5+ years in systems or ML systems, with real depth in at least one of: ML compilers, model porting/bringup, or high-performance kernels.
  • You've taken a model architecture you didn't design and made it run — and run correctly — on a target it wasn't written for. Numerics debugging doesn't scare you.
  • Strong on the internals of modern LLM inference: transformers, attention, KV cache, MoE routing, quantization, batching. You can read a new model's reference implementation and know what will be hard to lower.
  • Comfortable inside a compiler stack — MLIR/LLVM, XLA, or a vendor graph compiler — at the level of IR, lowering, and op coverage, not just calling into one.
  • Fluent with agentic tooling. You'd rather build the agent that runs the tedious bringup loop than run it by hand — and you have the taste to know where the loop still needs a human.
  • Self-directed. We don't assign tickets — you'll see the next model coming and have it half brought-up before anyone asks.
Nice to have
  • Have worked on a non-NVIDIA accelerator — TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar — at the compiler or model-bringup layer.
  • Kernel-level experience in CUDA, Triton, or a vendor kernel language. You know why a fused attention kernel beats three unfused ops.
  • Have built eval and numerics-verification harnesses (logit parity, activation diffing) for models in production.
  • Contributed to a graph compiler or serving runtime — XLA, TVM, MLIR, vLLM, TGI, TensorRT-LLM, or SGLang.
  • Have built agent loops or LLM-driven tooling that did real engineering work, not demos.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Platform
Software Engineer, Inference Platform

General Compute Inc. • New York (NY)

On-site
USD 130,000 - 160,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

General Compute Inc. • San Francisco (CA)

On-site
USD 200,000 - 260,000
Head of Infrastructure
Head of Infrastructure

General Compute Inc. • New York (NY)

On-site
USD 120,000 - 150,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Senior Model Bringup & ML Compiler Engineer
Senior Model Bringup & ML Compiler Engineer

General Compute Inc. • New York (NY)

On-site
USD 180,000 - 300,000
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda • California (MO)

On-site
USD 120,000 - 190,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda • California (MO)

On-site
USD 180,000 - 240,000
Compiler Optimization Engineer
Compiler Optimization Engineer

Lemurian Labs Inc. • Santa Clara (CA)

On-site
USD 180,000 - 230,000
Equity
Medical benefits
Retirement plan
+1
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000