Full-Stack AI Compute Architect

Oxmiq Labs

Campbell (CA)

On-site

USD 230,000 - 360,000

Full time

39 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oxmiq Labs in California seeks a Full-Stack AI Compute Architect to define the software stack—from model ingestion and graph capture through compiler, runtime, and kernels—driving high-performance, CUDA-compatible solutions.

You will own end-to-end performance, guide MLIR/LLVM strategies, and collaborate with silicon teams to scale workloads for inference at real customer scale. This role requires fluency across ML frameworks and deep hardware knowledge.

Qualifications

  • 12+ years building systems and/or AI-infrastructure software, owning architecture across multiple layers.
  • Breadth across the AI software stack from model to kernel and instruction level.
  • Hands-on experience with ML frameworks (PyTorch, JAX, ONNX) and models.
  • Strong expertise in C++ and compiler infra (LLVM/MLIR).
  • Deep expertise in SIMT and parallel compute architectures, including CUDA.

Responsibilities

  • Refine the technical direction of OxPython and the software stack from front-ends to kernels.
  • Own how AI models map onto the platform and ensure scalable, efficient inference.
  • Drive end-to-end performance, remove bottlenecks across layers, close gaps to peak hardware utilization.
  • Guide graph and compiler strategy: MLIR dialects, IR design, LLVM back-end codegen.
  • Architect work orchestration across OxCore’s scalar, tensor, and SIMT engines, maintaining day-zero support.
  • Advance SIMT execution and memory hierarchy, improve thread/warp scheduling and occupancy.
  • Refine kernel programming models, kernel fusion, tiling with kernel engineers.
  • Design for scale and production, iterate with hardware teams, co-verify stack and silicon.

Skills

C++
LLVM/MLIR
SIMT architectures
CUDA
Compiler design

Education

BS/MS/PhD in CS/CE/EE

Tools

PyTorch
ONNX
Triton

Job description

Updated Role | Now hiring: Full-Stack AI Compute Architect
About OXMIQ

OXMIQ provides complete hardware and software GPU IP that lets our customers build their own AI silicon. Founded by Raja Koduri, we recently closed a $35M Series A (co-led by Samsung Catalyst Fund and Fundomo, with MediaTek, Intel Capital, and others), bringing total funding to $60M, with Jim Keller joining our board.

Our OxCore architecture combines scalar, tensor, SIMT, and orchestration engines on one platform — delivering CUDA compatibility alongside a fully programmable architecture. On top of it runs OxPython, our software stack that runs existing CUDA and PyTorch code unmodified with day-zero support for new AI models. We have an established team and a working stack, and we're now shifting from development to production — hardening the platform and scaling it for inference at real customer scale.

You'll join the OXMIQ architecture team, which sets the direction of our software, silicon, and systems. You'll help refine our software strategy and deliver the highest-performance solutions by optimizing the whole stack — from AI models and frameworks at the top down to kernel and hardware-level code generation.

The Scope of This Role

This role spans the AI software stack from models on down — the full vertical path a workload travels: from a model authored in frameworks such as PyTorch, JAX, or ONNX, through graph capture and the compiler, into the runtime and scheduler, down to individual kernels and the instructions that execute on OxCore.

We're looking for an architect who is fluent across that whole range — someone who can reason about how a transformer or diffusion model is structured and about warp divergence and register allocation, and who can make the choices that connect those layers into one coherent, high-performance platform. You won't be the deepest specialist at every layer, but you should be able to hold the whole range in your head and drive decisions end to end.

Key Responsibilities
  • Refine the technical direction of OxPython, the OXMIQ software stack — from framework front-ends (PyTorch, JAX, ONNX) and model ingestion, through graph capture, the compiler, runtime, and driver layers, down to kernels — in close partnership with the existing software team.
  • Own how AI models map onto the platform: understand how modern workloads (LLMs, diffusion, vision, agentic inference, and beyond) are structured and expressed at the framework level, and shape how they are captured, partitioned, and lowered so they run unmodified with day-zero support and scale efficiently for inference.
  • Own whole-stack performance: drive optimization end to end, identify and remove bottlenecks at every layer — from graph-level and framework-level inefficiencies down through compiler, runtime, driver, and kernels — and close the gap to peak hardware utilization.
  • Guide the graph and compiler strategy: MLIR dialect and IR design, lowering pipelines, operator representation, and LLVM backend code generation targeting Oxmiq hardware IP.
  • Architect work orchestration across OxCore's scalar, tensor, and SIMT engines, and shape the compiler/runtime approach that runs CUDA-optimized models unmodified — preserving day-zero support for new AI models without sacrificing programmability.
  • Advance the SIMT execution and parallel compute strategy: thread/warp scheduling, divergence handling, occupancy, and the memory hierarchy, and how parallel workloads are expressed, lowered, and mapped onto the hardware.
  • Refine and optimize the strategy for kernel programming models (including Triton-based and custom op lowering), kernel fusion, and tiling, in collaboration with kernel engineers.
  • Design for scale and production: ensure solutions grow cleanly with customer workloads, deployment sizes, and model complexity, and harden them for real-world delivery.
  • Work with architecture and hardware teams to translate performance requirements into efficient code generation and runtime strategies, and feed software needs back into hardware definition.
  • Drive the software/hardware validation strategy — evolving how the stack and silicon are co-verified as we scale from bring-up through production.
  • Raise the technical bar through design reviews, code reviews, and architecture discussions, and mentor engineers across the frameworks, compiler, runtime, and kernel teams.
  • Guide performance profiling, benchmarking, and root cause analysis for compiler- and runtime-generated code.
Required Qualifications
  • 12+ years building systems and/or AI-infrastructure software, with a track record of owning architecture for a major component or a full stack that spans multiple layers.
  • Breadth across the AI software stack — you can reason about the problem from the model and framework level down to the kernel and instruction level, and understand how choices at one layer constrain the others.
  • Hands-on experience with ML frameworks (PyTorch, JAX, ONNX) and a solid understanding of how deep learning models are structured, expressed, captured, and optimized — not just at the operator level but as whole graphs and workloads.
  • Strong expertise in C++ and compiler infrastructure, including LLVM (IR, pass infrastructure, backend code generation) and MLIR (dialect design, conversion passes, progressive lowering).
  • Deep expertise in SIMT and parallel compute architectures — thread/warp execution models, divergence, synchronization, occupancy constraints, and memory hierarchies.
  • Proven experience with GPU programming models, including CUDA, and an understanding of what it takes to deliver CUDA compatibility on non-NVIDIA hardware.
  • Solid command of code generation concepts — instruction selection, scheduling, register allocation, vectorization, and memory hierarchy optimization.
  • Experience co-designing software alongside silicon or shaping a hardware-software interface.
  • Strong debugging and performance analysis skills, and a bias toward shipping: you prototype, measure, and make decisions with incomplete information.
  • Hands-on experience with Claude Code or equivalent AI-assisted development workflows.

Note: we expect exceptional depth in some of these areas and working fluency across the rest. The essential requirement is the ability to operate across the whole stack, not mastery of every layer.

Preferred Qualifications
  • Experience architecting compiler backends or runtimes for custom AI accelerators, NPUs, or GPU IP.
  • Experience with graph-level model optimization, framework integration internals (e.g., PyTorch compile/inductor, XLA), or serving/inference infrastructure for large models.
  • Experience with Triton or similar high-level GPU kernel frameworks, and with kernel fusion, operator tiling, and auto-tuning.
  • Exposure to quantization, mixed-precision, or sparsity-aware compilation techniques.
  • Experience with performance modeling and roofline analysis for GPU workloads.
  • Background in deep learning, computer vision, image processing, or video — the workloads our customers run.
  • Experience bringing up and scaling a software stack on new or pre-silicon hardware into production.
  • Contributions to open-source compiler, runtime, or ML-framework projects such as LLVM, MLIR, PyTorch, or Triton.
Education
  • BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
Why OXMIQ

You'll join a well-funded company — fresh off a $35M Series A backed by Samsung, MediaTek, Intel Capital, and others, with Jim Keller on our board — with an established team and a proven stack, at the point where we're scaling it into a production platform in customers' hands. You'll work directly with the people designing the hardware it runs on.

This is a role for an architect who wants to refine and optimize the AI software stack end to end — from models and frameworks down to kernels and silicon — rather than just one layer of it, and measure success in shipped, high-performance solutions.

OXMIQ is an equal opportunity employer. We welcome applicants of all backgrounds.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Architect | AI Platform
Principal Architect | AI Platform

Oxmiq Labs • San Francisco (CA)

On-site
USD 250,000 - 400,000
Equity participation
Medical coverage
Principal Performance Modeling Architect
Principal Performance Modeling Architect

Oxmiq Labs • Campbell (CA)

On-site
USD 180,000 - 240,000
Staff Design Engineer – TPU/GPU Design
Staff Design Engineer – TPU/GPU Design

Oxmiq Labs • Campbell (CA)

On-site
USD 180,000 - 240,000
Equity participation
Medical, dental, and vision coverage
Principal Software Architect
Principal Software Architect

Oho Group • San Francisco (CA)

On-site
USD 260,000 - 380,000
End-to-End AI Compute Architect — High-Performance Stack
End-to-End AI Compute Architect — High-Performance Stack

Oxmiq Labs • Campbell (CA)

On-site
USD 230,000 - 360,000
Software Technical Program Manager
Software Technical Program Manager

Socket.dev • Los Altos (CA)

Hybrid
USD 180,000 - 240,000
ML Research Engineer - Hardware Codesign OpenAI San Francisco
ML Research Engineer - Hardware Codesign OpenAI San Francisco

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Performance Architect - AI Hardware
Performance Architect - AI Hardware

TEEMA • United States

Remote
USD 140,000 - 210,000
Product Manager – AI SW Infrastructure (AI Software Stack / Compiler / Platform) - US
Product Manager – AI SW Infrastructure (AI Software Stack / Compiler / Platform) - US

Majestic Labs • Los Altos (CA)

On-site
USD 150,000 - 190,000
Architect
Architect

MatX • Mountain View (CA)

Hybrid
USD 120,000 - 600,000
4 weeks PTO + holidays
Health insurance
Remote work up to 3 weeks
+1