End-to-End AI Compute Architect — High-Performance Stack

Oxmiq Labs

Campbell (CA)

On-site

USD 230,000 - 360,000

Full time

42 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oxmiq Labs in California seeks a Full-Stack AI Compute Architect to define the software stack—from model ingestion and graph capture through compiler, runtime, and kernels—driving high-performance, CUDA-compatible solutions.

You will own end-to-end performance, guide MLIR/LLVM strategies, and collaborate with silicon teams to scale workloads for inference at real customer scale. This role requires fluency across ML frameworks and deep hardware knowledge.

Qualifications

  • 12+ years building systems and/or AI-infrastructure software, owning architecture across multiple layers.
  • Breadth across the AI software stack from model to kernel and instruction level.
  • Hands-on experience with ML frameworks (PyTorch, JAX, ONNX) and models.
  • Strong expertise in C++ and compiler infra (LLVM/MLIR).
  • Deep expertise in SIMT and parallel compute architectures, including CUDA.

Responsibilities

  • Refine the technical direction of OxPython and the software stack from front-ends to kernels.
  • Own how AI models map onto the platform and ensure scalable, efficient inference.
  • Drive end-to-end performance, remove bottlenecks across layers, close gaps to peak hardware utilization.
  • Guide graph and compiler strategy: MLIR dialects, IR design, LLVM back-end codegen.
  • Architect work orchestration across OxCore’s scalar, tensor, and SIMT engines, maintaining day-zero support.
  • Advance SIMT execution and memory hierarchy, improve thread/warp scheduling and occupancy.
  • Refine kernel programming models, kernel fusion, tiling with kernel engineers.
  • Design for scale and production, iterate with hardware teams, co-verify stack and silicon.

Skills

C++
LLVM/MLIR
SIMT architectures
CUDA
Compiler design

Education

BS/MS/PhD in CS/CE/EE

Tools

PyTorch
ONNX
Triton

Job description

Oxmiq Labs in California seeks a Full-Stack AI Compute Architect to define the software stack—from model ingestion and graph capture through compiler, runtime, and kernels—driving high-performance, CUDA-compatible solutions.

You will own end-to-end performance, guide MLIR/LLVM strategies, and collaborate with silicon teams to scale workloads for inference at real customer scale. This role requires fluency across ML frameworks and deep hardware knowledge.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Full-Stack AI Compute Architect
Full-Stack AI Compute Architect

Oxmiq Labs • Campbell (CA)

On-site
USD 230,000 - 360,000
Senior AI Compute Architect — GPUs, MLIR & CUDA
Senior AI Compute Architect — GPUs, MLIR & CUDA

Oho Group • San Francisco (CA)

On-site
USD 260,000 - 380,000
Principal Software Architect
Principal Software Architect

Oho Group • San Francisco (CA)

On-site
USD 260,000 - 380,000
AI Compute Architect: Next-Gen Silicon & ML-Driven Design
AI Compute Architect: Next-Gen Silicon & ML-Driven Design

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 180,000 - 300,000
GenAI Hardware Architect for High-Performance Silicon
GenAI Hardware Architect for High-Performance Silicon

MatX • Mountain View (CA)

Hybrid
USD 120,000 - 600,000
4 weeks PTO + holidays
Health insurance
Remote work up to 3 weeks
+1
Principal Performance Modeling Architect
Principal Performance Modeling Architect

Oxmiq Labs • Campbell (CA)

On-site
USD 180,000 - 240,000
Computer Architect
Computer Architect

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 180,000 - 300,000
Principal Architect | AI Platform
Principal Architect | AI Platform

Oxmiq Labs • San Francisco (CA)

On-site
USD 250,000 - 400,000
Equity participation
Medical coverage
Staff Design Engineer – TPU/GPU Design
Staff Design Engineer – TPU/GPU Design

Oxmiq Labs • Campbell (CA)

On-site
USD 180,000 - 240,000
Equity participation
Medical, dental, and vision coverage
AI Platform Architect
AI Platform Architect

Cerebras • Austin (TX)

On-site
USD 180,000 - 280,000