Principal Software Architect

Oho Group

San Francisco (CA)

On-site

USD 260,000 - 380,000

Full time

14 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oho Group is hiring a senior architect to shape the software stack for a programmable accelerator platform. You will connect models, frameworks, compilers, runtimes and hardware from high-level workload behavior to kernel execution and code generation.

You will work across graph capture, MLIR/LLVM lowering, runtime and driver architecture, with a focus on SIMT execution, memory hierarchy and performance profiling, guiding senior engineers through design reviews and pre-silicon validation.

Qualifications

  • 12+ years across AI, compiler, runtime or systems software.
  • Strong C++ with LLVM and MLIR experience.
  • Expertise in GPU/SIMT execution and CUDA programming.
  • Understanding of modern AI models at graph and operator level.
  • Experience connecting frameworks, compilers, runtimes and hardware.
  • Hands-on, measurement-led approach to architecture.
  • Experience guiding senior engineers through design and code reviews.

Skills

C++
LLVM
MLIR
GPU/SIMT execution
CUDA programming
Framework integration
Performance profiling

Tools

PyTorch
XLA
Triton
CUDA
MLIR

Job description

I’m working with a well-funded AI compute company building a programmable accelerator platform and the complete software stack required to run modern AI workloads efficiently.

They are hiring a senior architect to shape the full path from models and frameworks through graph capture, compiler, runtime and driver layers, down to kernels and hardware execution. This is a rare role for someone who can connect high-level workload behaviour with low-level GPU architecture and code generation.

You will work across:

  • PyTorch, JAX and ONNX integration
  • Graph capture, partitioning and optimization
  • MLIR/LLVM lowering and backend code generation
  • Runtime, scheduling and driver architecture
  • SIMT execution, occupancy and memory hierarchy
  • CUDA and Triton programming models
  • Kernel fusion, tiling and custom operator lowering
  • Hardware/software co-design and platform validation
  • Performance profiling from models down to instructions

They are looking for:

  • 12+ years across AI, compiler, runtime or systems software
  • Strong C++ with deep LLVM and MLIR experience
  • Expertise in GPU/SIMT execution and CUDA programming
  • Understanding of modern AI models at graph and operator level
  • Experience connecting frameworks, compilers, runtimes and hardware
  • A hands-on, measurement-led approach to architecture
  • Experience guiding senior engineers through design and code reviews

Experience with custom accelerators, PyTorch compilation, XLA, Triton, quantization, performance modelling or pre-silicon software development would be particularly valuable.

This is an opportunity to define the software architecture for a new compute platform as it moves from a working stack into large-scale production.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Architect — GPUs, MLIR & CUDA
Senior AI Compute Architect — GPUs, MLIR & CUDA

Oho Group • San Francisco (CA)

On-site
USD 260,000 - 380,000
AI Platform Architect
AI Platform Architect

Cerebras • Austin (TX)

On-site
USD 180,000 - 280,000
GPU Architect
GPU Architect

EngineersOfAI • Milpitas (CA)

On-site
USD 140,000 - 190,000
Computer Architect
Computer Architect

Acceler8 Talent • Mountain View (CA)

Hybrid
USD 180,000 - 300,000
Rack-Scale AI Platform Architect
Rack-Scale AI Platform Architect

Cerebras • Austin (TX)

On-site
USD 180,000 - 280,000
Full-Stack AI Compute Architect
Full-Stack AI Compute Architect

Oxmiq Labs • Campbell (CA)

On-site
USD 230,000 - 360,000
Senior Microarchitect
Senior Microarchitect

Acceler8 Talent • Santa Clara (CA), Northern (KY)

Hybrid
USD 250,000 - 420,000
Performance Architect - AI Hardware
Performance Architect - AI Hardware

TEEMA • United States

Remote
USD 140,000 - 210,000
Compiler Runtime Engineer
Compiler Runtime Engineer

Oho Group • San Francisco (CA)

On-site
USD 150,000 - 210,000
Lead Kernel Engineer/Architect (m/f/d)
Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems • Germany (OH)

Hybrid
USD 104,000 - 152,000