Senior AI Inference Optimization Engineer

DigitalOcean

Seattle (WA)

On-site

USD 191,000 - 239,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

DigitalOcean is seeking a Senior Engineer 2 to lead architectural decisions for inference performance and to optimize GPU kernel and inference engine performance. You will solve memory bandwidth bottlenecks and guide the technical roadmap for a high-performance inference fleet.

You will collaborate with product teams and TPMs to turn theoretical hardware limits into shippable features while mentoring engineers and driving open-source collaboration in the GPU AI space.

Qualifications

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Strong familiarity with Gen AI landscapes and model families.
  • Hands-on optimization experience for attention layers and parallelization.
  • Deep understanding of NVIDIA/AMD GPU architectures and software stacks.
  • Open-source software integration and contribution experience.
  • Systems design with low-level GPU programming and memory access patterns.
  • Leadership by influence with cross-functional collaboration.
  • Proficiency with Triton or CUDA toolchains.

Responsibilities

  • Lead benchmarking and performance optimizations for inference engines and GPU kernels.
  • Engineer solutions for memory bandwidth and compute utilization bottlenecks.
  • Develop and deploy FP8, BF16, FP4 quantization techniques.
  • Identify kernel fusion opportunities and multi-node GPU parallelization.
  • Collaborate with product teams to translate hardware limits into features.
  • Mentor through code reviews and elevate technical standards.

Skills

High-performance computing
Gen AI literacy
Attention optimization
Distributed GPU
CUDA proficiency
ROCm proficiency
Triton
Open source

Tools

CUDA
ROCm
TensorRT
Triton compiler

Job description

DigitalOcean is seeking a Senior Engineer 2 to lead architectural decisions for inference performance and to optimize GPU kernel and inference engine performance. You will solve memory bandwidth bottlenecks and guide the technical roadmap for a high-performance inference fleet.

You will collaborate with product teams and TPMs to turn theoretical hardware limits into shippable features while mentoring engineers and driving open-source collaboration in the GPU AI space.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,000 - 239,000
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,000 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,000 - 239,000
Senior Engineer II: Cloud AI Inference & Scalable Systems
Senior Engineer II: Cloud AI Inference & Scalable Systems

digitalocean98 • Seattle (WA)

Hybrid
USD 167,000 - 209,000
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
AI Inference Performance & Scale Engineer
AI Inference Performance & Scale Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance
Senior AI Systems Engineer — High-Performance Inference
Senior AI Systems Engineer — High-Performance Inference

DDN • United States

Remote
USD 150,000 - 210,000
Senior AI Systems Engineer: GPU Kernels & Inference Equity
Senior AI Systems Engineer: GPU Kernels & Inference Equity

NVIDIA AI • Michigan

On-site
USD 150,000 - 190,000
Equity
Health Insurance
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000