Senior AI Inference Optimization Engineer

DigitalOcean

Austin (TX)

On-site

USD 191,200 - 239,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

DigitalOcean is seeking a Senior Engineer 2 to drive architectural decisions that maximize throughput and minimize latency for large-model inference. You will lead benchmarking, performance optimization, and advanced kernel tuning across multi-node GPU clusters, working with CUDA, ROCm, and Triton toolchains.

You will mentor engineers, collaborate with product teams, and push for cutting-edge quantization and precision techniques to double throughput without sacrificing accuracy.

Qualifications

  • 5+ years in high-performance computing or AI infrastructure.
  • Deep familiarity with Gen AI landscape and major model families.
  • Hands-on optimization of attention layers and parallelization across distributed GPUs.
  • Comprehensive understanding of NVIDIA/AMD GPUs and software stacks.

Responsibilities

  • Lead technical strategy for benchmarking and performance optimization at inference engine and GPU kernel layers.
  • Engineer solutions for attention, memory, precision management, and multi-node GPU clusters.
  • Implement cutting-edge optimization techniques to stay at the forefront of Gen AI.
  • Mentor through code reviews and up-level engineering practices.
  • Collaborate with Product Management to translate hardware limits into shippable features.
  • Engage with GPU communities and contribute to open-source projects.

Skills

Technical depth
Gen AI literacy
Optimization expert
Hardware fluency
Open source mastery
Systems design
Leadership through influence
Low-level mastery
Triton/CUDA

Tools

CUDA
ROCm
TensorRT
OpenAI Triton

Job description

DigitalOcean is seeking a Senior Engineer 2 to drive architectural decisions that maximize throughput and minimize latency for large-model inference. You will lead benchmarking, performance optimization, and advanced kernel tuning across multi-node GPU clusters, working with CUDA, ROCm, and Triton toolchains.

You will mentor engineers, collaborate with product teams, and push for cutting-edge quantization and precision techniques to double throughput without sacrificing accuracy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,200 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,200 - 239,000
Senior AI Inference Optimization Engineer (Remote)
Senior AI Inference Optimization Engineer (Remote)

DigitalOcean • United States

Remote
USD 191,000 - 239,000
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Staff Engineer, Inference Optimizations
Staff Engineer, Inference Optimizations

DigitalOcean • United States

Remote
USD 191,000 - 239,000
Staff Engineer, Inference Optimizations
Staff Engineer, Inference Optimizations

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000