Senior AI Inference Optimization Engineer

DigitalOcean

Austin (TX)

On-site

USD 191,200 - 239,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DigitalOcean is seeking a Senior Engineer 2 to drive architectural decisions that maximize throughput and minimize latency for large-model inference. You will lead benchmarking, performance optimization, and advanced kernel tuning across multi-node GPU clusters, working with CUDA, ROCm, and Triton toolchains.

You will mentor engineers, collaborate with product teams, and push for cutting-edge quantization and precision techniques to double throughput without sacrificing accuracy.

Qualifications

  • 5+ years in high-performance computing or AI infrastructure.
  • Deep familiarity with Gen AI landscape and major model families.
  • Hands-on optimization of attention layers and parallelization across distributed GPUs.
  • Comprehensive understanding of NVIDIA/AMD GPUs and software stacks.

Responsibilities

  • Lead technical strategy for benchmarking and performance optimization at inference engine and GPU kernel layers.
  • Engineer solutions for attention, memory, precision management, and multi-node GPU clusters.
  • Implement cutting-edge optimization techniques to stay at the forefront of Gen AI.
  • Mentor through code reviews and up-level engineering practices.
  • Collaborate with Product Management to translate hardware limits into shippable features.
  • Engage with GPU communities and contribute to open-source projects.

Skills

Technical depth
Gen AI literacy
Optimization expert
Hardware fluency
Open source mastery
Systems design
Leadership through influence
Low-level mastery
Triton/CUDA

Tools

CUDA
ROCm
TensorRT
OpenAI Triton

Job description

DigitalOcean is seeking a Senior Engineer 2 to drive architectural decisions that maximize throughput and minimize latency for large-model inference. You will lead benchmarking, performance optimization, and advanced kernel tuning across multi-node GPU clusters, working with CUDA, ROCm, and Triton toolchains.

You will mentor engineers, collaborate with product teams, and push for cutting-edge quantization and precision techniques to double throughput without sacrificing accuracy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,000 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,000 - 239,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Data Plane Engineer
Senior AI Inference Data Plane Engineer

DigitalOcean • Seattle (WA)

Hybrid
USD 139,000 - 174,000
Equity compensation
Hybrid work model
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000
Senior AI Systems Engineer — High-Performance Inference
Senior AI Systems Engineer — High-Performance Inference

DDN • United States

Remote
USD 150,000 - 210,000