Remote Senior AI Inference Optimization Engineer

DigitalOcean

San Francisco (CA)

On-site

USD 191,200 - 239,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity compensation
Remote work

Job summary

DigitalOcean is seeking a Senior Engineer 2 to drive AI Inference Optimization. You will lead benchmarking and optimizations at the inference engine and GPU kernel levels, and guide the technical roadmap for high-performance inference fleets.

Bring 5+ years of HPC or AI infra experience, deep Gen AI knowledge, and hands-on CUDA/Triton expertise to push throughput and reduce latency across multi-node GPU clusters.

Qualifications

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Experience solving compute utilization and memory bandwidth bottlenecks.
  • Familiarity with Gen AI (LLM, VLM, LMM) landscape and model families.
  • Proficiency in GPU architectures and software ecosystems (CUDA, ROCm, etc.).

Responsibilities

  • Lead benchmarking and performance optimizations at the inference engine and GPU kernel layers.
  • Engineer solutions for attention layer optimizations and memory/precision management.
  • Implement cutting-edge optimization techniques for Gen AI workloads.
  • Collaborate with product and TPMs to translate hardware limits into features.
  • Mentor through code reviews and design discussions.

Skills

GPU optimization
Gen AI knowledge
CUDA/Triton
Parallel processing
Open source
System design
Leadership by influence
Low-level GPU programming

Tools

CUDA
ROCm
TensorRT
Triton

Job description

DigitalOcean is seeking a Senior Engineer 2 to drive AI Inference Optimization. You will lead benchmarking and optimizations at the inference engine and GPU kernel levels, and guide the technical roadmap for high-performance inference fleets.

Bring 5+ years of HPC or AI infra experience, deep Gen AI knowledge, and hands-on CUDA/Triton expertise to push throughput and reduce latency across multi-node GPU clusters.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,000 - 239,000
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,000 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,000 - 239,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Engineer 2: GPU Kernel and Performance
Senior Engineer 2: GPU Kernel and Performance

DigitalOcean • Seattle (WA)

On-site
USD 167,000 - 209,000
Flexible time off
Employee Stock Purchase Program
Career development resources
Senior AI Systems Engineer — High-Performance Inference
Senior AI Systems Engineer — High-Performance Inference

DDN • United States

Remote
USD 150,000 - 210,000
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000