Senior AI Inference Optimization Engineer (Remote)

DigitalOcean

United States

Remote

USD 191,000 - 239,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

DigitalOcean is seeking a Senior Engineer 2 in AI Inference Optimization to drive performance across inference engines and GPU kernels. You will lead optimization efforts, collaborate with product teams, and push the frontier of Gen AI efficiencies in a remote setting.

You will influence hardware choices, implement advanced quantization, and mentor engineers while delivering scalable, low-latency inference for cutting-edge large models.

Qualifications

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Gen AI literacy with experience in LLM/VLM/LMM landscape.
  • Hands-on experience with attention-layer optimizations and distributed GPU parallelization.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and software ecosystems.
  • Open source software integration, building, and contributions.

Responsibilities

  • Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers.
  • Engineer solutions for complex performance issues including attention optimizations and memory/precision management.
  • Implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of Gen AI.
  • Advise on hardware procurement and software integration for GPUs.
  • Develop and deploy quantization techniques (FP8, INT8, FP4) to boost throughput.
  • Provide technical mentorship and lead by example in design reviews.
  • Collaborate with Product Management to translate hardware limits into product features.
  • Maintain community leadership in GPU performance optimization.

Skills

High-performance computing
AI infrastructure
GPU architectures
CUDA / ROCm
Open Source
Leadership through influence

Tools

Triton
CUDA
ROCm
TensorRT

Job description

DigitalOcean is seeking a Senior Engineer 2 in AI Inference Optimization to drive performance across inference engines and GPU kernels. You will lead optimization efforts, collaborate with product teams, and push the frontier of Gen AI efficiencies in a remote setting.

You will influence hardware choices, implement advanced quantization, and mentor engineers while delivering scalable, low-latency inference for cutting-edge large models.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,200 - 239,000
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,200 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,200 - 239,000
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,200 - 209,000
Competitive salary
Flexible time off policy
Employee Assistance Program
+2
Senior AI Kernel Engineer — Remote/Hybrid Inference
Senior AI Kernel Engineer — Remote/Hybrid Inference

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Staff Engineer, Inference Optimizations
Staff Engineer, Inference Optimizations

DigitalOcean • United States

Remote
USD 191,000 - 239,000
Senior Engineer 2: GPU Kernel and Performance
Senior Engineer 2: GPU Kernel and Performance

DigitalOcean • San Francisco (CA)

On-site
USD 167,200 - 209,000
Competitive salary
Flexible time off policy
Employee Assistance Program
+2
Senior AI Performance Engineer - Remote GPU Systems
Senior AI Performance Engineer - Remote GPU Systems

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 258,750 - 431,250
Equity
Comprehensive benefits