Remote Senior AI Inference Optimization Engineer

DigitalOcean

San Francisco (CA)

On-site

USD 191,200 - 239,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity compensation
Remote work

Job summary

DigitalOcean is seeking a Senior Engineer 2 to drive AI Inference Optimization. You will lead benchmarking and optimizations at the inference engine and GPU kernel levels, and guide the technical roadmap for high-performance inference fleets.

Bring 5+ years of HPC or AI infra experience, deep Gen AI knowledge, and hands-on CUDA/Triton expertise to push throughput and reduce latency across multi-node GPU clusters.

Qualifications

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Experience solving compute utilization and memory bandwidth bottlenecks.
  • Familiarity with Gen AI (LLM, VLM, LMM) landscape and model families.
  • Proficiency in GPU architectures and software ecosystems (CUDA, ROCm, etc.).

Responsibilities

  • Lead benchmarking and performance optimizations at the inference engine and GPU kernel layers.
  • Engineer solutions for attention layer optimizations and memory/precision management.
  • Implement cutting-edge optimization techniques for Gen AI workloads.
  • Collaborate with product and TPMs to translate hardware limits into features.
  • Mentor through code reviews and design discussions.

Skills

GPU optimization
Gen AI knowledge
CUDA/Triton
Parallel processing
Open source
System design
Leadership by influence
Low-level GPU programming

Tools

CUDA
ROCm
TensorRT
Triton

Job description

DigitalOcean is seeking a Senior Engineer 2 to drive AI Inference Optimization. You will lead benchmarking and optimizations at the inference engine and GPU kernel levels, and guide the technical roadmap for high-performance inference fleets.

Bring 5+ years of HPC or AI infra experience, deep Gen AI knowledge, and hands-on CUDA/Triton expertise to push throughput and reduce latency across multi-node GPU clusters.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,200 - 239,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,200 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,200 - 239,000
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Engineer II — Serverless AI Inference (Hybrid)
Senior Engineer II — Serverless AI Inference (Hybrid)

DigitalOcean • Seattle (WA)

Hybrid
USD 167,000 - 209,000
Equity compensation
Bonus potential
Senior Dynamo-Triton AI Inference Engineer (Remote)
Senior Dynamo-Triton AI Inference Engineer (Remote)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 224,000 - 357,000
Equity
Comprehensive benefits
Hybrid work model
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Inference Engineer: AI-Driven GPU Kernel Optimization
Senior Inference Engineer: AI-Driven GPU Kernel Optimization

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000