Senior AI Inference Performance Engineer (Remote)

DigitalOcean

Denver (CO)

On-site

USD 191,200 - 239,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

DigitalOcean is seeking a Senior Engineer 2 to drive performance optimizations for AI inference across GPU architectures, focusing on memory bandwidth, throughput, and low-latency inference at scale.

You will lead benchmarking, optimize attention layers, and champion quantization techniques, while mentoring peers and aligning with product goals to deliver high-performance, developer-friendly AI infrastructure.

Qualifications

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Gen AI literacy with major model families.
  • Hands-on experience with attention-layer optimizations and distributed GPU environments.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems (CUDA, ROCm, etc.).
  • Open source experience: building with and contributing to OSS projects.
  • Systems design skills in low-level GPU programming, optimization, memory access patterns, and parallel execution.
  • Leadership experience as a technical lead driving design and delivery.
  • Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores).
  • Expert-level Triton or CUDA; contributions to Triton compiler or CUDA kernels are a plus.

Responsibilities

  • Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, maximizing TFLOP value.
  • Engineer solutions for complex performance issues, including attention, memory and precision management, with multi-node GPU parallelization.
  • Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of Gen AI.
  • Identify kernel fusion opportunities for GLM-5 kernels across Transformer blocks (FlashAttention, RMS Norm).
  • Tune gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5.
  • Advise on hardware procurement and software integration.
  • Develop state-of-the-art quantization techniques (FP8, INT8, FP4) to double throughput without accuracy loss.
  • Lead by example through high-quality code and design reviews.
  • Collaborate with Product Management to translate hardware limits into shippable features.
  • Contribute to open-source AI communities.

Skills

Performance Architecture
Deep-Dive Optimization
Technological Innovation
Hardware Fluency
Precision Optimization
Technical Mentorship
Strategic Collaboration
Low-Level Mastery
Triton/CUDA

Tools

CUDA
ROCm
TensorRT
OpenAI Triton

Job description

DigitalOcean is seeking a Senior Engineer 2 to drive performance optimizations for AI inference across GPU architectures, focusing on memory bandwidth, throughput, and low-latency inference at scale.

You will lead benchmarking, optimize attention layers, and champion quantization techniques, while mentoring peers and aligning with product goals to deliver high-performance, developer-friendly AI infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Optimization Engineer (Remote)
Senior AI Inference Optimization Engineer (Remote)

DigitalOcean • United States

Remote
USD 191,000 - 239,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,200 - 239,000
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,200 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,200 - 209,000
Competitive salary
Flexible time off policy
Employee Assistance Program
+2
Senior AI Performance Engineer - Remote GPU Systems
Senior AI Performance Engineer - Remote GPU Systems

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 258,750 - 431,250
Equity
Comprehensive benefits
Staff Engineer, Inference Optimizations
Staff Engineer, Inference Optimizations

DigitalOcean • United States

Remote
USD 191,000 - 239,000
Remote AI Inference Benchmark Engineer
Remote AI Inference Benchmark Engineer

Silicon Data • United States

Remote
USD 140,000 - 200,000
Remote AI Performance Engineer - Scale ML Inference
Remote AI Performance Engineer - Scale ML Inference

United States Digital Space LLC • United States

Remote
USD 75,000 - 100,000