Senior AI Inference Performance Engineer (Remote)

DigitalOcean

Denver (CO)

On-site

USD 191,200 - 239,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DigitalOcean is seeking a Senior Engineer 2 to drive performance optimizations for AI inference across GPU architectures, focusing on memory bandwidth, throughput, and low-latency inference at scale.

You will lead benchmarking, optimize attention layers, and champion quantization techniques, while mentoring peers and aligning with product goals to deliver high-performance, developer-friendly AI infrastructure.

Qualifications

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Gen AI literacy with major model families.
  • Hands-on experience with attention-layer optimizations and distributed GPU environments.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems (CUDA, ROCm, etc.).
  • Open source experience: building with and contributing to OSS projects.
  • Systems design skills in low-level GPU programming, optimization, memory access patterns, and parallel execution.
  • Leadership experience as a technical lead driving design and delivery.
  • Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores).
  • Expert-level Triton or CUDA; contributions to Triton compiler or CUDA kernels are a plus.

Responsibilities

  • Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, maximizing TFLOP value.
  • Engineer solutions for complex performance issues, including attention, memory and precision management, with multi-node GPU parallelization.
  • Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of Gen AI.
  • Identify kernel fusion opportunities for GLM-5 kernels across Transformer blocks (FlashAttention, RMS Norm).
  • Tune gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5.
  • Advise on hardware procurement and software integration.
  • Develop state-of-the-art quantization techniques (FP8, INT8, FP4) to double throughput without accuracy loss.
  • Lead by example through high-quality code and design reviews.
  • Collaborate with Product Management to translate hardware limits into shippable features.
  • Contribute to open-source AI communities.

Skills

Performance Architecture
Deep-Dive Optimization
Technological Innovation
Hardware Fluency
Precision Optimization
Technical Mentorship
Strategic Collaboration
Low-Level Mastery
Triton/CUDA

Tools

CUDA
ROCm
TensorRT
OpenAI Triton

Job description

DigitalOcean is seeking a Senior Engineer 2 to drive performance optimizations for AI inference across GPU architectures, focusing on memory bandwidth, throughput, and low-latency inference at scale.

You will lead benchmarking, optimize attention layers, and champion quantization techniques, while mentoring peers and aligning with product goals to deliver high-performance, developer-friendly AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,000 - 239,000
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,000 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000
Senior AI Inference Data Plane Engineer
Senior AI Inference Data Plane Engineer

DigitalOcean • Seattle (WA)

Hybrid
USD 139,000 - 174,000
Equity compensation
Hybrid work model
Senior AI Performance Engineer - Remote GPU Systems
Senior AI Performance Engineer - Remote GPU Systems

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 258,000 - 432,000
Equity
Comprehensive benefits
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior Engineer 2: GPU Kernel and Performance
Senior Engineer 2: GPU Kernel and Performance

DigitalOcean • Seattle (WA)

On-site
USD 167,000 - 209,000
Flexible time off
Employee Stock Purchase Program
Career development resources