Senior Engineer II: Cloud AI Inference & Scalable Systems

digitalocean98

Seattle (WA)

Hybrid

USD 167,000 - 209,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

DigitalOcean seeks a Senior Engineer II to design and optimize serverless inference infrastructure and APIs for large-scale AI workloads. You will tackle throughput, GPU utilization, and fault tolerance while delivering reliable, production-grade systems.

You’ll collaborate with platform and product teams, raise the bar on software design and incident response, and mentor junior engineers in a fast-paced, growth-focused environment. This hybrid role supports innovation and ownership.

Qualifications

  • 7+ years building and operating multi-tenant platforms or distributed backend systems.
  • Experience operating high-scale distributed services in production.
  • Deep understanding of SRE principles: observability, incident management, reliability engineering, capacity planning, and automation.
  • 1+ years hands-on Go/Golang in production systems.
  • 2+ years of Kubernetes experience.
  • Strong understanding of cloud-native architectures and microservices.

Responsibilities

  • Design and build scalable, multi-tenant services powering AI inference and intelligent routing workloads.
  • Improve platform resiliency through observability, capacity management, automation, and tooling.
  • Partner with platform, GPU infra, and product teams to deliver production-grade systems and highly available APIs.
  • Raise engineering bar with solid software design, incident management, and continuous improvement practices.
  • Make architectural decisions around traffic management, service orchestration, reliability, and scalability.
  • Participate in on-call rotations and lead efforts to reduce operator pain and prevent incidents.
  • Provide technical leadership by breaking down problems and guiding engineers to execution plans.
  • Support growth of junior and mid-level engineers through coaching and pairing.

Skills

Golang
Kubernetes
Observability
Distributed systems
Multi-tenant platforms

Tools

vLLM
TensorRT-LLM
Triton

Job description

DigitalOcean seeks a Senior Engineer II to design and optimize serverless inference infrastructure and APIs for large-scale AI workloads. You will tackle throughput, GPU utilization, and fault tolerance while delivering reliable, production-grade systems.

You’ll collaborate with platform and product teams, raise the bar on software design and incident response, and mentor junior engineers in a fast-paced, growth-focused environment. This hybrid role supports innovation and ownership.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,000 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior Engineer II, Serverless Inference
Senior Engineer II, Serverless Inference

digitalocean98 • Seattle (WA)

Hybrid
USD 167,000 - 209,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,000 - 239,000
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,000 - 239,000
Senior Engineering Manager, AI Inference & Kubernetes
Senior Engineering Manager, AI Inference & Kubernetes

DigitalOcean, LLC • Seattle (WA)

Hybrid
USD 200,000 - 251,000
Equity compensation
Education reimbursement
Flexible time-off policy
Staff Engineer, Inference Optimizations
Staff Engineer, Inference Optimizations

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Senior Engineer 2: GPU Kernel and Performance
Senior Engineer 2: GPU Kernel and Performance

DigitalOcean • San Francisco (CA)

On-site
USD 167,000 - 209,000
Competitive salary
Flexible time off policy
Employee Assistance Program
+2
Senior AI Cloud Platform Engineer - Serverless & GPU
Senior AI Cloud Platform Engineer - Serverless & GPU

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 288,000