Cloud Inference Launch Engineer: Scale & Optimize LLM Infra

United States Digital Space LLC

Washington

On-site

USD 320,000 - 485,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking an engineer to join the Cloud Inference team and scale Claude across AWS, GCP, Azure, and future CSPs. You will own end-to-end inference on each cloud platform, from API integration to deployment and daily operations.

You will focus on fast, cost-efficient validation, performance improvements, and reliability to ensure consistent behavior across providers and accelerate model delivery.

Qualifications

  • Have a strong interest in LLM serving; prior inference or ML experience is not required
  • Have significant software engineering experience, with a strong background in high-performance, large-scale distributed systems serving millions of users
  • Have a track record of building automation or test infrastructure that measurably improved release velocity or reliability
  • Have experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration
  • Thrive in cross-functional collaboration with both internal teams and external partners
  • Are a fast learner who can quickly ramp up on new technologies, hardware platforms and provider ecosystems
  • Are highly autonomous and take ownership of problems end-to-end, including work that falls outside your job description

Responsibilities

  • Be on the critical path for frontier model launches, bringing up inference for new model architectures and shipping them to cloud platforms in lockstep with our first-party platform
  • Work with the core inference team to bring new inference features (e.g. structured sampling, prompt caching, and more) to cloud platforms, owning the platform-specific integration that gets them to production
  • Identify and dive deep on the gaps that make inference behave differently across first-party and CSPs — config drift, observability, deployment patterns, hard cross-platform bugs — and fix them at the source rather than building platform-specific workarounds
  • Design, build, and own the CI/CD infrastructure for the inference server and load balancer across cloud platforms, with shadow traffic, performance baselines (throughput and latency), and correctness checks that catch regressions before production
  • Drive down merge-to-production cycle time by making validation faster, more parallel, and cost-effective enough to run on the same constrained accelerator pool that serves customers, without trading away reliability
  • Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads

Skills

Distributed systems
High-performance software
Automation / test infra
Cloud platforms AWS/GCP/Azure
Kubernetes exposure
Cross-functional collaboration
Fast learner
Autonomous ownership

Education

Bachelor's degree or equivalent

Tools

Kubernetes
Infrastructure as Code
Container orchestration

Job description

United States Digital Space LLC is seeking an engineer to join the Cloud Inference team and scale Claude across AWS, GCP, Azure, and future CSPs. You will own end-to-end inference on each cloud platform, from API integration to deployment and daily operations.

You will focus on fast, cost-efficient validation, performance improvements, and reliability to ensure consistent behavior across providers and accelerate model delivery.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Inference Engineer — Multi-Cloud Scale
Senior Cloud Inference Engineer — Multi-Cloud Scale

United States Digital Space LLC • Washington

On-site
USD 291,000 - 485,000
Staff Cloud Inference Engineer — Launch & Scale
Staff Cloud Inference Engineer — Launch & Scale

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Senior Cloud Backend Engineer — Multi-Cloud Inference
Senior Cloud Backend Engineer — Multi-Cloud Inference

Anthropic • Seattle (WA)

Hybrid
USD 320,000 - 485,000
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

United States Digital Space LLC • Washington

On-site
USD 320,000 - 485,000
Cloud Orchestration Engineer for Scalable ML Inference
Cloud Orchestration Engineer for Scalable ML Inference

Inferact • United States

Remote
USD 140,000 - 190,000
Visa sponsorship
Fully remote
Head of LLM Inference Platform & Architecture
Head of LLM Inference Platform & Architecture

United States Digital Space LLC • San Francisco (CA)

On-site
USD 240,000 - 360,000
Graduate Cloud-Native LLM Inference Engineer
Graduate Cloud-Native LLM Inference Engineer

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Staff + Sr. Software Engineer, Cloud Inference
Staff + Sr. Software Engineer, Cloud Inference

United States Digital Space LLC • Washington

On-site
USD 291,000 - 485,000
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Cloud Hardware Engineer — AI/ML Server Fleet
Cloud Hardware Engineer — AI/ML Server Fleet

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000