Cloud Orchestration Engineer for Scalable ML Inference

Inferact

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Fully remote

Job summary

Inferact is seeking a cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at scale. You’ll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction.

You’ll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.

Qualifications

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Strong experience with Kubernetes and container orchestration at scale.
  • Experience designing and implementing custom Kubernetes operators.
  • Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).
  • Experience managing GPU clusters and debugging hardware issues.
  • Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.

Responsibilities

  • Design cluster management, deployment automation, and production monitoring for large-scale environments.
  • Ensure deployments are observable, debuggable, and recoverable.
  • Collaborate with teams across the globe to keep vLLM running reliably at scale.
  • Turn operational complexity into reliable infrastructure and streamlined workflows.

Skills

Kubernetes
Container orchestration at scale
Custom Kubernetes operators
Python/Rust/Go
GPU cluster management
Multi-cloud & on-prem infrastructure

Education

Bachelor’s degree or equivalent

Tools

Terraform
Helm

Job description

Inferact is seeking a cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at scale. You’ll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction.

You’ll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Cloud Orchestration (Remote)
Member of Technical Staff, Cloud Orchestration (Remote)

Inferact • United States

Remote
USD 140,000 - 190,000
Visa sponsorship
Fully remote
Remote Kubernetes Cloud Orchestration Engineer
Remote Kubernetes Cloud Orchestration Engineer

Inferact • United States

Remote
USD 140,000 - 210,000
Health coverage
Member of Technical Staff, Site Reliability Engineer
Member of Technical Staff, Site Reliability Engineer

Inferact • United States

Hybrid
USD 200,000 - 400,000
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

On-site
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cloud Inference Launch Engineer: Scale & Optimize LLM Infra
Cloud Inference Launch Engineer: Scale & Optimize LLM Infra

United States Digital Space LLC • Washington

On-site
USD 320,000 - 485,000
Remote MLOps Engineer — Scalable AI Inference
Remote MLOps Engineer — Scalable AI Inference

Bright Vision Technologies • Maple Grove (MN)

Remote
USD 100,000 - 150,000
Backend AI Engineer: Scalable Inference & Orchestration
Backend AI Engineer: Scalable Inference & Orchestration

re-zoo-me • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • United States

Remote
USD 120,000 - 180,000
Visa sponsorship
Health coverage
Fully remote
ML Cloud Infra Engineer: Scale AI Pipelines
ML Cloud Infra Engineer: Scale AI Pipelines

HavocAI • Providence (RI)

On-site
USD 140,000 - 190,000
Health insurance
401k (Matching)
Equity package
+3