ML Inference Infrastructure Architect

Morph

California (MO)

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Morph is looking for candidates with exceptional skills in infrastructure to manage high-uptime systems focused on GPU loads. The role requires deep knowledge of inference engines and experience with various infrastructures, including load balancers.

Your contributions to projects like vLLM or SGLang will be a valuable asset. Join us in tackling complex challenges in custom inference stacks.

Qualifications

  • Deep experience with inference engines is essential.
  • Prior contributions to vLLM or SGLang are a plus.
  • Experience with traditional infrastructure like load balancers.

Responsibilities

  • Design and manage systems to achieve high uptime.
  • Serve custom inference stacks with irregular GPU loads.
  • Work with various infrastructure types to meet challenging requirements.

Skills

Infrastructure design
GPU load management
Load balancers
Contribution to inference frameworks

Job description

Morph is looking for candidates with exceptional skills in infrastructure to manage high-uptime systems focused on GPU loads. The role requires deep knowledge of inference engines and experience with various infrastructures, including load balancers.

Your contributions to projects like vLLM or SGLang will be a valuable asset. Join us in tackling complex challenges in custom inference stacks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Infrastructure Engineer
Senior Machine Learning Infrastructure Engineer

Morph • California (MO)

On-site
USD 90,000 - 120,000
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: GPU Fleet & Orchestration
ML Infra Engineer: GPU Fleet & Orchestration

Generalist AI • San Mateo (CA), Somerville (MA)

On-site
USD 180,000 - 240,000
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer (LLM inference)
Machine Learning Engineer (LLM inference)

GMI Cloud • Mountain View (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Sail • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1