AI Infrastructure Engineer, Model Serving Platform

Scale AI, Inc.

New York (NY)

On-site

USD 180,000 - 225,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive health coverage
Equity compensation
Learning and development stipend
Generous PTO
Commuter stipend

Job summary

Scale AI, Inc. in New York is looking for a Software Engineer on the ML Infrastructure team to design and build robust platforms for serving LLMs. This role involves collaborating with researchers and applying backend system design principles to launch solutions that enhance innovation.

The ideal candidate will possess experience in developing high-performance systems while ensuring scalability and reliability. Benefits include comprehensive health coverage, equity compensation, and generous PTO.

Qualifications

  • 4+ years of experience building large-scale, high-performance backend systems.
  • Strong programming skills in one or more languages (e.g., Python, Go, Rust, C++).
  • Experience with LLM serving and routing fundamentals.
  • Familiarity with cloud infrastructure (AWS, GCP) and infrastructure as code.

Responsibilities

  • Build and maintain fault-tolerant, high-performance systems for serving LLMs workloads at scale.
  • Collaborate with researchers and engineers to integrate and optimize models.
  • Conduct architecture and design reviews to ensure best practices.
  • Develop monitoring solutions to ensure system health.

Skills

Backend Systems
Python
Go
Rust
C++
Containers
Cloud Infrastructure
Terraform

Tools

Docker
Kubernetes

Job description

As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting‑edge research and production systems, supporting both internal and external use cases across various environments.

The ideal candidate combines strong ML fundamentals with deep expertise in backend system design. You'll work in a highly collaborative environment, bridging research and engineering to deliver seamless experiences to our customers and accelerate innovation across the company.

Responsibilities
  • Build and maintain fault‑tolerant, high‑performance systems for serving LLMs workloads at scale.
  • Build an internal platform to empower LLM capability discovery.
  • Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.
  • Conduct architecture and design reviews to uphold best practices in system design and scalability.
  • Develop monitoring and observability solutions to ensure system health and performance.
  • Lead projects end‑to‑end, from requirements gathering to implementation, in a cross‑functional environment.
Qualifications
  • 4+ years of experience building large‑scale, high‑performance backend systems.
  • Strong programming skills in one or more languages (e.g., Python, Go, Rust, C++).
  • Experience with LLM serving and routing fundamentals (e.g., rate limiting, token streaming, load balancing, budgets).
  • Experience with LLM capabilities and concepts such as reasoning, tool calling, prompt templates.
  • Experience with containers and orchestration tools (e.g., Docker, Kubernetes).
  • Familiarity with cloud infrastructure (AWS, GCP) and infrastructure as code (e.g., Terraform).
  • Proven ability to solve complex problems and work independently in fast‑moving environments.
Nice to haves
  • Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT‑LLM, or text‑generation‑inference.

Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job‑related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend.

For pay transparency purposes, the base salary range for this full‑time position in the locations of San Francisco, New York, Seattle is: $180,000 — $225,000 USD.

We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.

We comply with the United States Department of Labor's Pay Transparency provision.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer (Model Serving Platform)
Senior AI Infrastructure Engineer (Model Serving Platform)

Scale AI • New York (NY), San Francisco (CA)

On-site
USD 216,200 - 270,250
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+1
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Segment (Twilio) • San Francisco (CA)

On-site
USD 175,000 - 220,000
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+2
ML Research Engineer, ML Systems New York, NY Apply →
ML Research Engineer, ML Systems New York, NY Apply →

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Comprehensive health coverage
Dental and vision coverage
Retirement benefits
+3
ML Research Engineer, ML Systems
ML Research Engineer, ML Systems

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
ML Research Engineer, ML Systems
ML Research Engineer, ML Systems

Scale AI, Inc. • Seattle (WA)

On-site
USD 189,000 - 237,000
Equity-based compensation
Commuter stipend
Health benefits
ML Research Engineer, ML Systems
ML Research Engineer, ML Systems

Scale AI, Inc. • San Francisco (CA)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
Machine Learning Engineering Manager - LLM Serving (Remote - US)
Machine Learning Engineering Manager - LLM Serving (Remote - US)

Jobgether • United States

Remote
USD 176,000 - 252,000
Competitive salary range
Comprehensive health insurance
Paid parental leave
+4
Tech Lead Manager- MLRE, ML Systems
Tech Lead Manager- MLRE, ML Systems

Scale AI • San Francisco (CA)

On-site
USD 275,000 - 350,000
Health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
Tech Lead Manager- MLRE, ML Systems New York, NY Apply →
Tech Lead Manager- MLRE, ML Systems New York, NY Apply →

Scale AI, Inc. • New York (NY)

On-site
USD 264,000 - 331,000
Comprehensive health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Senior AI Infra Engineer — Scalable LLM Serving
Senior AI Infra Engineer — Scalable LLM Serving

Scale AI • San Francisco (CA)

On-site
USD 216,200 - 270,250
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+1