Senior Model API Engineer – Low-Latency AI Serving

Baseten

United States

Remote

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Baseten is seeking a backend-focused engineer for the Model Performance team to own low-latency APIs powering hosted model endpoints. You’ll work across distributed systems, model serving, and developer tooling to optimize throughput, reliability, and cost efficiency.

You’ll design and deploy benchmarking, profiling, and observability, implement platform fundamentals like authentication and quotas, and collaborate with multiple squads to deliver robust model serving experiences.

Qualifications

  • 3+ years experience building and operating distributed systems or large‑scale APIs.
  • Proven track record of owning low‑latency, reliable backend services (rate‑limiting, auth, quotas, metering).
  • Infra instincts with performance sensibilities: profiling, tracing, capacity planning, and SLO management.
  • Strong written communication; able to produce clear design docs and collaborate across functions.

Responsibilities

  • Design, build, and operate the Model APIs surface with focus on advanced inference capabilities and multi-GPU setups.
  • Productionize performance improvements across runtimes and frameworks.
  • Build benchmarking frameworks measuring real-world performance across model architectures and hardware.
  • Instrument observability (metrics, traces, logs) and develop repeatable benchmarks for speed and reliability.
  • Implement platform fundamentals: API versioning, validation, usage metering, quotas, and authentication.
  • Collaborate with other teams to deliver robust, developer-friendly model serving experiences.

Skills

Distributed systems
Large-scale APIs
Low-latency backend
Performance profiling
GPU performance
Technical writing

Job description

Baseten is seeking a backend-focused engineer for the Model Performance team to own low-latency APIs powering hosted model endpoints. You’ll work across distributed systems, model serving, and developer tooling to optimize throughput, reliability, and cost efficiency.

You’ll design and deploy benchmarking, profiling, and observability, implement platform fundamentals like authentication and quotas, and collaborate with multiple squads to deliver robust model serving experiences.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model API Engineer - High-Performance Inference
Model API Engineer - High-Performance Inference

Jobzhr • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation with equity
100% medical, dental, vision for you +
Flexible PTO including Winter Break
+4
Senior Model Serving Engineer - Low-Latency AI Platform
Senior Model Serving Engineer - Low-Latency AI Platform

Menlo Ventures • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Eligibility for annual performance bonus
Equity opportunities
Model API Engineer - High-Performance Inference
Model API Engineer - High-Performance Inference

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
Senior Backend Platform Engineer — Enterprise AI Systems
Senior Backend Platform Engineer — Enterprise AI Systems

Boson AI • Palo Alto (CA)

On-site
USD 150,000 - 400,000
Software Engineer - Model Products
Software Engineer - Model Products

The Consensus • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Platform Engineer – AI Infra & Low-Latency APIs
Platform Engineer – AI Infra & Low-Latency APIs

Boson AI • Santa Clara (CA)

On-site
USD 170,000 - 250,000
Software Engineer - Model APIs
Software Engineer - Model APIs

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation including equity
100% medical, dental, and vision insurance
Generous PTO including Winter Break
+2
Software Engineer - Model Products
Software Engineer - Model Products

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Medical, dental and vision insurance (
Flexible PTO including Winter Break
+4
Software Engineer - Model Products
Software Engineer - Model Products

Jobzhr • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation with equity
100% medical, dental, vision for you +
Flexible PTO including Winter Break
+4
Remote Model Serving Engineer — High-Performance AI
Remote Model Serving Engineer — High-Performance AI

Bright Vision Technologies • Novi (MI)

On-site
USD 74,000 - 98,000