Senior Model Serving Engineer (Distributed ML Inference)

Fundamental

United States

Remote

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary and equity
Comprehensive health coverage
Parental leave

Job summary

Fundamental is seeking a Model Serving Engineer to turn NEXUS into a reliable production system. You will own the inference stack across environments, optimize performance, and collaborate with researchers to evolve model architectures for production.

The role is deeply technical and Python-heavy, focusing on distributed systems and low-level infrastructure surrounding modern ML workloads. You will work on the end-to-end concurrency, deployment architecture, and performance tooling to improve

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field (or equivalent practical experience)
  • 5+ years of experience in model serving, ML infrastructure, or backend engineering
  • Deep expertise in Python concurrency, including GIL behavior, multi-threading, thread safety, multiprocessing
  • Experience building asynchronous and message-driven systems
  • High-performance, large-scale distributed systems
  • Ability to read and reason about ML model implementations at a computational level, including compute behavior, batching, memory usage, and inference characteristics
  • Experience profiling and optimizing performance across CPU, memory, I/O, and ideally GPU workloads

Responsibilities

  • Optimize Python inference code for performance under real concurrency constraints, including GIL contention, multi-threading, multiprocessing, async execution, and long-running production workloads
  • Work closely with research to understand model internals and support the continuous evolution of the architecture, especially around complex and non-obvious computational behavior under production load
  • Collaborate with research and infrastructure teams to reason about hardware utilization and serving tradeoffs across GPU, CPU, memory, networking, batching, and concurrency
  • Define and evolve the architecture behind our distributed inference and asynchronous execution stack, including orchestration, worker coordination, and end-to-end concurrency patterns
  • Own the Triton serving layer for NEXUS, including how models are packaged, configured, and executed as part of our production inference pipeline
  • Build observability and performance tooling across the serving stack, and use production metrics to drive tuning decisions around latency, throughput, and resource efficiency
  • Solve cross-cutting serving challenges that emerge from deploying the same model across environments with very different scale, isolation, and reliability constraints
  • Evaluate and integrate new inference runtimes, serving strategies, and infrastructure approaches as the model ecosystem evolves

Skills

Python concurrency
Asynchronous systems
Distributed systems
ML model internals
Performance profiling
GPU workloads
Kubernetes
DevOps tooling

Education

Bachelors/Masters in CS or related

Tools

Kubernetes
Triton Inference Server

Job description

Fundamental is seeking a Model Serving Engineer to turn NEXUS into a reliable production system. You will own the inference stack across environments, optimize performance, and collaborate with researchers to evolve model architectures for production.

The role is deeply technical and Python-heavy, focusing on distributed systems and low-level infrastructure surrounding modern ML workloads. You will work on the end-to-end concurrency, deployment architecture, and performance tooling to improve

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Senior ML Infrastructure Lead — Model Serving
Senior ML Infrastructure Lead — Model Serving

Cohere • United States

Hybrid
USD 210,000 - 260,000
Lunch stipend
Health and dental benefits
RRSP matching/401K
+4
Senior Model Serving Engineer – Remote AI Infra
Senior Model Serving Engineer – Remote AI Infra

United States Digital Space LLC • United States

Remote
USD 74,000 - 98,000
Senior ML Infra Engineer - Scalable Model Serving (Remote)
Senior ML Infra Engineer - Scalable Model Serving (Remote)

Etsy, Inc. • New York (NY)

Hybrid
USD 182,000 - 246,000
Equity package
Annual bonus
Comprehensive benefits
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Forward-Deployed ML Engineer: Transform Enterprise AI
Forward-Deployed ML Engineer: Transform Enterprise AI

Fundamental • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive compensation with salary and equity
Comprehensive health coverage
Paid parental leave
+2
Remote Senior ML Systems Engineer - Model Serving
Remote Senior ML Systems Engineer - Model Serving

Atlassian • Northern (KY)

Hybrid
USD 149,000 - 235,000
Benefits and bonuses
Equity
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Software Engineer - Model Serving & Systems
Staff Software Engineer - Model Serving & Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 260,000