Senior Machine Learning Engineer

Acceler8 Talent

San Francisco, Northern (CA, KY)

On-site

USD 170,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Acceler8 Talent is seeking a Member of Technical Staff to join our San Francisco onsite team building a next-gen AI inference platform. You will design and implement production-grade ML inference and model serving systems, optimizing latency and throughput for large-scale workloads.

You'll collaborate with compiler, kernel, networking, and distributed systems engineers to push performance and efficiency, supporting diverse hardware and heterogeneous compute in production environments.

Qualifications

  • Strong software engineering fundamentals.
  • Experience building ML inference or model serving systems in production.
  • Deep understanding of system performance, memory behaviour, and optimisation under production workloads.
  • Experience with inference runtimes such as vLLM, TensorRT-LLM, or custom serving frameworks.
  • Familiarity with batching, scheduling, concurrency, and KV cache management.
  • Experience profiling and optimising latency- and throughput-critical systems.
  • Strong Python and C++ development experience.
  • Comfortable working in a fast-moving, early-stage environment with significant ownership.

Responsibilities

  • Design and build production-grade ML inference and model serving systems
  • Optimise latency, throughput, and resource utilisation across large-scale AI workloads
  • Develop execution strategies around batching, scheduling, concurrency, and runtime optimisation
  • Improve KV cache management, memory efficiency, and model execution behaviour
  • Enable new model architectures and inference techniques to run efficiently in production
  • Partner closely with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance
  • Help shape the architecture of a next-generation AI inference platform powering production workloads at scale

Skills

Software engineering fundamentals
Production ML inference
Python
C++
Inference runtimes

Tools

vLLM
TensorRT-LLM
Custom serving frameworks

Job description

Member of Technical Staff – ML Systems & Inference

San Francisco, CA (Onsite)

I am seeking a Member of Technical Staff to join one of the most exciting AI infrastructure companies building the next generation of inference systems.

As AI models continue to grow in size and complexity, the challenge is no longer simply adding more GPUs, it's about making diverse hardware work together efficiently. This team is building the infrastructure that intelligently executes AI workloads across heterogeneous compute, delivering significant improvements in performance, efficiency, and scalability for production AI applications.

You'll join a small, highly technical engineering team solving some of the hardest problems in AI systems, working across inference runtimes, scheduling, memory management, and distributed infrastructure.

What You'll Do:

  • Design and build production-grade ML inference and model serving systems
  • Optimise latency, throughput, and resource utilisation across large-scale AI workloads
  • Develop execution strategies around batching, scheduling, concurrency, and runtime optimisation
  • Improve KV cache management, memory efficiency, and model execution behaviour
  • Enable new model architectures and inference techniques to run efficiently in production
  • Partner closely with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance
  • Help shape the architecture of a next-generation AI inference platform powering production workloads at scale

What We're Looking For:

  • Strong software engineering fundamentals
  • Experience building ML inference or model serving systems in production
  • Deep understanding of system performance, memory behaviour, and optimisation under production workloads
  • Experience with inference runtimes such as vLLM,TensorRT-LLM, or custom serving frameworks is highly desirable
  • Familiarity with batching, scheduling, concurrency, and KV cache management
  • Experience profiling and optimising latency- and throughput-critical systems
  • Strong Python and C++ development experience
  • Comfortable working in a fast-moving, early-stage environment with significant ownership

This is an opportunity to join an exceptionally well-funded AI infrastructure company with a small, world-class engineering team already supporting production deployments for Fortune 500 and AI-native organisations.

You'll work across compiler systems, GPU kernels, distributed scheduling, inference optimisation, and heterogeneous compute, solving difficult engineering challenges that directly impact how modern AI workloads are executed in production.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Machine Learning Engineer
Machine Learning Engineer

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior/Staff AI Engineer
Senior/Staff AI Engineer

Data Direct Networks • California (MO)

On-site
USD 150,000 - 230,000
Vacation plans
Paid holidays
Bonus programs
+5
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Relocation assistance
Visa sponsorship
Lunch stipend
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Cambridge (MA)

On-site
USD 190,000 - 230,000
Relocation assistance
Visa sponsorship
Daily lunch stipend
+2