Senior Software Engineer - ML/LLM Serving

Alldus

San Jose (CA)

On-site

USD 180,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A tech company in AI/ML is seeking a Senior Software Engineer specializing in ML Serving to build robust infrastructure for ML models. The ideal candidate has 5+ years of experience in software engineering, with a focus on ML serving. Proficiency in Python and knowledge of various serving frameworks are essential. This full-time role is located in San Jose, California and offers a competitive salary.

Qualifications

  • 5+ years of software engineering experience with an ML serving focus.
  • Proven experience deploying large language models in production.
  • Strong programming skills in Python, Go, or C++.

Responsibilities

  • Design and build ML serving infrastructure.
  • Optimize inference pipelines for efficiency.
  • Integrate models into varied customer environments.

Skills

ML serving
Python
Cloud platforms (AWS, GCP, Azure)
Distributed systems
Container orchestration (Kubernetes, Docker)

Tools

TensorFlow Serving
TorchServe
Triton Inference Server
BentoML
Ray Serve

Job description

Senior Software Engineer - ML/LLM Serving

This range is provided by Alldus. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.

Base pay range

$180,000.00/yr - $220,000.00/yr

Direct message the job poster from Alldus

About the Role

We are seeking a senior/staff Machine Learning Serving Software Engineer who thrives in a fast-paced, customer-focused environment and can build robust, flexible infrastructure to serve a diverse range of ML models including both LLMs and classical ML.

The Role

As an ML Serving Engineer, you will design, implement, and optimize infrastructure that powers the deployment and inference of machine learning models across varied customer environments. You’ll work closely with product, research, and customer engineering teams to deliver low-latency, secure, and scalable ML serving solutions.

Responsibilities
  • Design and build scalable, high-performance ML serving infrastructure capable of handling diverse model types (LLMs, recommendation systems, etc.).
  • Optimize inference pipelines for latency, throughput, and cost efficiency.
  • Integrate with a wide range of customer environments, adapting serving strategies to fit their infrastructure and compliance needs.
  • Deploy, monitor, and maintain ML models in production using modern deployment stacks.
  • Collaborate with ML researchers to operationalize new models and ensure seamless integration into customer workflows.
  • Ensure security and privacy best practices are applied to model deployment and inference, aligning with enterprise-grade data security requirements.
  • Stay up-to-date with the latest serving technologies and frameworks, evaluating and integrating them where relevant.
Qualifications

Required

  • 5+ years of professional software engineering experience, with at least 3+ years focused on ML serving, inference infrastructure, or similar domains.
  • Proven experience deploying and optimizing large language models (LLMs) in production.
  • Hands-on expertise with multiple ML serving frameworks (e.g., TensorFlow Serving, TorchServe, Triton Inference Server, BentoML, Ray Serve, vLLM, etc.).
  • Strong programming skills in Python, Go, or C++.
  • Experience with distributed systems and container orchestration tools (Kubernetes, Docker).
  • Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry) and performance profiling for inference workloads.
  • Solid understanding of secure data handling and privacy-preserving ML practices.
  • Knowledge of cloud platforms (AWS, GCP, Azure) and hybrid/on-prem deployment scenarios.

Preferred

  • Prior experience serving multiple model types beyond LLMs, e.g., recommendation engines and classical ML models.
  • Exposure to model quantization, distillation, caching, and other optimization techniques for inference efficiency.
  • Experience working with enterprise customers or within compliance-heavy environments.
Details
  • Seniority level: Mid-Senior level
  • Employment type: Full-time
  • Job function: Software Development

Referrals increase your chances of interviewing at Alldus by 2x

Related roles

AI/ML Engineer (Multiple roles and seniority levels) – San Jose, CA (examples of similar roles and salary ranges).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineering Manager - LLM Serving (Remote - US)
Machine Learning Engineering Manager - LLM Serving (Remote - US)

Jobgether • United States

Remote
USD 176,000 - 252,000
Competitive salary range
Comprehensive health insurance
Paid parental leave
+4
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Scale AI, Inc. • New York (NY)

On-site
USD 180,000 - 225,000
Comprehensive health coverage
Equity compensation
Learning and development stipend
+2
Sr AI/ML Engineer
Sr AI/ML Engineer

Vizient • Irving (TX)

On-site
USD 102,000 - 179,000
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000
Principal Machine Learning Engineer I
Principal Machine Learning Engineer I

RELX • Raleigh (NC)

On-site
USD 136,000 - 253,000
Country-specific benefits
Machine Learning Engineer Lead
Machine Learning Engineer Lead

RELX • Raleigh (NC)

On-site
USD 115,000 - 192,000
Annual incentive bonus
Country-specific benefits
Senior Software Engineer (AI Systems & Infrastructure)
Senior Software Engineer (AI Systems & Infrastructure)

Bytoa • Columbia (MD)

On-site
USD 235,000 - 255,000
Machine Learning Engineer Lead
Machine Learning Engineer Lead

LexisNexis • Raleigh (NC)

On-site
USD 115,000 - 193,000
Annual incentive bonus
Country-specific benefits
AI/Machine Learning Engineer
AI/Machine Learning Engineer

veritone • Irvine (CA)

On-site
USD 175,000 - 200,000
Principal Machine Learning Engineer I
Principal Machine Learning Engineer I

LexisNexis • Raleigh (NC)

On-site
USD 136,000 - 253,000