Backend ML Engineer: Scale AI Infra & LLM Pipelines

Sterling

North Sioux City (SD)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Sterling Computers is seeking a Backend ML Engineer to take AI/ML systems from prototype to production, designing inference APIs, building retrieval and orchestration pipelines, and integrating large language models. The role requires 3–5 years in backend or ML engineering, Python (FastAPI/Flask), cloud experience (AWS, GCP, or Azure), and hands-on work with LLMs and vector databases.

Travel up to 25%. Join a collaborative team delivering AI features for government and commercial clients,

Qualifications

  • Bachelor’s degree in Computer Science, Machine Learning, or related field or equivalent practical experience.

Responsibilities

  • Build, test, and maintain production ML services — inference APIs, retrieval pipelines, orchestration layers, evaluation components.
  • Design scalable RESTful and streaming APIs that serve ML model outputs under real-world load.
  • Integrate and tune LLMs, embedding models, and rerankers across hosted and self-hosted options; balance cost, latency, and quality.
  • Build ingestion and chunking pipelines for unstructured data and maintain vector store schemas for multi-tenant retrieval.
  • Implement evaluation harnesses to measure retrieval quality, generation faithfulness, and end-to-end accuracy; close loop from evals to improvements.
  • Containerize and deploy ML workloads with Docker and Kubernetes; manage resources and versioning.
  • Optimize queries, vector search, and caching to reduce latency and cost.
  • CI/CD for ML services and monitoring for system and ML-specific metrics.
  • Collaborate with frontend engineers, ML researchers, and product analysts to ship features.
  • Document backend and ML infrastructure, including model cards and decisions.
  • Travel - up to 25–50%.

Skills

Python (FastAPI/Flask)
PyTorch
Transformers
Sentence Transformers
AWS/GCP/Azure
LLM integration
Vector databases
RAG patterns
Async APIs

Education

Bachelor's in CS/ML

Tools

MLflow
Weights & Biases
Kubeflow
LangChain

Job description

Sterling Computers is seeking a Backend ML Engineer to take AI/ML systems from prototype to production, designing inference APIs, building retrieval and orchestration pipelines, and integrating large language models. The role requires 3–5 years in backend or ML engineering, Python (FastAPI/Flask), cloud experience (AWS, GCP, or Azure), and hands-on work with LLMs and vector databases.

Travel up to 25%. Join a collaborative team delivering AI features for government and commercial clients,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend ML Engineer
Backend ML Engineer

Sterling • North Sioux City (SD)

On-site
USD 120,000 - 180,000
Senior ML Platform & Infra Engineer - Scale AI Pipelines
Senior ML Platform & Infra Engineer - Scale AI Pipelines

Monograph • United States

Hybrid
USD 160,000 - 240,000
Competitive base pay
Equity (RSUs)
Benefits
AI/ML Infra Engineer: Scale LLM Serving & Pipelines
AI/ML Infra Engineer: Scale LLM Serving & Pipelines

Glean • Mountain View (CA)

Hybrid
USD 175,000 - 270,000
Home office stipend
Education stipend
Wellness stipend
+1
Senior Staff ML Engineer - Scalable LLM Infra
Senior Staff ML Engineer - Scalable LLM Infra

Moveworks • Mountain View (CA), Northern (KY)

Hybrid
USD 190,000 - 280,000
Lead ML Engineer: Scale Production AI Systems
Lead ML Engineer: Scale Production AI Systems

Salt Digital Recruitment • United States

On-site
USD 180,000 - 260,000
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Backend Engineer — Federal AI & Data Pipelines
Backend Engineer — Federal AI & Data Pipelines

Scale AI • Washington, St. Louis (MO), New York (NY), San Francisco (CA)

On-site
USD 162,400 - 225,000
Comprehensive health, dental andVision
Learning & development stipend
Generous PTO
+1
Senior AI Platform Engineer - LLM & MLOps
Senior AI Platform Engineer - LLM & MLOps

Zs Associates • Princeton (NJ)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
AI/ML Engineer
AI/ML Engineer

RiskForce • Northern (KY)

Hybrid
USD 120,000 - 155,000