Backend ML Engineer

Sterling

North Sioux City (SD)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Sterling Computers is seeking a Backend ML Engineer to take AI/ML systems from prototype to production, designing inference APIs, building retrieval and orchestration pipelines, and integrating large language models. The role requires 3–5 years in backend or ML engineering, Python (FastAPI/Flask), cloud experience (AWS, GCP, or Azure), and hands-on work with LLMs and vector databases.

Travel up to 25%. Join a collaborative team delivering AI features for government and commercial clients,

Qualifications

  • Bachelor’s degree in Computer Science, Machine Learning, or related field or equivalent practical experience.

Responsibilities

  • Build, test, and maintain production ML services — inference APIs, retrieval pipelines, orchestration layers, evaluation components.
  • Design scalable RESTful and streaming APIs that serve ML model outputs under real-world load.
  • Integrate and tune LLMs, embedding models, and rerankers across hosted and self-hosted options; balance cost, latency, and quality.
  • Build ingestion and chunking pipelines for unstructured data and maintain vector store schemas for multi-tenant retrieval.
  • Implement evaluation harnesses to measure retrieval quality, generation faithfulness, and end-to-end accuracy; close loop from evals to improvements.
  • Containerize and deploy ML workloads with Docker and Kubernetes; manage resources and versioning.
  • Optimize queries, vector search, and caching to reduce latency and cost.
  • CI/CD for ML services and monitoring for system and ML-specific metrics.
  • Collaborate with frontend engineers, ML researchers, and product analysts to ship features.
  • Document backend and ML infrastructure, including model cards and decisions.
  • Travel - up to 25–50%.

Skills

Python (FastAPI/Flask)
PyTorch
Transformers
Sentence Transformers
AWS/GCP/Azure
LLM integration
Vector databases
RAG patterns
Async APIs

Education

Bachelor's in CS/ML

Tools

MLflow
Weights & Biases
Kubeflow
LangChain

Job description

Job Description:

Sterling Computers is a technology company that provides IT solutions to a variety of clients, including the federal government, state and local governments, education, and commercial entities. Sterling's Strategic Technologies Group is responsible for learning and becoming subject matter experts in new and emerging technologies. Our team uses this expertise to broaden the portfolio of products and solutions that the company sells, delivers, and manages. Our engineers work on a range of AI-integrated systems, from production RAG platforms and LLM orchestration layers to digital human solutions and intelligent automation pipelines. We are looking for a Backend ML Engineer who is interested in taking AI/ML systems from prototype to production, designing inference APIs, building retrieval and orchestration pipelines, integrating large language models, and operating ML infrastructure at scale. If you thrive in a collaborative, client-focused environment and enjoy shipping AI features that real users depend on, we'd love to have you on our team.

Required Technical Skills:
  • 3–5 years of experience in backend or ML engineering
  • Strong working knowledge of Python, including FastAPI or Flask
  • Experience with modern ML libraries such as PyTorch, Hugging Face Transformers, and sentence-transformers
  • Proficiency with cloud platforms including AWS, GCP, or Azure
  • Hands-on experience integrating LLMs (OpenAI, Anthropic, Gemini, or open-source models) into production systems
  • Familiarity with vector databases such as Weaviate, pgvector, Pinecone, or similar
  • Experience with retrieval-augmented generation (RAG) patterns
  • Self-motivated with a positive and professional attitude
Required Education/Experience
  • Bachelor’s degree in Computer Science, Machine Learning, or a related field (minimum requirement), or equivalent practical experience
  • Graduate-level coursework or specialization in ML/AI is a plus
  • Relevant cloud certifications are a plus
  • Demonstrated experience shipping ML systems to production is a plus
  • US DoD Clearance preferred or willingness to obtain such
Qualifications:
  • Strong experience building backend services with Python (FastAPI/Flask); comfort working with async APIs and request/response patterns for ML inference workloads.
  • Hands-on experience integrating LLMs and embedding models into production applications, including prompt engineering, context management, and handling rate limits, retries, and streaming responses.
  • Familiarity with RAG architectures: chunking strategies, embedding pipelines, vector search, reranking, and evaluation metrics (Recall@k, MRR, faithfulness, answer relevance).
  • Experience with vector databases (Weaviate, pgvector, Pinecone, Qdrant, or similar) and traditional databases (PostgreSQL, MariaDB) for hybrid retrieval and metadata filtering.
  • Cloud experience (AWS/GCP/Azure) for deploying ML services — including managed inference endpoints, GPU instances, or serverless model hosting.
  • Strong understanding of API authentication, secure handling of model inputs/outputs, and PII/PHI-aware design where applicable.
  • Experience with ML observability: tracking latency, token usage, cost-per-query, retrieval quality, and model drift in production.
  • Background in data pipelines, document ingestion/parsing, or evaluation frameworks (Ragas, TruLens, Docling, custom harnesses) is needed.
  • Familiarity with fine-tuning, LoRA/PEFT, or model distillation is appreciated.
  • Experience with MLOps tooling (MLflow, Weights & Biases, Kubeflow) or LLM orchestration frameworks (LangChain, LlamaIndex, Haystack, or custom orchestrators) is a plus.
Responsibilities:
  • Build, test, and maintain production ML services — inference APIs, retrieval pipelines, orchestration layers, and guardrail/evaluation components.
  • Design scalable RESTful and streaming APIs that serve ML model outputs reliably under real-world load.
  • Integrate and tune LLMs, embedding models, and rerankers; evaluate trade-offs across hosted (Anthropic, OpenAI, Vertex) and self-hosted (HF, vLLM) options on cost, latency, and quality.
  • Build ingestion and chunking pipelines for unstructured data (PDFs, HTML, transcripts) and maintain vector store schemas for multi-tenant or multi-domain retrieval.
  • Implement evaluation harnesses to measure retrieval quality, generation faithfulness, and end-to-end answer correctness; close the loop from evals back into pipeline improvements.
  • Containerize and deploy ML workloads with Docker and Kubernetes; manage GPU/CPU resource allocation and model versioning.
  • Optimize database queries, vector search performance, and caching strategies (including LLM prompt caching) to reduce latency and cost.
  • Implement CI/CD pipelines for ML services and instrument monitoring for both system metrics (latency, error rate) and ML-specific metrics (retrieval quality, hallucination rate, drift)
  • Collaborate with frontend engineers, ML researchers, and product analysts to translate model capabilities into shipped features.
  • Document backend and ML infrastructure, including model cards, evaluation results, and architectural decisions
  • Travel - must be willing to travel 25% and periodically up to 50%.

Sterling Computers Corporation (“Sterling”) is an Equal Opportunity Employer. Qualified applicants will receive consideration for employment without regard to age, race, color, creed, religion, disability, medical condition, economic status or status with regard to public assistance, citizenship status, national or social or ethnic origin, past or present membership in the uniformed services, protected veteran status, sex, pregnancy, marital or civil union or domestic partnership status, family or parental status, sexual orientation, gender expression or identity, family medical history or genetic information, HIV status, political belief, or any other status or characteristic protected by applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend ML Engineer: Scale AI Infra & LLM Pipelines
Backend ML Engineer: Scale AI Infra & LLM Pipelines

Sterling • North Sioux City (SD)

On-site
USD 120,000 - 180,000
Machine Learning Engineer
Machine Learning Engineer

Samson Rose • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Cloudflare • United States

Hybrid
USD 130,000 - 160,000
Equal Opportunity Employer
Diversity and Inclusiveness Initiatives
Reasonable accommodations for applicants with disabilities
ML engineer with Gen AI
ML engineer with Gen AI

MACHINE LEARNING TECHNOLOGIES LLC • Austin (TX)

On-site
USD 83,000 - 165,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Hoplon InfoSec, LLC • Oak Brook (IL), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI/ML Software Developer
AI/ML Software Developer

InterImage, Inc. • Arlington (VA)

On-site
USD 90,000 - 130,000
401K plan
20 days PTO
Healthcare coverage
Senior ML Infrastructure Engineer
Senior ML Infrastructure Engineer

Rebar • New York (NY)

On-site
USD 120,000 - 160,000
Comprehensive medical, dental, and vision coverage
Free lunches and dinners
AI/ML Software Developer
AI/ML Software Developer

InterImage • Arlington (TX)

On-site
USD 90,000 - 120,000
401K with 100% match on the first 7% of pay
20 days PTO per year
Healthcare, dental, vision
+1
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)

Match Group • West Hollywood (CA)

On-site
USD 190,000 - 246,000
ML Ops Engineer
ML Ops Engineer

Bana Solutions • Chantilly (VA)

On-site
USD 150,000 - 210,000