Senior Lead AI Engineer - LLM & RAG Systems

ZS

Bellevue (WA)

Hybrid

USD 180,000 - 260,000

Full time

34 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Travel opportunities
Total rewards package

Job summary

ZS is seeking an Applied AI Engineer to design, build, and optimize production-grade LLM systems and RAG pipelines. You will work on prompt engineering, embeddings, and tool-augmented LLMs, delivering scalable AI capabilities within a consulting-teams environment.

Responsibilities include developing RESTful services with FastAPI, evaluating models, and improving latency, cost, and reliability. Hybrid work and travel are common in this role.

Qualifications

  • A master's or bachelor's degree in Computer Science or related field from a top university.
  • 4+ years' hands-on experience in Machine Learning (ML) with production LLM systems.
  • Strong fundamentals of ML/DL and fine tuning models (LLM) including transformers, prompt engineering, embeddings and vector search.
  • Experience in backend API design with FastAPI, async patterns, rate limiting.

Responsibilities

  • Design and implement LLM-powered applications using state-of-the-art transformer models.
  • Build and optimize RAG pipelines using embeddings, chunking strategies, and vector search.
  • Experiment with prompt engineering, structured outputs (JSON schemas/function calling), and tool-augmented LLMs (agents/workflows).
  • Fine-tune models using LoRA/PEFT/instruction tuning.
  • Develop and evaluate embedding models for similarity search and semantic retrieval.
  • Conduct LLM evaluation with automated and human-in-the-loop methods (offline + online).
  • Optimize inference workflows for latency, GPU utilization, and cost (quantization, batching, caching).
  • Build and maintain REST API Services (FastAPI etc.) to deploy LLM/RAG endpoints and support scalable inference.
  • Integrate AI systems into production software environments (CI/CD, monitoring, reliability).
  • Research and prototype cutting-edge approaches in Generative AI and share learnings.

Skills

Prompt engineering
Vector search
Python programming
Async programming
FastAPI
ML Ops
NLP/Computer Vision
English fluency

Education

Bachelor's or Master's in CS or related

Tools

Pinecone/Weaviate/Chroma
Langfuse
MLFlow
LLM frameworks
REST APIs

Job description

ZS is seeking an Applied AI Engineer to design, build, and optimize production-grade LLM systems and RAG pipelines. You will work on prompt engineering, embeddings, and tool-augmented LLMs, delivering scalable AI capabilities within a consulting-teams environment.

Responsibilities include developing RESTful services with FastAPI, evaluating models, and improving latency, cost, and reliability. Hybrid work and travel are common in this role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead, AI Engineering — LLM & GenAI Systems
Senior Lead, AI Engineering — LLM & GenAI Systems

Zs Associates • Bellevue (WA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Travel opportunities
Senior AI/ML Engineer: Lead LLM Systems & RAG Pipelines
Senior AI/ML Engineer: Lead LLM Systems & RAG Pipelines

Vizient • Centennial (CO)

On-site
USD 102,000 - 179,000
Senior AI Engineer: LLM & MLOps Architect
Senior AI Engineer: LLM & MLOps Architect

ZS • Princeton (NJ)

Hybrid
USD 150,000 - 210,000
Senior AI Lead Architect: RAG, LLMs & Python
Senior AI Lead Architect: RAG, LLMs & Python

Synechron • Charlotte (NC)

On-site
USD 120,000 - 140,000
Competitive package
Work abroad opportunities
Paid annual leave
+7
AI/ML Engineer — LLM & RAG Architect (Hybrid)
AI/ML Engineer — LLM & RAG Architect (Hybrid)

Pulsora, Inc. • United States

Hybrid
USD 100,000 - 150,000
Equity in company
Comprehensive benefits
Opportunity to work on impactful projects
Remote AI/ML Engineer: LLMs, RAG & GovCloud Systems
Remote AI/ML Engineer: LLMs, RAG & GovCloud Systems

YO AI Labs • Houston (TX)

Remote
USD 140,000 - 190,000
Remote AI/ML Engineer — LLMs, RAG & Secure Cloud
Remote AI/ML Engineer — LLMs, RAG & Secure Cloud

YO AI Labs • Chicago (IL)

Remote
USD 120,000 - 180,000
Production AI Engineer: LLMs, RAG & Observability
Production AI Engineer: LLMs, RAG & Observability

GeniusXLab LLC • United States

On-site
USD 140,000 - 210,000
Learning budget
Remote work
Competitive pay
Senior AI Engineer: LLM & Generative AI Backend
Senior AI Engineer: LLM & Generative AI Backend

TechWize • New York (NY), Northern (KY)

Hybrid
USD 180,000 - 230,000
Senior AI Engineer: RAG & LLM Stack
Senior AI Engineer: RAG & LLM Stack

Fanisko • United States

On-site
USD 120,000 - 160,000