Senior LLMOps Engineer

Steampunk

Bloomington (IL)

On-site

USD 145,000 - 185,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Steampunk is seeking an experienced Senior LLMOps Engineer to design, implement, and maintain production-grade LLM and RAG pipelines across enterprise environments. You will build hosting, inference, retrieval, and context management, with a focus on security, reliability, and governance.

The role requires expert Python, hands-on experience with Hugging Face Transformers, vLLM, and vector databases, plus CI/CD for AI workloads.

Qualifications

  • U.S. government security clearance capability is required or achievable.
  • Bachelor’s degree with 10+ years total experience.
  • 5+ years in software/data engineering, MLOps, or cloud engineering with 2+ years on LLM/GenAI ops.
  • Hands-on RAG pipeline design, build and maintenance experience.
  • Fluency in Python and production-grade model deployment experience.

Responsibilities

  • Architect and maintain scalable LLM and RAG pipelines with hosting, inference, and context management.
  • Design secure GenAI infrastructure across cloud environments for reliability and cost efficiency.
  • Build automated evaluation systems for quality, safety, latency and governance compliance.
  • Develop CI/CD workflows for AI applications including dataset versioning and model lineage.
  • Collaborate with AI Product Engineers to productionize prototypes into enterprise-grade systems.
  • Integrate vector databases, model gateways, content filters, and guardrails end-to-end.
  • Implement observability to track performance, hallucinations, costs, and usage patterns.
  • Lead troubleshooting and RCA for deployment and inference issues.
  • Produce Python-based services and integrations to support LLM/RAG pipelines.
  • Stay current with MLSecOps trends and governance practices.

Skills

Security clearance
Software engineering
MLOps
LLM/GenAI operations
RAG pipelines
Python
CI/CD for AI/ML
Observability & monitoring
Technical leadership
Cloud architectures
Docker & Kubernetes
Terraform/CloudFormation
AI governance & safety
Agile project mgmt

Education

Bachelor’s degree

Tools

Hugging Face Transformers
vLLM
TensorRT-LLM
LangChain
LlamaIndex
FAISS
Milvus
Pinecone
FastAPI
PyTorch

Job description

Overview

We are looking for an experienced Senior LLMOps Engineer to design, implement, and maintain production-grade large language model (LLM) pipelines, deployment architectures, and monitoring systems across enterprise environments. The Senior LLMOps Engineer will play a critical role in operationalizing generative AI capabilities, ensuring that LLM-based applications are scalable, secure, reliable, and compliant with emerging AI risk and governance frameworks. This role spans model deployment, orchestration, evaluation, optimization, and Retrieval-Augmented Generation (RAG) pipelines.

Contributions
  • Architect, build, and maintain scalable LLM and Retrieval-Augmented Generation (RAG) pipelines, including model hosting, inference optimization, retrieval layers, and context management frameworks.

  • Lead the design and implementation of secure GenAI infrastructure across cloud environments, ensuring reliability, performance, scalability, and cost efficiency.

  • Build and manage automated evaluation systems that assess LLM output quality, safety, latency, and adherence to AI governance requirements.

  • Develop CI/CD workflows tailored for LLM- and GenAI-based applications, including dataset versioning, model lineage, and automated testing of prompt and model behaviors.

  • Collaborate with AI Product Engineers and Data Scientists to productionize LLM-based prototypes into enterprise-grade, maintainable systems.

  • Integrate vector databases, model gateways, content filters, and guardrail frameworks into end-to-end LLM and RAG solutions.

  • Implement observability and monitoring solutions that track performance metrics, hallucination rates, cost profiles, and user interaction patterns.

  • Lead troubleshooting and root-cause analysis for issues related to LLM deployment, inference performance, retrieval pipelines, or overall pipeline reliability.

  • Develop and maintain Python-based services, integrations, and automation supporting LLM and RAG pipelines.

  • Stay current with emerging LLM architectures, inference optimizations, fine-tuning techniques, and relevant MLSecOps patterns.

  • Ensure compliance with data privacy, ethical AI, security, and AI governance frameworks throughout pipeline design and operations.

  • Mentor junior engineers and contribute to AI engineering best practices, tooling, and reusable infrastructure patterns.

  • You will contribute to the growth of our AI & Data Exploitation Practice!

Qualifications
  • Ability to obtain and maintain a U.S. government security clearance.

  • Bachelor’s degree and 10+ years of total experience.

  • 5+ years of experience in software engineering, data engineering, MLOps, or cloud engineering, with 2+ years focused specifically on LLM or GenAI operations.

  • Hands-on experience designing, building, and maintaining Retrieval-Augmented Generation (RAG) pipelines.

  • Fluency in Python.

  • Strong experience deploying models using frameworks such as Hugging Face Transformers, vLLM, TensorRT-LLM, or similar.

  • Experience with operational tooling and frameworks such as FastAPI, PyTorch, LangChain, LlamaIndex, and vector databases such as FAISS, Milvus, Pinecone, or similar.

  • Advanced knowledge of cloud platforms such as AWS, Azure, or GCP, including model hosting, distributed compute, and secure networking patterns.

  • Hands-on experience building CI/CD pipelines, automated testing frameworks, and environment provisioning for AI/ML workloads.

  • Experience with Docker, Kubernetes, and infrastructure-as-code tools such as Terraform or CloudFormation.

  • Familiarity with MLSecOps, AI governance, model hardening, prompt injection defenses, and content safety monitoring.

  • Strong understanding of logging, observability, and performance profiling for high-throughput LLM inference systems.

  • Excellent written and verbal communication skills, with the ability to explain trade-offs and architectural decisions to technical and non-technical stakeholders.

  • Demonstrated ability to balance long-term platform thinking with hands-on operations and rapid problem solving.

  • Experience working in Agile development environments and using modern project management tools.

About steampunk

Steampunk relies on several factors to determine salary, including but not limited to geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this position is $145,000 to $185,000. The estimate displayed represents a typical annual salary range for this position. Annual salary is just one aspect of Steampunk’s total compensation package for employees. Learn more about additional Steampunk benefits here.

Identity Statement

As part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.

Steampunk is a Change Agent in the Federal contracting industry, bringing new thinking to clients in the Homeland, Federal Civilian, Health and DoD sectors. Through our Human-Centered delivery methodology, we are fundamentally changing the expectations our Federal clients have for true shared accountability in solving their toughest mission challenges.If you want to learn more about our story, visit http://www.steampunk.com.

We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, or any other characteristic protected by law.Steampunk participates in the E-Verify program.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLMOps Engineer
Senior LLMOps Engineer

Steampunk, Inc. • McLean (VA)

On-site
USD 145,000 - 185,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Steampunk • McLean (VA)

On-site
USD 140,000 - 190,000
MLOps Engineer
MLOps Engineer

Steampunk • Bloomington (IL)

On-site
USD 115,000 - 150,000
Data Scientist
Data Scientist

Steampunk • Bloomington (IL)

On-site
USD 125,000 - 160,000
Machine Learning / Federated-Learning Engineer
Machine Learning / Federated-Learning Engineer

Steampunk, Inc. • McLean (VA)

On-site
USD 115,000 - 150,000
Health insurance
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk • McLean (VA)

On-site
USD 105,000 - 145,000
MLOps Engineer
MLOps Engineer

Steampunk, Inc. • McLean (VA)

On-site
USD 115,000 - 150,000
AI Tech Lead
AI Tech Lead

Steampunk • Bloomington (IL)

On-site
USD 140,000 - 180,000
AI Tech Lead
AI Tech Lead

Steampunk • McLean (VA)

On-site
USD 140,000 - 180,000
Program Manager
Program Manager

Steampunk • Bloomington (IL)

Hybrid
USD 140,000 - 200,000