Overview
We are looking for a skilledLLMOpsEngineerto help build,operate, andoptimizethe infrastructure and pipelines that power our predictive and generative AI capabilities. TheLLMOpsEngineer will work closely with AI Product Engineers, Data Scientists, and seniorLLMOpsstaff to support the deployment, monitoring, and continuous improvement of LLM-based systems. This role is highly technical and hands-on, with opportunities to grow into senior roles through increasing independence, system ownership, and architectural contributions.
Responsibilities
- Implement andmaintaincore LLM pipelines, including model hosting endpoints, embedding pipelines, retrieval layers, and vector database integrations.
- Build automation for model versioning, dataset management, and configuration tracking to support reproducible and reliable GenAI development.
- Develop monitoring and observability components for LLM-based applications, includingmetricsdashboards, alerting, and logging of model outputs and performance.
- Collaborate with AI Product Engineers to move prototypes into production environments, supporting scalability, testing, and integration with mission systems.
- Assistin configuring andoptimizinginference runtimes, containers, and microservices to ensure responsive and cost-efficient model operations.
- Contribute to CI/CD pipelines that support automated testing,evaluationworkflows, and safe deployment of updated prompts, models, and context strategies.
- Help implement andmaintainbasicMLSecOpsguardrails, such as input validation, prompt protections, and output filtering.
- Participate in troubleshooting efforts, root-cause analysis, and issue resolution for operational incidents involving LLM pipelines or infrastructure.
- Stay current with emerging tools, libraries, and practices for operationalizing LLMs, including orchestration frameworks and inference accelerators.
- Document system workflows, runbooks, and operational best practices to support teamknowledgegrowth and onboarding.
- You will contribute to the growth of our AI & Data Exploitation Practice!
Qualifications
- Ability to hold aposition of public trustwith the U.S. government.
- Bachelor’s or Master’s degree in Computer Science, Data Engineering, Machine Learning, or a related field and 5+ years of experience; OR
- Master’s degree in Computer Science, Data Engineering, Machine Learning, or a related field and 3+ years of experience.
- 2+ yearsof experience in software engineering, DevOps,MLOps, cloud engineering, or data engineering, with exposure to LLM or ML model operations.
- Proficiencyin Python and familiarity with LLM-related tools and frameworks such asHugging Face Transformers,LangChain,LlamaIndex, or similar.
- Experience withcontainerization (Docker)and basic orchestration usingKubernetesor serverless model hosting environments.
- Hands-on knowledge of cloud platforms (AWS, Azure, or GCP), includingcompute, storage, and networking fundamentals for AI workloads.
- Familiarity withCI/CD pipelines, automated testing, and environment provisioning for AI or data systems.
- Exposure tovector databases, embedding models,and RAG pattern implementations preferred.
- Understanding ofmodernDevSecOpsprinciples, security basics, and safe handling of data used in AI pipelines.
- Strong debugging and problem-solving skills withanability to collaborate in cross-functional engineering teams.
- Strong written and verbal communication skills with the ability to document workflows and explain operational concepts.
- Experience working in agile or iterative development environments is a plus.