Overview
We are seeking aMLOpsEngineerto design, build, and support the infrastructure, tooling, and automation that enable scalable and reliable machine learning systems across our client engagements. This roleis responsible foroperationalizing ML models, implementing robust pipelines, and ensuring smooth transitions from experimentation to production. TheMLOpsEngineer works closely with Data Scientists, AI Developers, Data Engineers, and cloud engineering teams to streamline model deployment, monitoring, and lifecycle management in alignment with mission needs.
Responsibilities
- Develop andmaintainend-to-end ML pipelines, including data ingestion, feature engineering, model training, model packaging, deployment, and monitoring workflows.
- Implement CI/CD pipelines for ML assets, enabling automated testing, versioning, promotion, and reproducibility across environments.
- Integrate ML models into production services using APIs, microservices, serverless functions, or container orchestration frameworks like Kubernetes.
- Build and manage core ML platform components such as model registries, experiment tracking systems, feature stores, datasets, job schedulers, and lineage tools.
- Monitor model performance, system health, and data drift using logging, observability frameworks, dashboards, and alerting systems; partner with Data Scientists to refine retraining strategies.
- Collaborate with Data Engineers to ensure data pipelines and data quality support high-performing ML systems.
- ImplementDevSecOpsbest practices—includingsecretsmanagement, environment hardening, and secure deployment patterns—to ensure compliance and operational resilience.
- Help define and enforceMLOpsstandards, documentation, and reusable patterns that improve efficiency and reduce technical debt across teams.
- Support troubleshooting and root-cause analysis of pipeline issues, infrastructure problems, or performance degradation in deployed ML models.
- Stay current with emergingMLOpstools, cloud-native ML technologies, distributed training methodologies, and best practices in ML lifecycle management.
- You will contribute to the growth of our AI & Data Exploitation Practice!
Qualifications
- Ability to hold a position of public trust with the U.S. government.
- Bachelors or Master’s degree in Computer Science, Data Engineering, Machine Learning, Information Systems, ora relatedtechnical discipline.
- Masters Degree and 0 years of experience OR Bachelors Degree and 2 years of experience OR No degree and 6 years of experience.
- 2+ years of experience inMLOps, ML engineering, DevOps, cloud engineering, or applied ML development.
- Proficiencyin Python and familiarity with ML frameworks such as scikit-learn, TensorFlow,PyTorch, orXGBoost.
- Hands-on experience with at least one cloud platform (AWS, Azure, or GCP) and associated ML/DevOps services (e.g., SageMaker, Azure ML, Vertex AI, EKS/AKS/GKE).
- Practical experience with CI/CD tools (GitHub Actions, GitLab CI, Jenkins) and containerization (Docker, Kubernetes).
- Strong understanding of ML lifecycle management, including versioning, packaging, deployment, monitoring, and retraining.
- Familiarity with infrastructure-as-code tools such as Terraform or CloudFormation.
- Experience with logging, observability, and monitoring frameworks (CloudWatch, Prometheus, Grafana, ELK stack, Datadog, etc.).
- Ability to collaborate with Data Scientists, Engineers, and mission stakeholders to ensure ML systems deliver operational value.
- Strong communicationskills and the ability to document workflows, architecture decisions, and runbooks.
- Preferred certifications:
- AWS ML Specialty
- AWS DevOps Engineer
- Azure Data Scientist Associate
- Google Professional Machine Learning Engineer
- Databricks Machine Learning Associate/Professional