ML Ops Engineer

Bana Solutions

Chantilly (VA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bana Solutions in Chantilly, VA, seeks an ML Ops Engineer to own the machine-learning lifecycle from trained artifacts to live, low-latency serving. You will monitor models, manage automatic retraining, and ensure governance and reproducibility across experiments.

Responsibilities include packaging and optimizing PyTorch/TensorFlow models, deploying via Kubernetes/Docker, and maintaining model lineage with MLflow and Evidently. Experience with Feast and model monitoring is required.

Qualifications

  • End-to-end ML lifecycle ownership from trained artifact to live serving.
  • Experience with MLflow or equivalent for versioning and lineage.
  • Real-time model serving with low latency.
  • Drift monitoring with Evidently or similar.
  • Automated retraining pipelines via Airflow on GPUs.
  • PyTorch or TensorFlow in production; optimization.
  • Feature store usage with Feast.
  • Kubernetes and Docker for deployment.

Responsibilities

  • Own the ML lifecycle from trained artifact to live serving.
  • Manage model release, versioning, and rollback.
  • Serve models for real-time inference with low latency.
  • Monitor models and predictions; detect drift.
  • Automate retraining pipelines with gates and promotions.
  • Ensure reproducibility and governance of experiments.
  • Coordinate feature store usage to prevent training-serving skew.

Skills

ML lifecycle ownership
Python
Kubernetes
Docker
Airflow
Model monitoring
Git

Tools

MLflow
TorchServe
Triton
Seldon
Feast
Prometheus
Grafana
S3/MinIO

Job description

OVERVIEW:

We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible for getting the five detection models from trained artifact to live, low-latency serving, then keeping them healthy -monitored, versioned, and retrained. Your product is the models running well in production, not the data pipeline underneath them.

GENERAL DUTIES:
  • Model release management in MLflow - versioning, aliasing, promotion and rollback, champion/challenger across the five models.
  • Serving models for real-time inference - package and optimize PyTorch models, run them in the low-latency inference workers, hold the <60s SLO; batching, CPU/GPU tradeoffs, inference correctness.
  • Model and prediction monitoring with Evidently - data, concept, and prediction drift; performance decay; alerting - and closing the loop back to retraining.
  • Automated retraining / continuous training - Airflow pipelines that retrain (including GPU training on EKS), validate against gates, and promote new model versions safely.
  • Training/serving consistency - manage the Feast online/offline boundary to prevent training-serving skew.
  • Reproducibility and governance - experiment tracking, model lineage/provenance, and model cards / approval gates for federal AI accountability.
REQUIRED QUALIFICATIONS:
  • Owned the full production ML lifecycle - trained artifact to live serving to be monitored/retrained. Not model-building only, and not data-pipeline-building only.
  • Model registry and experiment tracking - MLflow or equivalent (SageMaker, Weights & Biases, Vertex): versioning, promotion, rollback, lineage.
  • Model serving for real-time/low-latency inference - embedded serving or a model server (TorchServe, Triton, KServe, Seldon, BentoML): model loading, optimization, latency debugging.
  • Model and data drift monitoring - Evidently or equivalent; defining model-quality metrics and acting on decay.
  • Automated retraining / CT pipelines and model CI/CD - validation gates, champion/challenger, shadow or canary rollouts for models.
  • PyTorch (or TensorFlow) in production - packaging, optimizing (ONNX/quantization a plus), serving; debugging inference correctness and latency.
  • Feature store consumption (Feast or equivalent) with real focus on training/serving skew.
  • Kubernetes and Docker to package and deploy model workloads (Helm); Prometheus/Grafana for model and inference metrics.
  • Strong Python and solid software engineering (tests, reproducibility) - not notebook-only.
DESIRED QUALIFICATIONS:
  • The streaming pipeline you serve models into - Kafka + Bytewax (or Flink, Spark Streaming, Kafka Streams). You integrate with it; the data engineer owns it.
  • Apache Airflow used specifically for ML orchestration (training, promotion, drift jobs).
  • GPU training/serving on Kubernetes/EKS (CUDA/NVIDIA images).
  • OpenShift and/or air-gapped model deployment.
  • AWS GovCloud / FedRAMP / FIPS 140-2 / IL4-5, and federal AI governance - model cards, provenance, OSCAL, explainable scoring.
  • Graph ML, autoencoders, and anomaly detection (our detection approach); security/behavioral feature work.
  • Model artifacts in object storage (S3/MinIO); a warehouse (Redshift or equivalent) for offline evaluation data.
CLEARANCE:
  • Active U.S. Citizenship with the eligibility to gain a clearance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops Senior Engineer
ML Ops Senior Engineer

Compunnel, Inc. • California (MO)

On-site
USD 120,000 - 160,000
MLOps Engineer
MLOps Engineer

ERT, Inc. • Arlington (VA)

On-site
USD 120,000 - 150,000
ML Ops Engineer
ML Ops Engineer

TechDigital Group • San Leandro (CA)

On-site
USD 120,000 - 150,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
ML Operations Engineer
ML Operations Engineer

NextGen Healthcare • Georgia

On-site
USD 80,000 - 120,000
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

Aisquared • Washington

Hybrid
USD 120,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

ControlRooms.ai • United States

On-site
USD 120,000 - 170,000
Machine Learning Engineer
Machine Learning Engineer

Evlo AI • Chicago (IL)

On-site
USD 120,000 - 160,000