Stand out for this role — generate a tailored resume and cover letter in about a minute.
Get past ATS filters
Job summary
Deepstreamtech is seeking an ML Ops Engineer to manage the reliability, scalability, and operational integrity of machine-learning systems in both research and production. In this hands-on role, you will design and maintain data pipelines, implement model monitoring, and work with cross-functional teams to productionize ML models. The ideal candidate should have 3–7 years of experience, strong Python skills, and a solid background in data engineering and ML infrastructure.
Qualifications
3–7 years of experience in ML Ops or Data Engineering.
Strong production Python skills required.
Experience deploying ML models in production.
Responsibilities
Own reliability and operational integrity of ML systems.
Design and operate data pipelines for ML models.
Implement model monitoring and ensure data quality.
Skills
Production Python skills
Machine learning operational integrity
Data engineering
System design reasoning
Cloud infrastructure
Tools
CI/CD pipelines
Orchestration systems
Job description
Requirements
3–7 years of professional experience in ML Ops, Data Engineering, or adjacent backend roles
Strong production Python skills (clean APIs, testing, performance awareness)
Experience deploying and operating ML models in production environments
Solid understanding of:
Model training vs. inference requirements
Batch vs. streaming data pipelines
Failure modes in data-driven systems
Hands‑on experience with at least one modern orchestration or workflow system
Comfort working with cloud infrastructure and containerized workloads
Ability to reason about system design, not just tool usage
(Desirable) Experience operating systems at TB-scale data volumes or higher
(Desirable) Prior ownership of model monitoring, drift detection, or automated retraining
(Desirable) Familiarity with feature stores or online/offline feature consistency problems
(Desirable) Experience supporting multiple models or teams on a shared ML platform
(Desirable) Exposure to regulated or high‑reliability production environments
What the job involves
We’re hiring an ML Ops Engineer to own the reliability, scalability, and operational integrity of our machine‑learning systems in research & production
This role sits at the intersection of data engineering and ML infrastructure: you’ll design and operate data pipelines that feed models, and you’ll build the tooling that trains, deploys, monitors, and retrains them
You’ll work closely with research engineers and product teams, taking models from experimentation to production‑grade systems with clear SLAs, reproducibility guarantees, and observable behaviour
This is not a research role; it is a hands‑on engineering role focused on making ML systems work reliably at scale
Productionizing models: packaging, deployment, versioning, and rollback
Designing CI/CD pipelines for ML (training → validation → deployment)
Implementing model monitoring (data drift, prediction drift, performance decay)
Managing experiment tracking and reproducibility
Building and maintaining batch and near‑real‑time data pipelines
Ensuring data quality, schema evolution, and lineage across systems
Designing datasets and feature pipelines that support both training and inference
Operating pipelines with clear reliability and latency expectations
Defining and meeting availability, latency, and freshness targets for ML services
Debugging production issues across data, infrastructure, and model layers
Improving system robustness through automation and observability
Collaborating with platform and security teams on access, secrets, and compliance
Writing production‑grade Python used in long‑running services and pipelines
Establishing testing, validation, and release practices for ML systems
Making trade‑offs explicit between research flexibility and production stability