Stand out for this role — generate a tailored resume and cover letter in about a minute.
Careervitablr seeks a seasoned ML Systems Engineer to design, build, and operate production ML pipelines at scale. You will own model serving infra, CI/CD for ML, and feature pipelines, collaborating with ML scientists to productionize research into reliable services.
You will lead design reviews, mentor engineers, and drive reliability, performance, and cost optimizations across cloud-based ML platforms.
You own the path to production and the production system itself turning modeling work into reliable, scalable, observable services that run at client's traffic volumes.
You partner closely with ML Specialists/Data Scientists, but the accountability for latency, uptime, cost, retraining pipelines, and CI/CD for models is yours.
This role suits someone who is as comfortable in distributed systems and MLOps as in ML itself think software engineer who
specializes in ML systems rather than data scientist who can code.
Design, build, and operate production ML pipelines: feature engineering, training pipelines, model serving, and monitoring at client's scale (millions of requests/day).
Build and maintain low-latency, high-availability model serving infrastructure (real-time inference for search /personalization; batch scoring for pricing/forecasting).
Own CI/CD for ML: automated retraining, model versioning, shadow deployments, canary releases, and rollback strategies.
Build robust feature pipelines (batch and streaming) using Spark and Kafka, and maintain feature stores for training /serving consistency.
Instrument models and pipelines with observability: drift detection, data quality checks, latency/throughput SLAs, and alerting.
Collaborate with ML Specialists to productionize research prototypes translating notebook code into tested, maintainable, scalable services.
Optimize training and inference cost/performance (GPU utilization, batching, caching, model compression /quantization where relevant).
Contribute to platform-level decisions: build vs. buy for MLOps tooling, standards for reproducibility, and shared infrastructure across ML teams.
Participate in on-call rotation for production ML services; drive postmortems and reliability improvements.
Mentor mid-level ML/software engineers and lead design reviews for ML platform and serving infrastructure components.
6 to 9 years of software engineering experience, with 4+ years specifically building and operating production ML systems.
Strong software engineering fundamentals: Java, Scala, or Python at a production quality bar (testing, code review, design patterns), plus working knowledge of the other two.
Deep experience with cloud infrastructure (AWS preferred) EC2, EKS/Kubernetes, S3, Lambda, IAM and infrastructure-as-code (Terraform or similar).
Event-Driven Systems: Experience building and operating streaming/event-driven pipelines (Kafka) and distributed batch processing (Spark), including microservices that consume and produce events at scale.
Hands-on experience with model serving frameworks (e.g., TorchServe, TensorFlow Serving, Triton, or custom microservices) and CI/CD for ML (MLflow, SageMaker Pipelines, Airflow, or similar).
Testing & Data Quality: Experience with data validation and testing practices for ML pipelines unit/integration tests for data and models, and contract tests between training and serving to prevent train/serve skew.
Security & Compliance: Experience handling PII and other sensitive data securely within ML pipelines; familiarity with data governance, access controls (IAM policies), and encryption at rest/in transit.
Solid understanding of ML fundamentals (enough to have real technical conversations with data scientists) even if you don't build models from scratch day-to-day.
Technical Leadership: Track record of mentoring engineers and leading design/architecture reviews for distributed or ML-platform systems.
Track record of owning production systems: SLAs, on-call, incident response, capacity planning.
Experience with feature stores (Feast, Tecton, or internal equivalents) and real-time feature computation.
Familiarity with AWS SageMaker, Databricks, or Kubeflow for orchestration.
Experience with GPU infrastructure and inference optimization (quantization, distillation, batching strategies).
Background in search/ranking, personalization, pricing, or fraud systems at marketplace scale.
Experience with A/B testing infrastructure for ML models (shadow traffic, interleaving, canary analysis).
Java/Scala/Python (production services), SQL
Kubernetes (EKS), Airflow, internal EGAP platform services, TorchServe/TensorFlow Serving or gRPC-based custom serving
Kafka, Spark, S3 data lake, Presto/Trino, Databricks
MLflow, SageMaker Pipelines, Terraform, Docker, CI/CD via Jenkins/GitHub Actions
AWS (EC2, EKS, Lambda, S3, IAM)
CloudWatch/Datadog-style metrics, custom drift/data-quality monitors
Jira/Confluence (Atlassian), Git-based workflows