Senior MLOps Engineer

EPAM Systems

México

Remote

MXN 600,000 - 1,000,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off
Upskilling and certifications
LinkedIn Learning access
Global career opportunities

Job summary

EPAM Systems, Inc. seeks a Senior MLOps Engineer to join an MVP engagement with a major AAA game publisher. You will build a test intelligence platform, ensure reproducible model lineage, and maintain a calibration harness for historical data and live feeds.

You will lead signal catalog operations, index versioning, and experiment tracking, collaborating with AI developers and data engineers to optimize catch-rate vs. scope across two game franchises.

Qualifications

  • 3+ years of experience in MLOps or ML platform engineering.
  • Expertise in ML model lifecycle management, versioning, and rollback.
  • Background in signal computation pipelines and provenance capture.
  • Proficiency in Python and SQL for ML pipeline automation.
  • Knowledge of as-of temporal data systems or back-test design is a plus.

Responsibilities

  • Own Signal Catalogue operations and provenance across signals.
  • Operate Semantic Vector Index versioning and calibration harness.
  • Design and maintain back-test and calibration workflows.
  • Enforce holdout discipline and configuration sweeps for evaluation.
  • Lead experimentation tracking for signal configurations and performance.

Skills

MLOps
Python
SQL
Model monitoring
Drift detection
Provenance
Calibration
Configuration sweep
Back-test harness
Vector store
pgvector
Snowflake ML
Snowpark
MLflow
Kubeflow

Tools

Snowflake ML
Snowpark
pgvector

Job description

We are seeking a Senior MLOps Engineer to join an MVP engagement with a major AAA game publisher, building a test intelligence platform for two game franchises in parallel. A core design principle is full re-derivability and model lineage from day one — every run must be replayable from its stored configuration version and feed read positions. The signal catalog feeds a scoring strategy engine with versioned configurations, and calibration sweeps over historical data produce suggested weight updates surfaced directly in the Settings View.This role ensures the ML and signal components are production-ready, reproducible, and improvable over time, forming the foundation of the system's long-term value as franchise history accumulates and models are refined.ResponsibilitiesOwn Signal Catalogue operations: signal refresh orchestration triggered by feed read-position advances, grain translation between per-test, per-area, and per-run signal families, and provenance capture across all 8 signalsOperate the Semantic Vector Index versioning: coordinate with the Senior AI Developer on model and dimension stamp conventions; design and execute the controlled reindex path when the enterprise AI gateway model changesDesign, implement, and own the Back-test & Calibration Harness: as-of temporal filtering across all record families, replay runner, look-ahead spot audit, and configuration sweep runnerEnforce holdout patch-set discipline, configuration sweep over route limits, thresholds, weights, and Composition setting; produce catch-rate vs. scope tables per candidate configuration and publish winning configurations as suggested-weight proposals into the Settings ViewLead Model Generation & Experimentation: systematic experimentation framework over scoring strategy configurations, tracking which signal weights and route combinations yield the best catch-rate vs. scope trade-offMaintain model lineage across configuration versions for both franchisesImplement the MLOps Monitor and Data-Health Monitor: catch-rate floor monitoring, run-behaviour drift counters (per-run candidate volumes per route, score distribution vs. usual range), data-health telemetry across all ingestion channelsManage exploration cadence support: unbiased random-sample injection with provenance ensuring exploration entries are never counted as model recommendationsContribute to operator runbook sections covering signal refresh, calibration campaigns, model generation runs, and embedding reindex proceduresRequirements3+ years of experience in MLOps or ML platform engineeringExpertise in ML model lifecycle management, including versioning, configuration management, and rollbackBackground in signal computation pipeline design, covering scheduled refresh, provenance capture, and grain translationProficiency in calibration methodology: holdout discipline, configuration sweep design, and catch-rate vs. scope measurementKnowledge of as-of temporal data systems or back-test harness design and operationSkills in Python and SQL for ML pipeline automationCompetency in model monitoring, including drift detection, catch-rate floor monitoring, and run-behavior drift countersCapability to collaborate with data engineers and AI developers on feature alignmentQualifications in documenting calibration procedures, signal definitions, and operational runbooksFamiliarity with Spec Driven DevelopmentEnglish proficiency at an Upper-Intermediate level (B2) or higherNice to haveUnderstanding of MLflow, Kubeflow, or equivalent experiment tracking platformsFamiliarity with Snowflake ML or SnowparkShowcase of embedding model versioning and controlled reindex orchestrationSkills in pgvector or vector store operational managementBackground in gaming domain or QA toolingWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead MLOps Engineer
Lead MLOps Engineer

EPAM Systems • Mexico

Remote
MXN 1,200,000 - 2,100,000
Healthcare benefits
Paid time off & sick leave
Upskilling & certification courses
+1
Senior MLOps Engineer — Reproducible ML for AAA Games
Senior MLOps Engineer — Reproducible ML for AAA Games

EPAM Systems • Mexico

Remote
MXN 600,000 - 1,000,000
Healthcare benefits
Paid time off
Upskilling and certifications
+2
Lead AI Engineer
Lead AI Engineer

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,500,000
Healthcare benefits
Upskilling & certification courses
Paid time off & sick leave
+1
Senior AI Engineer
Senior AI Engineer

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,300,000
Senior MLOps Engineer: Reproducible ML for AAA Games
Senior MLOps Engineer: Reproducible ML for AAA Games

EPAM Systems • Mexico

On-site
MXN 900,000 - 1,500,000
Healthcare benefits
Global career opportunities
Paid time off and sick leave
+3
Lead MLOps Architect for Game AI Platform
Lead MLOps Architect for Game AI Platform

EPAM Systems • Mexico

Remote
MXN 1,200,000 - 2,100,000
Healthcare benefits
Paid time off & sick leave
Upskilling & certification courses
+1
Senior GenAI Engineer
Senior GenAI Engineer

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,300,000
Lead Python AI Solution Engineer
Lead Python AI Solution Engineer

EPAM Systems • Mexico

Remote
MXN 1,000,000 - 1,600,000
Senior QA Automation Engineer
Senior QA Automation Engineer

EPAM Systems • Mexico

Remote
MXN 550,000 - 900,000
Healthcare benefits
Paid time off
Upskilling and certifications
+2
Data Engineer
Data Engineer

EPAM Systems • Mexico

Remote
MXN 520,000 - 760,000
Healthcare benefits
Paid time off
Learning & development
+1