Lead MLOps Engineer

EPAM Systems

Argentina

On-site

ARS 3,000,000 - 6,000,000

Full time

18 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Global career opportunities
Upskilling and certification

Job summary

EPAM Systems seeks a Lead MLOps Engineer to head an MVP engagement with a major AAA game publisher, building a test intelligence platform for two game franchises in parallel. The role emphasizes full model lineage and replayable runs from stored configurations.

The Lead MLOps Engineer will set technical direction for ML and signal components, ensuring production-readiness, reproducibility, and long-term value as models evolve across franchises. Mentorship and cross-team leadership are key.

Qualifications

  • 5+ years in MLOps or ML platform engineering leadership.
  • Experience with ML model lifecycle management including versioning and configuration management.
  • Experience architecting signal computation pipelines with provenance capture.
  • Advanced Python and SQL for ML automation; setting coding standards.

Responsibilities

  • Define MLOps architecture and standards for reproducibility and lineage across franchises.
  • Own signal catalogue operations end-to-end and establish provenance across signals.
  • Set strategy for vector index versioning and reindex pathways when models update.
  • Architect and own back-test, calibration harness and temporal filtering.
  • Lead calibration procedures and configuration sweeps; review and approve winning configurations.
  • Lead model generation and experimentation strategy across scoring configurations.
  • Govern model lineage across configurations for cross-franchise consistency.
  • Design and oversee MLOps monitors and data-health telemetry.
  • Define exploration cadence with provenance to avoid bias in results.
  • Author operator runbooks for signal refresh and calibration campaigns.
  • Mentor engineers and represent MLOps in cross-team discussions.

Skills

MLOps
Python
SQL
ML lifecycle
Signal pipelines
Model monitoring
Leadership
English (B2+)
Mentorship
Cross-team collaboration

Tools

Snowflake ML
pgvector

Job description

We are seeking a Lead MLOps Engineer to head an MVP engagement with a major AAA game publisher, building a test intelligence platform for two game franchises in parallel. A core design principle is full re-derivability and model lineage from day one — every run must be replayable from its stored configuration version and feed read positions. The signal catalog feeds a scoring strategy engine with versioned configurations, and calibration sweeps over historical data produce suggested weight updates surfaced directly in the Settings View.

This role sets the technical direction for the ML and signal components, ensuring they are production-ready, reproducible, and improvable over time, forming the foundation of the system's long‑term value as franchise history accumulates and models are refined.

The Lead MLOps Engineer will define standards, mentor engineers, and act as the primary technical authority for all MLOps practices across both franchise workstreams.

Responsibilities
  • Define the overall MLOps architecture and strategy for the platform, establishing standards for reproducibility, lineage, and model lifecycle management across both franchises
  • Own Signal Catalogue operations end-to-end: architect signal refresh orchestration triggered by feed read-position advances, define grain translation policies between per-test, per-area, and per-run signal families, and establish provenance capture standards across all 8 signals
  • Set the strategy for Semantic Vector Index versioning: partner with the Lead AI Developer on model and dimension stamp conventions; design and govern the controlled reindex path when the enterprise AI gateway model changes
  • Architect, deliver, and own the Back-test & Calibration Harness: as-of temporal filtering across all record families, replay runner, look‑ahead spot audit, and configuration sweep runner
  • Establish and enforce holdout patch-set discipline, define configuration sweep methodology over route limits, thresholds, weights, and Composition setting; oversee catch‑rate vs. scope analysis per candidate configuration and approve winning configurations as suggested‑weight proposals into the Settings View
  • Lead Model Generation & Experimentation strategy: define the systematic experimentation framework over scoring strategy configurations, guiding the team on which signal weights and route combinations yield the best catch‑rate vs. scope trade‑off
  • Govern model lineage across configuration versions for both franchises, ensuring cross‑franchise consistency and auditability
  • Design and oversee the MLOps Monitor and Data-Health Monitor: catch‑rate floor monitoring, run‑behaviour drift counters (per‑run candidate volumes per route, score distribution vs. usual range), data‑health telemetry across all ingestion channels
  • Define exploration cadence policy: unbiased random‑sample injection with provenance ensuring exploration entries are never counted as model recommendations
  • Author and own operator runbook sections covering signal refresh, calibration campaigns, model generation runs, and embedding reindex procedures
  • Mentor engineers on MLOps best practices, review technical designs, and represent the MLOps function in cross‑team architecture discussions with data engineering, AI development, and product stakeholders
Requirements
  • 5+ years of experience in MLOps or ML platform engineering, with a proven track record of leading technical initiatives end‑to‑end
  • Deep expertise in ML model lifecycle management, including versioning, configuration management, and rollback, with experience defining organization‑wide standards
  • Strong background in architecting signal computation pipelines, covering scheduled refresh, provenance capture, and grain translation
  • Advanced proficiency in calibration methodology: holdout discipline, configuration sweep design, and catch‑rate vs. scope measurement
  • Expert knowledge of as‑of temporal data systems or back‑test harness design and operation
  • Advanced skills in Python and SQL for ML pipeline automation, with experience setting coding and design standards for a team
  • Deep competency in model monitoring, including drift detection, catch‑rate floor monitoring, and run‑behaviour drift counters
  • Proven ability to lead cross‑functional collaboration with data engineers and AI developers on feature alignment and shared roadmaps
  • Strong track record of authoring calibration procedures, signal definitions, and operational runbooks that scale across teams
  • Solid experience with Spec Driven Development, ideally as a methodology advocate or champion
  • Demonstrated mentorship and technical leadership experience, including code review, design review, and guiding mid‑to‑senior engineers
  • Excellent written and verbal communication skills in English (B2+ level)
Nice to have
  • Familiarity with Snowflake ML or Snowpark in production settings
  • Proven showcase of embedding model versioning and controlled reindex orchestration at scale
  • Operational leadership with pgvector or vector store management
  • Background in gaming domain or QA tooling
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award‑winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead MLOps Engineer – Reproducible ML for Game Franchises
Lead MLOps Engineer – Reproducible ML for Game Franchises

EPAM Systems • Argentina

On-site
ARS 3,000,000 - 6,000,000
Healthcare benefits
Global career opportunities
Upskilling and certification
Lead AI Engineer
Lead AI Engineer

EPAM Systems • Argentina

Remote
ARS 136,660,000 - 197,397,000
Healthcare benefits
Paid time off and sick leave
LinkedIn Learning access
+2
Remote Senior Software Engineer — Flask/React
Remote Senior Software Engineer — Flask/React

Bluelight Consulting, Llc • Neuquén

Remote
ARS 1,800,000 - 3,000,000
Long-term contract
High visibility role
Influence MLOps strategy
+2
Data Engineering & Infrastructure Lead ID88014
Data Engineering & Infrastructure Lead ID88014

AgileEngine • Ciudad de Mendoza

On-site
ARS 135,892,000 - 226,487,000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Senior MLOps Engineer - Remote - Latin America
Senior MLOps Engineer - Remote - Latin America

FullStack • Córdoba

On-site
ARS 181,190,000 - 256,685,000
100% remote work
Opportunity with startups and Fortune
Continuing education opportunities
Remote Dataops Engineer: Snowflake, Airflow & Ci/Cd
Remote Dataops Engineer: Snowflake, Airflow & Ci/Cd

Unavailable • Buenos Aires

Hybrid
ARS 30,182,000 - 81,492,000
Flextime
Remote work options
Education budget and professional成长
Python Engineer: Data & Gis Automation
Python Engineer: Data & Gis Automation

Capgemini Engineering • Buenos Aires

Hybrid
ARS 182,771,000 - 243,694,000
Professional growth
Competitive USD compensation
A selection of exciting projects
+1
Lead AI/ML Engineer
Lead AI/ML Engineer

N-iX • Argentina

On-site
ARS 83,351,856 - 111,135,808
Flexible working format
Competitive salary
Education reimbursement
+1
Senior Python Backend Developer / ML Engineer (IR-535)
Senior Python Backend Developer / ML Engineer (IR-535)

Intellectsoft • Argentina

On-site
ARS 1,800,000 - 3,200,000
Awesome projects with an impact
Udemy courses of your choice
Team-building events
+2
Mobile Engineering Manager: Lead Ios & Android Delivery
Mobile Engineering Manager: Lead Ios & Android Delivery

Epam Systems, Inc. • Ciudad de Mendoza

Hybrid
ARS 182,310,000 - 273,465,000
Professional growth
USD-based compensation
Exciting projects
+1