Lead MLOps Engineer

EPAM Systems

Colombia

On-site

COP 388,375,000 - 582,562,000

Full time

20 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off
Upskilling programs
Global career opportunities

Job summary

EPAM Systems in Colombia seeks a Lead MLOps Engineer to head an MVP engagement with a major AAA game publisher, building a test intelligence platform with two franchises. You will set technical direction for ML and signal components, ensuring production-readiness and replayable configurations from day one.

You will mentor engineers, define standards, and serve as the primary MLOps authority across franchises, driving model lineage, calibration harnesses, and monitoring for scalable, auditable

Qualifications

  • 5+ years of experience in MLOps or ML platform engineering.
  • Deep expertise in ML model lifecycle management, including versioning and rollback.
  • Strong background in architecting signal computation pipelines and provenance capture.
  • Advanced calibration methodology and configuration sweep design.
  • Expert knowledge of as-of temporal data systems or back-test harness design.
  • Advanced skills in Python and SQL for ML pipeline automation.
  • Deep competency in model monitoring and drift detection.
  • Proven ability to lead cross-functional collaboration with data engineers and AI developers.
  • Excellent written and verbal English (B2+).

Responsibilities

  • Define the MLOps architecture and strategy for the platform across franchises.
  • Own signal catalogue operations end-to-end and provenance standards.
  • Set strategy for model lineage and controlled reindexing when gateway models change.
  • Architect, deliver, and own the Back-test & Calibration Harness.
  • Oversee configuration sweeps and holdout discipline across routes and weights.
  • Lead model generation and experimentation strategy for scoring configurations.
  • Govern model lineage and cross-franchise auditability.
  • Design and oversee MLOps Monitor and Data-Health Monitor.

Skills

MLOps
Python
SQL
Model lifecycle
Versioning
Leadership
Cross-functional
Back-test design

Tools

Snowflake ML
Snowpark
pgvector

Job description

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.

We are seeking a Lead MLOps Engineer to head an MVP engagement with a major AAA game publisher, building a test intelligence platform for two game franchises in parallel. A core design principle is full re-derivability and model lineage from day one — every run must be replayable from its stored configuration version and feed read positions. The signal catalog feeds a scoring strategy engine with versioned configurations, and calibration sweeps over historical data produce suggested weight updates surfaced directly in the Settings View.

This role sets the technical direction for the ML and signal components, ensuring they are production-ready, reproducible, and improvable over time, forming the foundation of the system's long-term value as franchise history accumulates and models are refined.

The Lead MLOps Engineer will define standards, mentor engineers, and act as the primary technical authority for all MLOps practices across both franchise workstreams.

Responsibilities
  • Define the overall MLOps architecture and strategy for the platform, establishing standards for reproducibility, lineage, and model lifecycle management across both franchises
  • Own Signal Catalogue operations end-to-end: architect signal refresh orchestration triggered by feed read-position advances, define grain translation policies between per-test, per-area, and per-run signal families, and establish provenance capture standards across all 8 signals
  • Set the strategy for Semantic Vector Index versioning: partner with the Lead AI Developer on model and dimension stamp conventions; design and govern the controlled reindex path when the enterprise AI gateway model changes
  • Architect, deliver, and own the Back-test & Calibration Harness: as-of temporal filtering across all record families, replay runner, look-ahead spot audit, and configuration sweep runner
  • Establish and enforce holdout patch-set discipline, define configuration sweep methodology over route limits, thresholds, weights, and Composition setting; oversee catch-rate vs. scope analysis per candidate configuration and approve winning configurations as suggested-weight proposals into the Settings View
  • Lead Model Generation & Experimentation strategy: define the systematic experimentation framework over scoring strategy configurations, guiding the team on which signal weights and route combinations yield the best catch-rate vs. scope trade-off
  • Govern model lineage across configuration versions for both franchises, ensuring cross-franchise consistency and auditability
  • Design and oversee the MLOps Monitor and Data-Health Monitor: catch-rate floor monitoring, run-behaviour drift counters (per-run candidate volumes per route, score distribution vs. usual range), data-health telemetry across all ingestion channels
  • Define exploration cadence policy: unbiased random-sample injection with provenance ensuring exploration entries are never counted as model recommendations
  • Author and own operator runbook sections covering signal refresh, calibration campaigns, model generation runs, and embedding reindex procedures
  • Mentor engineers on MLOps best practices, review technical designs, and represent the MLOps function in cross-team architecture discussions with data engineering, AI development, and product stakeholders
Requirements
  • 5+ years of experience in MLOps or ML platform engineering, with a proven track record of leading technical initiatives end-to-end
  • Deep expertise in ML model lifecycle management, including versioning, configuration management, and rollback, with experience defining organization-wide standards
  • Strong background in architecting signal computation pipelines, covering scheduled refresh, provenance capture, and grain translation
  • Advanced proficiency in calibration methodology: holdout discipline, configuration sweep design, and catch-rate vs. scope measurement
  • Expert knowledge of as-of temporal data systems or back-test harness design and operation
  • Advanced skills in Python and SQL for ML pipeline automation, with experience setting coding and design standards for a team
  • Deep competency in model monitoring, including drift detection, catch-rate floor monitoring, and run-behavior drift counters
  • Proven ability to lead cross-functional collaboration with data engineers and AI developers on feature alignment and shared roadmaps
  • Strong track record of authoring calibration procedures, signal definitions, and operational runbooks that scale across teams
  • Solid experience with Spec Driven Development, ideally as a methodology advocate or champion
  • Demonstrated mentorship and technical leadership experience, including code review, design review, and guiding mid-to-senior engineers
  • Excellent written and verbal communication skills in English (B2+ level)
Nice to have
  • Familiarity with Snowflake ML or Snowpark in production settings
  • Proven showcase of embedding model versioning and controlled reindex orchestration at scale
  • Operational leadership with pgvector or vector store management
  • Background in gaming domain or QA tooling
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer
Senior MLOps Engineer

EPAM Systems • Colombia

On-site
COP 120,000,000 - 210,000,000
Healthcare benefits
Paid time off
Upskilling and certifications
+2
Lead MLOps Engineer
Lead MLOps Engineer

EPAM Systems • Colombia

Remote
COP 287,954,000 - 415,933,000
Healthcare benefits
Paid time off
Learning opportunities
Senior MLOps Engineer
Senior MLOps Engineer

EPAM Systems • Colombia

Remote
COP 144,000,000 - 216,000,000
Healthcare benefits
Paid time off
Upskilling and certifications
+2
Lead GenAI Engineer
Lead GenAI Engineer

EPAM Systems • Colombia

On-site
COP 120,000,000 - 190,000,000
Healthcare benefits
Paid time off
Upskilling courses
+2
Lead Platform Engineering
Lead Platform Engineering

EPAM Systems • Colombia

On-site
COP 120,000,000 - 200,000,000
Healthcare benefits
Paid time off
Upskilling programs
+1
Lead MLOps Architect: End-to-End ML Platform
Lead MLOps Architect: End-to-End ML Platform

EPAM Systems • Colombia

On-site
COP 388,375,000 - 582,562,000
Healthcare benefits
Paid time off
Upskilling programs
+1
Senior/Lead AI Engineer
Senior/Lead AI Engineer

EPAM Systems • Colombia

On-site
COP 120,000,000 - 180,000,000
Learning culture
Health coverage
Visual health budget
+6
Lead Operational Intelligence Engineer
Lead Operational Intelligence Engineer

EPAM Systems • Colombia

On-site
COP 120,000,000 - 210,000,000
Learning culture
Health coverage
Medical leave coverage
+2
Senior Data Engineer
Senior Data Engineer

EPAM Systems • Colombia

On-site
COP 180,000,000 - 240,000,000
Healthcare benefits
Paid time off
LinkedIn Learning access
+3
Senior MLOps Engineer for Game AI Platform
Senior MLOps Engineer for Game AI Platform

EPAM Systems • Colombia

On-site
COP 120,000,000 - 210,000,000
Healthcare benefits
Paid time off
Upskilling and certifications
+2