Senior Machine Learning Operations Engineer

BetMGM

Nevada (IA)

On-site

USD 135,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision, Life, and Dis
Disability Insurance
401(k) with company match
Pre‑tax spending accounts including H3
Flexible paid time off
Professional development reimbursement
Employee resource groups
Swag, ticket giveaways, and more!

Job summary

BetMGM is seeking a Senior MLOps Engineer to treat ML systems as software and own the path from training to production endpoints. You will balance latency, cost, and reliability across batch and real‑time inference on AWS SageMaker and Snowflake Cortex pipelines.

The role requires a strong software‑engineering mindset, experience with CI/CD for ML, feature stores, and drift/monitoring. GenAI integration is a plus, and collaboration with data scientists and engineers is essential.

Qualifications

  • BS or MS in Computer Science, Math, Statistics, Machine Learning, or other STEM field; practical experience is valued.
  • 5+ years shipping software in production with Python, Docker, Kubernetes or ECS, CI/CD, and distributed systems debugging.
  • 3+ years operating ML in production with real traffic and defined latency/cost budgets.
  • AWS SageMaker depth (Training, Endpoints, Batch Transform, Model Registry, Pipelines) and supporting services.
  • Snowflake fluency (Snowpark ML, Cortex, dbt‑orchestrated batch scoring).
  • IaC for ML (Terraform + SageMaker Pipelines or equivalent).
  • Experience with feature stores (SageMaker Feature Store, Tecton, Feast) and online/offline parity.
  • Champion/challenger, shadow, and canary deployment patterns as standard platform capabilities.
  • Drift/monitoring tools (Evidently, Arize, SageMaker Model Monitor) integrated with paging.

Responsibilities

  • Stand up and operate BetMGM’s ML platform on AWS and Snowflake with Terraform‑managed infra.
  • Build self‑service scaffolds for end‑to‑end model deployment with CI, drift monitoring, alerting, and connectivity baked in.
  • Design and operate batch scoring pipelines (SageMaker Batch Transform, dbt‑orchestrated scoring) and real‑time inference paths (SageMaker endpoints, Lambda + Bedrock).
  • Own the feature store with online/offline parity; treat training‑serving skew as an incident.
  • Implement CI/CD for ML: model registry, retraining triggers, and lineage from features to deployed models to live predictions.
  • Support drift detection, data quality, and model performance monitoring with paging to humans; own incident response.
  • Integrate GenAI (Bedrock, Anthropic, OpenAI) into production paths when applicable.
  • Collaborate with data engineers, scientists, and partners across BetMGM and external vendors on standards and interfaces.

Skills

Python
Docker
Kubernetes/ECS
CI/CD
Distributed systems debugging
On-call

Education

BS or MS in Computer Science/Math/Statistics/Machine Learning

Tools

SageMaker
Snowflake
Terraform
Snowpark ML
Cortex
Bedrock
Lambda
S3

Job description

Discover What’s Possible At BetMGM

Ready to make your career legendary? Join us as we bring the magic of Vegas to our players. The BetMGM team has over 1,400 talented members, revolutionizing sports betting and online gaming in the United States and Canada. We’re a brand with technology at our hearts and the most driven and focused talent in the business.

Benefits
  • Medical, Dental, Vision, Life, and Disability Insurance
  • 401(k) with company match
  • Pre‑tax spending accounts including health care FSA and commuter savings
  • Flexible paid time off
  • Professional development reimbursement and ongoing skills training opportunities
  • Employee resource groups
  • Swag, ticket giveaways, and more!
About The Role

The Senior MLOps Engineer treats ML systems as software systems and owns the path from a trained model to a production endpoint that meets its latency, cost, and reliability budgets — across both batch scoring (SageMaker Batch Transform, Snowflake Cortex / Snowpark ML, dbt‑orchestrated scoring) and real‑time inference (SageMaker real‑time endpoints, Lambda + Bedrock, sub‑second feature serving). The Senior Engineer builds the platform that data scientists and ML engineers ship on: feature store with guaranteed online/offline parity, model registry, CI/CD for ML, drift and quality monitoring, champion/challenger and shadow deployment scaffolding. This requires a software‑engineering‑first mindset — distributed systems, observability, and on‑call instincts are the foundation; ML literacy makes them effective for this role. GenAI integration experience is a plus, not a requirement.

Responsibilities
ML Production Platform
  • Stand up and operate BetMGM’s ML platform on AWS (SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (Snowpark ML, Cortex), with Terraform‑managed infrastructure.
  • Build self‑service scaffolds that let data scientists ship a model end‑to‑end without a ticket queue — cookie‑cutter project templates with CI, drift monitoring, alerting, IaC, and Snowflake connectivity pre‑baked.
Batch and Real‑Time Inference
  • Design and operate batch scoring pipelines — SageMaker Batch Transform, dbt‑orchestrated scoring against Snowflake, Snowpark ML — with explicit freshness and cost SLAs.
  • Design and operate real‑time inference paths — SageMaker real‑time endpoints, Lambda + Bedrock for GenAI, API Gateway — with stated latency budgets (typically sub‑100ms) and graceful degradation under load.
  • Own the feature store (SageMaker Feature Store, Tecton, or Feast) with guaranteed online/offline parity — training‑serving skew is treated as an incident, not a trade‑off.
CI/CD and Deployment Patterns
  • Build CI/CD for ML — model registry, automated retraining triggers, model versioning, lineage from feature → training run → deployed model → live prediction.
  • Implement champion/challenger, shadow deployments, and canary releases as platform primitives so individual model teams do not reinvent them per project.
Monitoring, Drift & Reliability
  • Stand up drift detection, data quality, and model performance monitoring (Evidently, Arize, or SageMaker Model Monitor — pick one and standardize) with paging that routes to humans who can fix it.
  • Own MLOps incident response — production model failures are SEV events with post‑mortems.
Cost and Performance
  • Right‑size endpoints, batch caching, request batching, and autoscaling. State cost‑per‑prediction targets up front and meet them.
GenAI Integration (Plus, Not Required)
  • Integrate LLM APIs (Bedrock, Anthropic, OpenAI) into production paths — RAG pipelines, agent eval frameworks, prompt versioning, cost and latency observability.
  • Partner with the Helix team on AI personalization workloads as they ramp toward March Madness 2027.
AI in the Engineering Loop
  • Direct AI coding agents (Claude Code, Cursor, GitHub Copilot, dbt Copilot) as a force multiplier across infrastructure code, eval suites, and model‑serving glue — designing work for agents to do, not just accepting their suggestions.
Collaboration
  • Partner with the data engineering team on shared standards (Terraform modules, CI/CD patterns, observability, lineage).
  • Work alongside data scientists and analytics partners to land the right interfaces between research and production — opinionated about the boundary.
  • Coordinate with Entain India and contractor ML partners as workloads consolidate onto the BetMGM‑owned platform.
Qualifications
BS or MS Requirement
  • BS or MS in Computer Science, Math, Statistics, Machine Learning, or other STEM field — or equivalent practical experience. Practical experience wins ties; a PhD is neither required nor a tiebreaker.
Must‑Haves
  • 5+ years shipping software in production — Python, Docker, Kubernetes or ECS, CI/CD, distributed systems debugging — including time on‑call.
  • 3+ years operating ML in production — you have owned a model in prod that served real traffic, with stated latency and cost budgets and a runbook you wrote.
  • AWS depth across the SageMaker surface (Training, Endpoints, Batch Transform, Model Registry, Pipelines) plus the supporting cast (IAM, Lambda, ECS, S3, Secrets Manager, VPC).
  • Snowflake fluency — Snowpark ML, Cortex, dbt‑orchestrated batch scoring, RBAC for ML workloads.
  • IaC for ML — Terraform + SageMaker Pipelines or equivalent. No manual console deployments to production.
  • Feature store experience — SageMaker Feature Store, Tecton, or Feast — with explicit ownership of online/offline parity.
  • Champion/challenger, shadow, and canary deployment patterns as production muscle, not blog‑post familiarity.
  • Drift and model monitoring — Evidently, Arize, WhyLabs, or SageMaker Model Monitor — wired to a paging path.
  • Software‑engineering‑first mindset — you treat ML systems as systems, not notebooks.
Nice‑to‑Haves
  • GenAI in production — Bedrock, Anthropic, or OpenAI APIs integrated into live systems; RAG pipelines; vector DBs (Snowflake Cortex Search, pgvector, Pinecone); evaluation frameworks (Langfuse or in‑house).
  • Snowflake‑native ML — Snowpark Container Services, Cortex AISQL, Cortex Agents — for workloads that do not need to leave the warehouse.
  • Streaming feature engineering — Kafka, Flink, or Snowpipe Streaming — for sub‑second features.
  • Fine‑tuning experience — LoRA, QLoRA, instruction tuning, eval‑driven iteration — with an honest read on when fine‑tuning beats prompting.
  • A track record of shipping more with AI in the engineering loop than without.
  • Regulated‑industry experience (gaming, fintech, healthcare) — comfort with model risk, audit, and lineage requirements.
Salary & Benefits

The annual salary range for this position is $135,000 to $170,000. Factors which may affect starting pay within this range may include geography/market, skills, education, experience and other qualifications of the successful candidate. This position is also eligible for participation in a performance‑based bonus plan.

Legal Authorization

Applicants must possess legal authorization to work for our company in the U.S. without the need for immigration sponsorship. At this time, this role is not eligible for immigration‑related employment authorization sponsorship including H‑1B, O‑1, E‑3, TN, OPT, etc.

Gaming Compliance & Licensing Requirements

As an online gaming company, BetMGM is required to comply with state gaming regulations which includes licensing obligations. Applicable employees must be licensed by at least one jurisdictional agency, although certain positions require licensing by multiple agencies. Failure to become licensed or maintain licensure with each agency as required for the role may result in termination of employment. Please note that the licensing process includes comprehensive background checks which may include a review of criminal records, financial history, and personal background verification.

In addition, candidates must comply with and support BetMGM’s responsible gambling policies, procedures, and initiatives.

About BetMGM

BetMGM is revolutionizing sports betting and online gaming in the United States and Canada. We are a partnership between two powerhouse organizations—MGM Resorts International and Entain Group. You know our name through our exciting portfolio of brands including BetMGM Casino, BetMGM Sportsbook, Borgata Online, Party Casino and Party Poker. We aim to bring our ideas into action and find ways to deliver the best quality in gaming platforms.

Equal Opportunity Employment

BetMGM LLC is an Equal Opportunity Employer. We provide equal employment opportunities to all qualified individuals, regardless of race, religion, gender, gender identity, age, marital status, national origin, sexual orientation, citizenship status, veteran status, disability, or any other legally protected status. As an organization, we are unwavering in our commitment to maintaining a discrimination‑free work environment, and fostering a culture of inclusivity, belonging and equal opportunity for all employees and applicants.

Accessibility

If you need assistance or accommodation with your application due to a disability, you may contact us at recruitment@betmgm.com.

Disclaimer

This job description is not an exclusive or exhaustive list of duties a person in this position may be asked to perform from time to time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Missouri

On-site
USD 135,000 - 170,000
Medical insurance
401(k) with company match
Flexible PTO
+3
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Nebraska

On-site
USD 135,000 - 170,000
Health insurance
401(k) match
Pre‑tax spending accounts
+4
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Tennessee

On-site
USD 135,000 - 170,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • North Carolina

On-site
USD 135,000 - 170,000
Medical Insurance
Dental Insurance
Vision Insurance
+4
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Arizona

On-site
USD 135,000 - 170,000
Medical/Dental/Vision/Life Insurance
401(k) with company match
Pre‑tax FSA and commuter benefits
+4
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Massachusetts

On-site
USD 135,000 - 170,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Connecticut

On-site
USD 135,000 - 170,000
Medical, Dental, Vision, Life
401(k) with company match
FSA and commuter savings
+4
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Town of Montana (WI)

On-site
USD 135,000 - 170,000
Medical insurance
401(k) with company match
Flexible PTO
+2
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • Louisiana (MO)

On-site
USD 135,000 - 170,000
Medical, dental, vision insurance
401(k) with company match
Pre‑tax FSAs & commuter accounts
+4
Senior Machine Learning Operations Engineer
Senior Machine Learning Operations Engineer

BetMGM • New York (NY)

On-site
USD 135,000 - 170,000
Health benefits
401(k) match
Pre‑tax accounts
+4