Senior Machine Learning Ops (AI Engineering)

Mastercard

Ireland

On-site

EUR 120,000 - 160,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mastercard is seeking a hands-on ML/AI platform engineer in Ireland to design and operate end-to-end ML pipelines, from training to production

including experiment tracking, model registries, and secure release processes. The role emphasizes observability, governance, and cost-aware workload orchestration on Databricks with strong collaboration across engineering, data, and security teams.

Qualifications

  • Experience with ML model deployment pipelines and MLOps practices.
  • Familiarity with cloud security and CI/CD security controls.
  • Experience with Databricks or a comparable platform for orchestration.
  • Ability to design scalable AI/ML data pipelines and governance.

Responsibilities

  • Build and operate pipelines, deployment workflows, and production-readiness practices that turn trained models into reliable, governed services.
  • Own experiment tracking and model registry practices using MLflow (or equivalent).
  • Implement drift and model-performance monitoring to detect data drift and degradation.
  • Design safe model release and rollout with canary/shadow patterns and rollback procedures.
  • Orchestrate training and inference workloads on Databricks and manage recurring jobs.
  • Develop observability: logging, metrics, tracing with SLIs/SLOs for real-time and batch workloads.
  • Set up automated evaluation gates for offline metrics and monitor costs for GPU workloads.
  • Design and build CI/CD pipelines for AI/data workloads and enforce security standards.

Skills

Problem solving
Cross-functional collaboration
Architecture & design
Security best practices

Tools

AWS
Databricks
MLflow
Terraform
Docker
Kubernetes

Job description


  • Responsible for building and operating the pipelines, deployment workflows, and production-readiness practices that turn trained models into reliable, governed services

  • AI Model Lifecycle & Deployment:

  • Own experiment tracking and model registry practices: using MLflow (or equivalent) to manage model versioning and the staging/production/archived lifecycle

  • Implement drift and model-performance monitoring: detecting data drift, embedding/representation drift, and downstream task performance degradation

  • Implement safe model release and rollout mechanisms: including canary or shadow deployment patterns for new model versions, version-gated promotion criteria, and rollback procedures, so downstream consumers are never broken by an untested release

  • Orchestrate training and inference workloads on Databricks: configuring and maintaining Databricks Workflows/Jobs for recurring training cycles and on-demand inference/embedding generation

  • Monitoring & Governance:

  • Design and implement observability for AI/ML services: logging, metrics, and distributed tracing across both real-time and batch workloads, with SLIs/SLOs appropriate to each

  • Set up automated evaluation gates for offline metrics and model performance degradation

  • Track cost and resource utilization for compute-intensive workloads: particularly GPU-based training and inference, flagging inefficiencies or budget risk

  • Pipeline & Infrastructure Development:

  • Design and build CI/CD pipelines for AI and data workloads: supporting model training, evaluation, and deployment, and recommending which tools and patterns to use within the organization’s existing supporting technology

  • Embed security best practices into every pipeline: secrets management, least-privilege access control, and secure configuration, integrating correctly with existing organizational identity and security standards rather than defining new ones

  • Onboard platform services onto centrally-owned infrastructure: such as API gateways and cross-environment data pipelines: meeting their existing security and integration requirements

  • Support incident response and post-incident improvement: contributing to troubleshooting production issues and helping drive follow-up actions after incidents

  • Abide by Mastercard’s security policies and practices

  • Ensure the confidentiality and integrity of the information being accessed

  • Report any suspected information security violation or breach

  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines


Working knowledge of monitoring and observability practices: logging, metrics, tracing, and how they apply differently to latency-sensitive versus batch AI workloadsStrong understanding of software delivery practices: version control, automated testing, and release disciplineStrong problem-solving skills and comfort owning technical design decisions, working effectively across engineering, data, and AI teams without requiring extensive oversightHands-on experience with cloud platforms, particularly AWS, as a consumer of managed services rather than an infrastructure architect. Experience with Azure or GCP also valuableExperience with infrastructure-as-code tools (e.g., Terraform) sufficient to provision and configure resources within an existing account/platform structureExperience with Databricks or a similar unified data/AI platform: job orchestration, workflow scheduling, and integration with governed data pipelines. Strong plus if not already presentFamiliarity with containerization (Docker; Kubernetes exposure a plus), particularly for packaging and deploying model-serving workloadsExperience supporting AI/ML workloads specifically: model deployment pipelines, batch or streaming inference, and the operational differences between training and serving workloadsStrong, hands-on experience building and maintaining CI/CD pipelines in production environments, including the judgment to recommend appropriate tools and patterns rather than simply operating an existing pipelineFamiliarity with security best practices in cloud and CI/CD environments: secrets management, IAM, least-privilege access: with the ability to implement these correctly within an existing security frameworkExperience with MLOps-specific tooling and practices: experiment tracking, model registries, and safe model deployment/rollout patterns (e.g., MLflow or equivalent)

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Machine Learning Ops - AI Engineering
Senior Machine Learning Ops - AI Engineering

Mastercard • Blackrock

On-site
EUR 110,000 - 160,000
Senior Machine Learning Ops - AI Engineering
Senior Machine Learning Ops - AI Engineering

Mastercard • Blanchardstown

On-site
EUR 110,000 - 150,000
Senior Machine Learning Ops - AI Engineering
Senior Machine Learning Ops - AI Engineering

Engg • Dublin

On-site
EUR 120,000 - 150,000
Senior Machine Learning Ops - AI Engineering
Senior Machine Learning Ops - AI Engineering

Mastercard • Dunboyne

On-site
EUR 120,000 - 150,000
Senior Machine Learning Ops - AI Engineering
Senior Machine Learning Ops - AI Engineering

Mastercard • Donabate

On-site
EUR 100,000 - 180,000
Senior Machine Learning Ops - AI Engineering
Senior Machine Learning Ops - AI Engineering

Mastercard • Malahide

On-site
EUR 110,000 - 150,000
Senior Machine Learning Ops - AI Engineering
Senior Machine Learning Ops - AI Engineering

Mastercard • Rathcoole

On-site
EUR 120,000 - 180,000
Lead Analytics Engineer/Data Analyst – AI & Foundation Models
Lead Analytics Engineer/Data Analyst – AI & Foundation Models

Mastercard Inc. • Dublin

On-site
EUR 90,000 - 120,000
Senior ML Ops Engineer: AI Deployment & Governance
Senior ML Ops Engineer: AI Deployment & Governance

Mastercard • Blanchardstown

On-site
EUR 110,000 - 150,000
Lead Site Reliability Engineer (AI/ML)
Lead Site Reliability Engineer (AI/ML)

Mastercard • Malahide

On-site
EUR 120,000 - 160,000