Senior MLOps Engineer

MUFG Investor Services

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MUFG Investor Services in London is seeking a highly skilled MLOps / Platform Engineer to join our AI initiative. You will design, deploy, and maintain scalable AI agent platforms, collaborating with Research, Data Science, backend and frontend teams to deliver secure, cost-optimized infrastructure.

The role focuses on observability, security, and high availability using AWS, Terraform, OpenTelemetry and Datadog, with CI/CD automation and IaC practices to enable production deployments of

Qualifications

  • 5+ years in Platform Engineering or DevOps roles.
  • Strong knowledge of DevOps principles, workflows, and best practices.
  • Proven platform engineering and full-stack development experience.
  • API design and integration proficiency.
  • Hands-on AWS experience; observability tooling expertise.
  • Proficiency in Python, Node.js, or Go.
  • Familiarity with microservices and cloud-native architectures.
  • Security, cost optimization and performance testing awareness.
  • Interest in MCPs (Model Context Protocol) and AI agent frameworks.

Responsibilities

  • Design, deploy, and maintain AI agents on core MCP servers and gateways.
  • Implement observability with OpenTelemetry and Datadog.
  • Ensure security, high availability, and cost optimization across AI platform components.
  • Provide infrastructure and deployment support to AI researchers and engineers.
  • Conduct load testing, token cost measurement, and optimize resources.
  • Facilitate vulnerability assessments and security best practices.
  • Troubleshoot platform issues to maintain operational stability.
  • Contribute to DevOps workflows, CI/CD pipelines, and automation.
  • Assist in evaluating third‑party hosting AI agents and related tooling.
  • Develop automation scripts using AWS Boto3; implement IaC with Terraform.

Skills

Platform engineering
DevOps practices
API design
AWS
Kubernetes/EKS
CI/CD pipelines
Programming (Python/Node.js/Go)
Observability tooling
IaC (Terraform)
Security best practices

Tools

Terraform
Kubernetes
Datadog
OpenTelemetry
AWS

Job description

MUFG Investor Services is a trusted partner to many of the world’s largest public and private funds, providing asset servicing and operational solutions built for alternatives. With over $1 trillion in client assets under administration, we offer fund administration, banking, payments, fund financing, foreign exchange overlay, corporate and regulatory services, custody, business consulting, and more. Operating from 17 locations worldwide, we help clients mitigate risk, enhance efficiency, and navigate the operational complexities of today’s investment management landscape. As a division of Mitsubishi UFJ Financial Group (MUFG), one of the world’s largest financial institutions with approximately $3 trillion in assets, we combine deep expertise with the strength and stability of a leading financial institution. To learn more, visit us at www.mufg-investorservices.com.

Job Description

We are seeking a highly skilled MLOps / Platform Engineer with a strong background in DevOps workflows and platform engineering best practices to join our AI initiative. This is a high-visibility project focused on deploying and managing AI agents across our infrastructure. You will work closely with the Research & Data Science team, backend and frontend engineers, and other technical teams to build a secure, scalable, and cost-optimized platform for AI workloads.

This position supports AI Engineering and Data Science initiatives by focusing on infrastructure, operations, and platform reliability. The Platform Engineer will work closely with AI Engineers and Data Scientists to ensure they have robust, scalable infrastructure to deploy their work.

You Will:

  • Design, deploy, and maintain AI agents on Agent Core MCP servers and MCP gateways.
  • Implement and manage observability using OpenTelemetry for logs and traces, integrating with Datadog.
  • Ensure security, high availability, and cost optimization across all AI platform components.
  • Provide infrastructure and deployment support to AI researchers and engineering teams, enabling integration of cutting‑edge technologies into production.
  • Perform load testing, token cost measurement, and optimize resource utilization.
  • Facilitate external vulnerability assessments and ensure compliance with security best practices.
  • Troubleshoot and resolve platform issues promptly to maintain operational stability.
  • Contribute to DevOps workflows, CI/CD pipelines, and automation for AI deployments.
  • Support evaluation of third‑party products related to hosting AI agents or enhancing project capabilities.
  • Assist in external audits and maintain documentation for platform architecture and processes.
  • Develop and execute automation scripts using the AWS Boto3 SDK to deploy, test, and validate AI platform components across multiple environments.
  • Implement Infrastructure as Code (IaC) using Terraform to provision and manage cloud resources for AI workloads, ensuring consistency and scalability.
Qualifications

You Have:

  • 5+ Years of experience in Platform Engineering / DevOps practice
  • Deep understanding of DevOps principles, workflows, and best practices.
  • Proven experience in platform engineering and full‑stack development.
  • Proficiency in API design and integration.
  • Hands‑on experience with AWS services
  • Familiarity with OpenTelemetry, Datadog, and observability tooling.
  • Solid coding skills in languages commonly used for backend and automation (e.g., Python, Node.js, Go).
  • Knowledge of microservices, container orchestration (Kubernetes/EKS), and cloud‑native architectures.
  • Extensive knowledge of security practices, cost optimization, and performance testing.
  • Interest and familiarity with latest trends in MCPs (Model Context Protocol) and AI agent frameworks.
Preferred Experience
  • Working with AI/ML platforms or deploying AI agents in production environments.
  • Exposure to high‑scale distributed systems and cloud infrastructure.
  • Experience in observability and monitoring for complex systems.
  • AWS certifications
  • High visibility within the organization.
  • Opportunity to work with cutting‑edge AI technologies and collaborate with leading experts.

MUFG is an equal opportunity employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer: AI Platform & Observability Lead
Senior MLOps Engineer: AI Platform & Observability Lead

MUFG Investor Services • Greater London

On-site
GBP 90,000 - 130,000
DevOps Engineer
DevOps Engineer

DGH Recruitment • City Of London

On-site
GBP 85,000 - 115,000
Senior ML Platform Engineer - Artificial Intelligence
Senior ML Platform Engineer - Artificial Intelligence

Bloomberg LP • Greater London

On-site
GBP 80,000 - 100,000
AI Platform Engineer
AI Platform Engineer

Tempest Vane Partners • Greater London

Hybrid
GBP 120,000 - 180,000
Competitive compensation
Discretionary bonus
Benefits package
AI Engineer - Platform
AI Engineer - Platform

Crisil • Greater London

On-site
GBP 80,000 - 120,000
AI Engineer, VP
AI Engineer, VP

WeAreTechWomen • Greater London

On-site
GBP 120,000 - 180,000
Senior AI Engineer
Senior AI Engineer

United States Digital Space LLC • Greater London

Hybrid
GBP 70,000 - 120,000
Hybrid working
Home office budget
Profit sharing
+2
Senior AI Platform Engineer - Production AI & MLOps
Senior AI Platform Engineer - Production AI & MLOps

Ubds-Group • Greater London

On-site
GBP 100,000 - 140,000
Training programs
Private medical cover
Discretionary bonus
+4
Senior AI Platform Engineer - Agentic ML & MCP Tools
Senior AI Platform Engineer - Agentic ML & MCP Tools

London Stock Exchange • Greater London

On-site
GBP 90,000 - 130,000
Agentic AI Platform Engineer
Agentic AI Platform Engineer

ON.com • Greater London

On-site
GBP 120,000 - 170,000