Senior Manager of SRE

Next Frontier Capital

Glasgow

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

JPMorgan Chase is seeking a Senior Manager - SRE to lead a 3-7 person SRE team and scale the AIML Data Platforms. You will tackle complex cloud data challenges, focusing on Data Lake Tools, and work in an agile environment with cross-functional teams.

You will drive reliability, security, and automation across data platforms, empowering AI capabilities with careful data governance and incident response best practices.

Qualifications

  • Experience with formal training or certification on software engineering concepts.
  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and incident management.
  • Ability to manage highly technical team of Site Reliability Engineers.
  • Experience leading teams in the safe use of enterprise AI capabilities within the work environment, with data sensitivity.
  • Ability to set and reinforce organization-level practices for reviewing AI-assisted recommendations while maintaining security and auditability.
  • Extensive experience with AWS, Databricks, or Snowflake platform administration and engineering support.
  • Experience with monitoring tools, automation frameworks, and CI/CD pipelines.
  • Proficient in Python application program development with automated unit testing.
  • Experience with Terraform development and understanding of Terraform enterprise.
  • Experience in delivering system design, application development, testing, and operational stability.
  • Knowledge of Big Data distributed compute frameworks like Spark, Glue, MapReduce etc.
  • Excellent troubleshooting, analytical, and communication skills.

Responsibilities

  • Leads a team of Site Reliability Engineers and support critical applications 24x7.
  • Implements SRE best practices to ensure reliability, scalability, and performance of data platforms.
  • Identifies opportunities to automate remediation of recurring issues to improve stability.
  • Scale and maintain a managed AWS Databricks platform and provide engineering support for teams.
  • Drives reuse-first adoption of enterprise AI capabilities with human-in-the-loop validation.
  • Leads evaluation sessions with external vendors and internal teams on architecture suitability.
  • Drives continuous improvement in observability, alerting, and capacity planning.
  • Collaborates with data teams to optimize infrastructure and deployment processes.
  • Performs platform design, setup, workspace administration, resource monitoring.
  • Develops secure production code and reviews code written by others.
  • Contributes to team culture of diversity, opportunity, and respect.
  • Develops and maintains incident response procedures and postmortems.

Skills

AWS
Databricks
Snowflake
SRE principles
Python
Terraform
CI/CD
Containerization
Kubernetes
Monitoring
Automation
Data platforms

Education

Formal training on software engineering concepts

Tools

Terraform Enterprise
Docker
Kubernetes
Spark

Job description

We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.

The Chief Data & Analytics Office (CDAO) at JPMorgan Chase is responsible for accelerating the firm’s data and analytics journey. This includes ensuring the quality, integrity, and security of the company’s data, as well as leveraging this data to generate insights and drive decision-making. The CDAO is also responsible for developing and implementing solutions that support the firm’s commercial goals by harnessing artificial intelligence and machine learning technologies to develop new products, improve productivity, and enhance risk management effectively and responsibly.

As a Senior Manager - SRE at JPMorgan Chase within the AIML Data Platforms and Chief Data and Analytics Team, you will lead a team of 3-7 SRE engineers and scale the AI/ML Data Platform, deliver advanced technology products focused on data and analytics. You will tackle complex cloud data platform challenges, especially around Data Lake Tools. In this role you will work in an agile environment, collaborating with cross-functional teams.

Job Responsibilities:
  • Leads a team of Site Reliability Engineers and support critical application 24x7
  • Implements Site Reliability Engineering (SRE) best practices to ensure reliability, scalability, and performance of data platforms.
  • Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems.
  • Scale and maintain a managed AWS Databricks platform, and provides engineering and operational support for the platform to Application/Engineering teams.
  • Drives reuse-first adoption of enterprise-authorized AI capabilities within the work environment to improve reliability operations and customer experience outcomes, with human-in-the-loop validation and appropriate handling of sensitive data.
  • Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture.
  • Drives continuous improvement in system observability, alerting, and capacity planning.
  • Collaborates with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
  • Performs platform design, set-up and configuration, workspace administration, resource monitoring, providing engineering support to Data Engineering teams, Data Science/ML, and Application/Integration teams.
  • Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or break down technical problems.
  • Develops secure high-quality production code, and reviews and debugs code written by others.
  • Adds to team culture of diversity, opportunity, and respect.
  • Develops and maintains incident response procedures, including root cause analysis and postmortem documentation.
Required Qualifications, Capabilities, and Skills:
  • Experience with formal training or certification on software engineering concepts.
  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and incident management.
  • Ability to manage highly technical team of Site Reliability Engineers.
  • Experience leading teams in the safe use of enterprise-authorized AI capabilities within the work environment for reliability engineering workflows, including validation habits and awareness of data sensitivity.
  • Ability to set and reinforce organization-level practices for reviewing AI-assisted recommendations and escalating uncertain decisions while maintaining resiliency, security, and auditability outcomes.
  • Extensive experience with AWS, Databricks, or Snowflake platform administration and engineering support is a MUST.
  • Experience with monitoring tools, automation frameworks, and CI/CD pipelines.
  • Proficient in Python application program development with use of automated unit testing.
  • Experience with Terraform development and understanding of Terraform enterprise.
  • Experience in delivering system design, application development, testing, and operational stability.
  • Knowledge of Big Data distributed compute frameworks like Spark, Glue, MapReduce etc.
  • Excellent troubleshooting, analytical, and communication skills.
Preferred Qualifications, Capabilities, and Skills:
  • Multi Region disaster recovery setup, monitoring and testing.
  • Experience in Data pipelines using Spark.
  • Exposure to AWS & Databricks Platform administration.
  • Knowledge of containerization (Docker, Kubernetes) and orchestration.
  • Familiarity with distributed systems and large-scale data processing.

J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

Influence your team’s strategic planning while driving continual site reliability improvements

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Manager of SRE
Senior Manager of SRE

Fairygodboss • Glasgow

On-site
GBP 120,000 - 180,000
Senior Manager of SRE
Senior Manager of SRE

Hackajob Ltd • Cumbernauld

On-site
GBP 90,000 - 140,000
Senior Manager of SRE
Senior Manager of SRE

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 120,000 - 180,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Site Reliability Engineer II - AI & Corporate Risk Tech
Site Reliability Engineer II - AI & Corporate Risk Tech

Next Frontier Capital • Glasgow

On-site
GBP 65,000 - 95,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE- Azure & GCP
Lead SRE- Azure & GCP

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Lead Software Engineer - Software Reliability
Lead Software Engineer - Software Reliability

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Site Reliability Engineer II - AI & Corporate Risk Tech
Site Reliability Engineer II - AI & Corporate Risk Tech

JPMorganChase • Glasgow

On-site
GBP 70,000 - 110,000
Lead SRE - AWS,Python
Lead SRE - AWS,Python

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000