Salary: £100,000 - 100,000 per year
Requirements:
- Formal training or certification on site reliability engineering concepts and advanced applied experience
- Demonstrated hands-on experience with Amazon Web Services (AWS), including deploying, operating, and maintaining resilient, highly available workloads in a cloud environment
- Demonstrated proficiency in reliability, scalability, performance, and enterprise system architecture, with hands-on experience conducting resiliency design reviews and implementing resiliency best practices
- Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET
- Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments
- Proficiency with continuous integration and continuous delivery practices and tooling
- Proficiency with container technologies and container orchestration
- Experience troubleshooting common networking technologies and issues
- Advanced knowledge of software applications and technical processes with emerging depth in one or more technical disciplines, with a demonstrated ability to evaluate and recommend suitable new technologies
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve site reliability engineering workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
- Experience with cloud platforms and infrastructure-as-code tooling in an enterprise environment
- Familiarity with chaos engineering principles and proactive resiliency testing practices
- Experience contributing to communities of practice, internal knowledge-sharing forums, or engineering guilds
- Exposure to advanced observability platforms and distributed tracing in large-scale production environments
Responsibilities:
- Consistently champion site reliability culture and practices, documenting and sharing knowledge across our organization through internal forums and communities of practice
- Drive initiatives to improve the reliability and stability of our teams applications and platforms using data-driven analytics to improve service levels, proactively identifying and resolving technology-related bottlenecks
- Collaborate with our team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets
- Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response
- Serve as the primary point of contact during major incidents for our application, applying strong diagnostic skills to identify and resolve issues quickly and minimize business impact
- Apply deep technical expertise within one or more technical domains, sharing knowledge and providing guidance to peers across the team
- Drive reuse-first adoption of AI-assisted reliability workflows across the software development lifecycle and toolchain practices (e.g., continuous integration/continuous delivery quality checks, test and validation automation, and operational readiness), ensuring traceability, auditability, resiliency, and security controls
- Use enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements
Technologies:
- AI
- AWS
- Cloud
- Support
- Java
- Marketing
- Python
- Security
- Spring
- Spring Boot
- Web
- ASP.NET
More:
We are JPMorganChase, partnering directly with hackajob to hire for this Lead Site Reliability Engineer role within Infrastructure Platforms. This role gives you the opportunity to help define the future of a globally recognized firm, make a direct impact on reliability outcomes, and act as a technical authority for medium to large-sized products. We offer a first-class business in a first-class way approach, a strong culture of diversity and inclusion, and a corporate functions environment where our teams support finance, risk, human resources, marketing, and other essential business areas.
last updated 36 week of 2026