IAM Engineering Director

DTCC

Greater London

On-site

GBP 110,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Health & life insurance
Pension
Paid time off
Hybrid work model

Job summary

DTCC in the United Kingdom seeks an experienced Director of IAM Site Reliability Engineering (SRE) & Observability to lead the reliability and resilience of our IAM ecosystem.

You will set monitoring standards, drive automation, partner with IAM engineering and security teams, and shape enterprise-wide observability practices to ensure maximum uptime.

This leadership role demands strategic thinking, strong communication with executives, and a track record of delivering scalable IAM solutions.

Qualifications

  • Bachelor's degree or equivalent in a related field.
  • 10+ years in SRE/Platform/ IAM related roles.
  • Experience with PAM, Active Directory, PKI, Secrets Management, and Cloud Identity platforms.
  • Strong observability, telemetry design, and instrumentation experience.

Responsibilities

  • Lead the IAM Site Reliability Engineering (SRE) function across all IAM platforms and services.
  • Own platform availability, service health, resiliency, observability, and operational readiness objectives.
  • Establish reliability engineering practices and operational excellence standards across the IAM ecosystem.
  • Partner with engineering teams to embed reliability-by-design principles.
  • Define and implement enterprise observability standards across all IAM platforms.
  • Develop monitoring architecture standards including infrastructure, application, user experience, and security event monitoring.
  • Drive implementation of centralized observability platforms with tools like Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, App Insights.
  • Define platform-specific monitoring requirements and instrumentation standards for all IAM services.
  • Ensure IAM applications meet monitoring and observability requirements prior to production deployment.
  • Create operational architecture diagrams, service dependency maps, data flow diagrams, and resiliency models for IAM services.
  • Review designs with IAM Engineering to identify availability and scalability risks.
  • Establish architecture standards for HA, DR, multi-site resiliency, failover, capacity, fault tolerance, and service recovery.
  • Conduct reliability design reviews and production readiness assessments for IAM platforms.

Skills

IAM
SRE leadership
Observability
Monitoring architecture
Incident management
Automation
IaC
Cloud platforms
Executive communication
Stakeholder mgmt

Education

Bachelor's degree or equivalent

Tools

Splunk
Dynatrace
Datadog
Grafana
Azure Monitor
App Insights

Job description

Job Description

Are you ready to make an impact at DTCC?

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We're committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

Pay and Benefits:
  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits
  • Pension
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (Onsite Tuesdays, Wednesdays and a third day of your choosing)
The impact you will have in this role:

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Role Summary:

We are seeking an experienced Director of IAM Site Reliability Engineering (SRE) & Observability to lead the reliability, availability, operational architecture, monitoring strategy, and resiliency of the enterprise Identity and Access Management (IAM) ecosystem.

This leader will be responsible for ensuring the health, performance, scalability, recoverability, and operational readiness of critical IAM platforms including Privileged Access Management (PAM), Active Directory, Certificate Management Infrastructure (PKI), Secrets Management, Authentication Services, Cloud Identity Platforms, and related Identity Security services.

The role will establish enterprise-wide observability standards, develop service reliability architectures, define monitoring frameworks, and partner closely with Engineering teams to ensure IAM platforms are designed, instrumented, and operated for maximum availability and resilience.

The successful candidate will drive a proactive reliability culture focused on monitoring, telemetry, automation, service health, incident prevention, operational architecture, and continuous improvement.

Your Primary Responsibilities:
IAM Site Reliability Engineering Leadership
  • Lead the IAM Site Reliability Engineering (SRE) function across all IAM platforms and services.
  • Own platform availability, service health, resiliency, observability, and operational readiness objectives.
  • Establish reliability engineering practices and operational excellence standards across the IAM ecosystem.
  • Partner with engineering teams throughout the software and platform lifecycle to embed reliability-by-design principles.
Observability & Monitoring Strategy
  • Define and implement enterprise observability standards across all IAM platforms.
  • Develop monitoring architecture standards covering:
    • Infrastructure Monitoring
    • Application Monitoring
    • User Experience Monitoring
    • Transaction Monitoring
    • Dependency Monitoring
    • Security Event Monitoring
    • Cloud Service Monitoring
  • Establish standards for logging, metrics, tracing, dashboards, alerting, correlation, and telemetry collection.
  • Drive implementation of centralized observability platforms leveraging tools such as Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, App Insights, or equivalent solutions.
  • Define platform-specific monitoring requirements and operational instrumentation standards for all IAM services.
  • Ensure all IAM applications meet monitoring, alerting, and observability requirements prior to production deployment.
Reliability Architecture & Engineering
  • Create operational architecture diagrams, service dependency maps, data flow diagrams, and platform resiliency models for IAM services.
  • Work closely with IAM Engineering teams to review solution designs and identify availability, scalability, and resiliency risks.
  • Establish architecture standards for:
    • High Availability (HA)
    • Disaster Recovery (DR)
    • Multi-site Resiliency
    • Failover Design
    • Capacity Planning
    • Fault Tolerance
    • Service Recovery
  • Conduct reliability design reviews and production readiness assessments for IAM platforms and applications.
Availability & Resilience Management
  • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and Availability Targets.
  • Drive initiatives to improve platform uptime, stability, recoverability, and performance.
  • Identify and eliminate single points of failure and operational bottlenecks.
  • Lead resiliency testing exercises, failover validation, disaster recovery testing, and business continuity preparedness.
Incident & Service Health Management
  • Lead major incident management, executive communications, and post-incident reviews.
  • Analyze incident trends, recurring failures, and reliability gaps to drive platform improvements.
  • Establish proactive service health review processes and operational risk assessments.
  • Ensure corrective actions are tracked and implemented to prevent repeat incidents.
Automation & Operational Excellence
  • Drive automation of monitoring, alert triage, remediation, service validation, and operational workflows.
  • Promote Infrastructure as Code (IaC), automated health checks, self-healing capabilities, and operational engineering practices.
  • Improve alert quality by reducing noise and focusing on actionable service indicators.
  • Establish reliability scorecards and operational maturity frameworks across IAM platforms.
Cross-Functional Leadership
  • Partner with IAM Engineering, Security Engineering, Enterprise Architecture, Cloud, Infrastructure, Network, and Vendor teams.
  • Serve as the primary authority for IAM platform observability, availability, and operational architecture standards.
  • Influence engineering roadmaps by incorporating reliability, monitoring, and resiliency requirements into platform design and development processes.
NOTE: The Primary Responsibilities of this role are not limited to the details above.
Qualifications:
  • Bachelor's degree preferred or equivalent experience
Talents Needed For Success:
  • Minimum of 10 years related experience
  • 12+ years of experience in Site Reliability Engineering, Platform Engineering, IAM, Infrastructure Engineering, or Cybersecurity.
  • Experience supporting complex IAM environments including PAM, Active Directory, PKI, Authentication Services, Secrets Management, and Cloud Identity platforms.
  • Strong expertise in observability, monitoring architecture, telemetry design, and platform instrumentation.
  • Experience creating architecture diagrams, dependency maps, operational blueprints, and resiliency designs.
  • Hands-on experience with enterprise monitoring and observability platforms.
  • Strong knowledge of High Availability, Disaster Recovery, Business Continuity, and resilience engineering.
  • Experience leading major incident management and operational transformation programs.
  • Proven ability to influence architecture and engineering teams without direct ownership.
  • Strong executive communication and stakeholder management skills.
We offer top class training and development for you to be an asset in our organization!

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

IAM Engineering Director
IAM Engineering Director

The Depository Trust & Clearing Corporation (DTCC) • Greater London

Hybrid
GBP 120,000 - 180,000
Base salary + annual incentive
Health & life insurance
Pension
+2
IAM SRE Director: Reliability, Observability & Security
IAM SRE Director: Reliability, Observability & Security

The Depository Trust & Clearing Corporation (DTCC) • Greater London

Hybrid
GBP 120,000 - 180,000
Base salary + annual incentive
Health & life insurance
Pension
+2
IAM SRE Director - Reliability & Observability Leader
IAM SRE Director - Reliability & Observability Leader

DTCC • Greater London

Hybrid
GBP 110,000 - 160,000
Competitive compensation
Health & life insurance
Pension
+2
Lead IT Security Engineer
Lead IT Security Engineer

The Depository Trust & Clearing Corporation (DTCC) • Greater London

Hybrid
GBP 70,000 - 110,000
Health insurance
Pension
Time off & leaves
+1
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Identity & Privileged Access Management Engineering Director / Leader - London/ Financial Services
Identity & Privileged Access Management Engineering Director / Leader - London/ Financial Services

Entasis Partners • Greater London

On-site
GBP 140,000 - 180,000
Performance bonus
Comprehensive benefits
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
Senior IAM Engineer
Senior IAM Engineer

Selby Jennings • City Of London

On-site
GBP 80,000 - 120,000
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Senior IAM Engineer
Senior IAM Engineer

NCC Group • Greater Manchester

On-site
GBP 70,000 - 110,000
Flexible Working
25 days holiday + bank holidays
Medicash & Critical Illness
+2