Principal SRE: Cloud Reliability & Observability Lead

The Depository Trust & Clearing Corporation (DTCC)

Jersey City (NJ)

Hybrid

USD 140,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Base pay + annual incentive
Health and life insurance
Pension / Retirement benefits
Paid Time Off & family care leaves
Hybrid model: 3 days onsite, 2 days 1)

Job summary

DTCC is seeking an experienced Site Reliability Engineer to ensure reliability and performance of enterprise applications across ITP and ECS business lines. You will drive observability, automation, and resiliency initiatives in a hybrid cloud environment.

The role requires strong incident management, collaboration with cross-functional teams, and the ability to design and measure SLOs/SLIs. Hybrid work schedule and comprehensive benefit programs are offered.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent.
  • 8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines.

Responsibilities

  • Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications.
  • Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms.
  • Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs.
  • Lead major incident response, root cause analysis, and continuous service improvement initiatives.
  • Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies.
  • Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle.
  • Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives.
  • Identify operational risks and deliver strategic reliability improvements across the technology ecosystem.
  • Collaborate with technical and business stakeholders to improve service reliability and operational outcomes.

Skills

AWS
Python
Java
Go
Linux
Observability
Incident management
Distributed systems
Communication

Education

Bachelor's degree in Computer Science, Engineering, or equivalent

Tools

Splunk
Grafana
Dynatrace
ITSI
Amazon Q
Kiro

Job description

DTCC is seeking an experienced Site Reliability Engineer to ensure reliability and performance of enterprise applications across ITP and ECS business lines. You will drive observability, automation, and resiliency initiatives in a hybrid cloud environment.

The role requires strong incident management, collaboration with cross-functional teams, and the ability to design and measure SLOs/SLIs. Hybrid work schedule and comprehensive benefit programs are offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Lead: Reliability, Automation & Observability
Senior SRE Lead: Reliability, Automation & Observability

Truist • Charlotte (NC)

On-site
USD 140,000 - 190,000
Medical, dental, vision
401k plan
Paid time off
Senior Application Reliability Engineer (Hybrid)
Senior Application Reliability Engineer (Hybrid)

The Depository Trust & Clearing Corporation (DTCC) • Jersey City (NJ)

Hybrid
USD 120,000 - 180,000
Base pay
Health & life insurance
Pension
+2
Principal Site Reliability Engineer (Cloud, Observability & Automation)
Principal Site Reliability Engineer (Cloud, Observability & Automation)

The Depository Trust & Clearing Corporation (DTCC) • Jersey City (NJ)

Hybrid
USD 140,000 - 190,000
Base pay + annual incentive
Health and life insurance
Pension / Retirement benefits
+2
Senior SRE Lead: Reliability, Automation & Incidents
Senior SRE Lead: Reliability, Automation & Incidents

Truist • Atlanta (GA)

On-site
USD 150,000 - 190,000
Principal SRE: Enterprise Reliability & Automation
Principal SRE: Enterprise Reliability & Automation

Early Warning • San Francisco (CA)

Hybrid
USD 194,000 - 284,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Principal SRE — Observability & Cloud Reliability
Principal SRE — Observability & Cloud Reliability

T. Rowe Price • Washington

Hybrid
USD 159,000 - 339,000
Competitive compensation
Annual bonus eligibility
Hybrid work schedule
+2
Senior SRE: Cloud Reliability, IaC & DR Leader
Senior SRE: Cloud Reliability, IaC & DR Leader

Castleton Commodities International • Stamford (CT)

On-site
USD 150,000 - 190,000
Competitive compensation
Comprehensive benefits
Tuition assistance
Principal SRE: Build Resilient, Scalable Cloud Systems
Principal SRE: Build Resilient, Scalable Cloud Systems

Worky • Durham (NC)

On-site
USD 150,000 - 210,000
Principal SRE: Hybrid Cloud Reliability & Observability
Principal SRE: Hybrid Cloud Reliability & Observability

Hewlett Packard Enterprise • San Juan (PR)

Hybrid
USD 120,000 - 160,000
Senior SRE: Automation & Resilience Leader
Senior SRE: Automation & Resilience Leader

Socket.dev • Atlanta (GA)

On-site
USD 140,000 - 180,000