Principal Site Reliability Engineer (Cloud, Observability & Automation)

The Depository Trust & Clearing Corporation (DTCC)

Jersey City (NJ)

Hybrid

USD 140,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Base pay + annual incentive
Health and life insurance
Pension / Retirement benefits
Paid Time Off & family care leaves
Hybrid model: 3 days onsite, 2 days 1)

Job summary

DTCC is seeking an experienced Site Reliability Engineer to ensure reliability and performance of enterprise applications across ITP and ECS business lines. You will drive observability, automation, and resiliency initiatives in a hybrid cloud environment.

The role requires strong incident management, collaboration with cross-functional teams, and the ability to design and measure SLOs/SLIs. Hybrid work schedule and comprehensive benefit programs are offered.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent.
  • 8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines.

Responsibilities

  • Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications.
  • Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms.
  • Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs.
  • Lead major incident response, root cause analysis, and continuous service improvement initiatives.
  • Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies.
  • Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle.
  • Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives.
  • Identify operational risks and deliver strategic reliability improvements across the technology ecosystem.
  • Collaborate with technical and business stakeholders to improve service reliability and operational outcomes.

Skills

AWS
Python
Java
Go
Linux
Observability
Incident management
Distributed systems
Communication

Education

Bachelor's degree in Computer Science, Engineering, or equivalent

Tools

Splunk
Grafana
Dynatrace
ITSI
Amazon Q
Kiro

Job description

Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

Pay And Benefits
  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits, based on location
  • Pension / Retirement benefits
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).
The Impact You Will Have in This Role

The Enterprise Application Support (EAS) team supports critical applications across the ITP and ECS business lines, ensuring the reliability, scalability, and performance of enterprise platforms.

Your Primary Responsibilities
  • Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications.
  • Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms.
  • Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs.
  • Lead major incident response, root cause analysis, and continuous service improvement initiatives.
  • Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies.
  • Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle.
  • Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives.
  • Identify operational risks and deliver strategic reliability improvements across the technology ecosystem.
  • Collaborate with technical and business stakeholders to improve service reliability and operational outcomes.
Qualifications
  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines.
Talent Needed for Success
  • Strong hands-on experience with AWS and cloud-native architectures.
  • Proficiency in Python, Java, Go, or similar programming languages.
  • Strong Linux/Unix systems administration and troubleshooting experience.
  • Expertise in observability and monitoring platforms including Splunk, Grafana, Dynatrace, and ITSI.
  • Experience leading major incident management and root cause investigations in complex production environments.
  • Strong understanding of distributed systems, resiliency engineering, performance tuning, automation, and operational excellence.
  • Excellent communication and stakeholder management skills with the ability to influence technical and business partners.
Preferred Qualifications
  • Experience with AI-assisted engineering tools such as Amazon Q, Kiro, or similar technologies.
  • Experience designing and measuring SLOs, SLIs, and operational KPIs.
  • Experience supporting large-scale enterprise applications in financial services or other highly regulated environments.

The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations.

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Application Support Engineer (SRE)
Senior Application Support Engineer (SRE)

The Depository Trust & Clearing Corporation (DTCC) • Tampa (FL)

Hybrid
USD 120,000 - 170,000
Competitive compensation
Health & life insurance
Pension / Retirement benefits
+2
Senior Application Support Engineer (SRE)
Senior Application Support Engineer (SRE)

The Depository Trust & Clearing Corporation (DTCC) • Jersey City (NJ)

On-site
USD 120,000 - 180,000
Base pay
Health & life insurance
Pension
+2
Lead Delivery and Implementation Engineer
Lead Delivery and Implementation Engineer

The Depository Trust & Clearing Corporation (DTCC) • Tampa (FL)

Hybrid
USD 120,000 - 180,000
Competitive compensation
Comprehensive health and life insurance
Pension / Retirement benefits
+2
Lead Application Support Engineer
Lead Application Support Engineer

The Depository Trust & Clearing Corporation (DTCC) • Boston (MA)

Hybrid
USD 120,000 - 180,000
Competitive compensation incl. base +
Health and life insurance
Pension / Retirement benefits
+2
Associate Director Cloud Operations
Associate Director Cloud Operations

The Depository Trust & Clearing Corporation (DTCC) • Tampa (FL)

Hybrid
USD 140,000 - 190,000
Health insurance
Pension
Paid Time Off
Lead Application Support Engineer (12-8pm EST Shift)
Lead Application Support Engineer (12-8pm EST Shift)

The Depository Trust & Clearing Corporation (DTCC) • Town of Texas (WI)

Hybrid
USD 100,000 - 120,000
Competitive compensation
Comprehensive health and life insurance
Pension / Retirement benefits
+1
Director Infrastructure Resiliency Engineering
Director Infrastructure Resiliency Engineering

The Depository Trust & Clearing Corporation (DTCC) • Jersey City (NJ)

Hybrid
USD 180,000 - 240,000
Hybrid work model (3 onsite / 2 remote
Competitive compensation
Health & life insurance
+2
Principal DevOps Engineering Manager
Principal DevOps Engineering Manager

The Depository Trust & Clearing Corporation (DTCC) • Jersey City (NJ)

Hybrid
USD 130,000 - 160,000
Competitive compensation
Comprehensive health benefits
Pension / Retirement benefits
+1
Senior Full Stack Software Engineer - Java/AWS
Senior Full Stack Software Engineer - Java/AWS

The Depository Trust & Clearing Corporation (DTCC) • Tampa (FL)

Hybrid
USD 140,000 - 170,000
Competitive compensation
Health and life insurance
Pension / Retirement benefits
+2
Application Support Engineer
Application Support Engineer

The Depository Trust & Clearing Corporation (DTCC) • Town of Texas (WI)

Hybrid
USD 90,000 - 140,000
Annual incentive
Health insurance
Pension plan
+2