Technical Operations Lead

First Citizens Bank

Dallas (TX)

On-site

USD 140,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

First Citizens Bank is seeking a Senior Site Reliability Engineer to join the Enterprise Observability and SRE team in Dallas. You will design, implement, and maintain highly available platforms, manage SLI/SLOs, and drive proactive reliability improvements across critical services.

You will lead monitoring strategies, automate toil, and collaborate with application, cloud, security, and production teams to raise operational excellence and reduce incidents.

Qualifications

  • Bachelor's degree in a technical field or equivalent experience.
  • 7+ years of SRE/DevOps/Platform Engineering experience in production environments.
  • Hands-on with observability and monitoring platforms (Dynatrace, Splunk, Databahn).
  • Strong understanding of APM, distributed tracing, log analytics, and infrastructure monitoring.
  • Experience with AWS, Azure, or Google Cloud Platform; scripting and automation proficient.

Responsibilities

  • Design, implement, and maintain highly available, scalable platforms.
  • Manage SLIs/SLOs and error budgets for critical services.
  • Lead reliability reviews and drive post-incident RCA actions.
  • Develop monitoring strategies, dashboards, alerts, and synthetic tests.
  • Embed observability in the SDLC and partner with cross-functional teams.
  • Automate toil and improve CI/CD for deployment reliability.
  • Participate in incident management and drive MTTR reductions.
  • Provide technical leadership and mentoring across the org.

Skills

SRE & DevOps
Production environments
APM & Observability
Scripting & Automation
Troubleshooting
Analytical thinking
Problem solving
Financial sector experience
Cross-team collaboration
OpenTelemetry

Education

Bachelor's degree in Computer Science / Engineering / Information Technology or equivalent

Tools

Dynatrace
Splunk
Databahn
Datadog
AWS
Azure
GCP
Kubernetes
CI/CD tooling
OpenTelemetry

Job description

Overview

We are seeking a highly skilled and motivated Senior Site Reliability Engineer (SRE) to join our Enterprise Observability and Site Reliability Engineering team. This role is responsible for improving the reliability, availability, performance, and operational excellence of critical enterprise platforms and applications.

The ideal candidate combines strong software engineering and infrastructure expertise with a passion for automation, observability, and operational excellence. You will partner closely with application development, cloud engineering, infrastructure, security, and production support teams to establish reliability standards, implement monitoring strategies, and drive proactive risk reduction across the technology landscape.

Responsibilities & Qualifications
Key Responsibilities
Reliability Engineering
  • Design, implement, and maintain highly available, resilient, and scalable technology platforms.
  • Manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across critical services.
  • Lead reliability reviews and identify opportunities to improve system stability and performance.
  • Drive root cause analysis and corrective actions for high-severity incidents during the problem process.
Observability & Monitoring
  • Serve as a subject matter expert for enterprise observability platforms, including Dynatrace and related monitoring technologies.
  • Develop monitoring, alerting, synthetic testing, and dashboard strategies.
  • Improve visibility into application health, infrastructure performance, user experience, and business transaction monitoring.
  • Partner with engineering teams to embed observability practices throughout the software development lifecycle.
Automation & Platform Engineering
  • Identify and eliminate operational toil through automation.
  • Develop scripts, tooling, integrations, and self-service capabilities.
  • Enhance CI/CD processes to improve deployment reliability and operational efficiency.
  • Support infrastructure-as-code and automation-first operating models.
Incident Management & Operational Excellence
  • Participate in major incident response and problem management activities.
  • Establish and improve operational runbooks, standards, and best practices.
  • Drive reduction in Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR).
  • Implement proactive monitoring and predictive alerting capabilities.
Leadership & Collaboration
  • Provide technical leadership and mentoring to engineers across the organization.
  • Influence reliability and observability strategy at the enterprise level.
  • Collaborate with application owners, cloud teams, security teams, and executive stakeholders.
  • Champion a culture of ownership, accountability, and continuous improvement.
Required Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience.
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or related disciplines.
  • Strong experience supporting mission-critical production environments.
  • Hands-on experience with observability and monitoring platforms such as Dynatrace, Splunk, Databahn or similar technologies.
  • Strong understanding of application performance monitoring (APM), distributed tracing, log analytics, and infrastructure monitoring.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Proficiency with scripting and automation.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Large financial institution experience in a complex environment.
Preferred Qualifications
  • Deep expertise with Dynatrace platform administration and implementation.
  • Experience in enterprise-scale financial services or highly regulated environments.
  • Knowledge of cloud-native observability patterns and OpenTelemetry.
  • Experience implementing SRE practices including SLOs, Error Budgets, reliability reviews, and operational readiness assessments.
  • Familiarity with enterprise event management and AIOps platforms.
  • Certifications in cloud platforms, Dynatrace, Kubernetes, or related technologies.
  • Experience leading large-scale monitoring transformation initiatives.
Additional Information

Benefits are an integral part of total rewards and First Citizens Bank is committed to providing a competitive, thoughtfully designed and quality benefits program to meet the needs of our associates. More information can be found at https://jobs.firstcitizens.com/benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Operations Lead
Technical Operations Lead

First Citizens Bank • Phoenix (AZ)

On-site
USD 140,000 - 190,000
Benefits program
Technical Operations Lead
Technical Operations Lead

First Citizens • Raleigh (NC)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Systems Engineer Consultant
Systems Engineer Consultant

First Citizens Bank • Scottsdale (AZ)

On-site
USD 180,000 - 240,000
Benefits package
401(k)
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • United States

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

100 CRC Insurance Group, LLC • Charlotte (NC)

On-site
USD 130,000 - 170,000
Health insurance
401(k) with match
Generous PTO
+1
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000