Site Reliability Engineer

Crump Life Insurance Svcs Inc

Atlanta (GA)

On-site

USD 120,000 - 160,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Life insurance
Disability insurance
401k plan
Vacation days
Sick days
Pension plan
Restricted stock units
Deferred compensation plan

Job summary

Truist is seeking a Senior Site Reliability Engineer to enhance reliability across hybrid cloud and on-prem environments. You will lead incident responses, drive problem management, and implement automation to reduce downtime, while standardizing observability practices and mentoring SREs.

Expect collaboration with multiple business and tech teams and strategic reliability initiatives. The role requires 7+ years in SRE/DevOps, expertise in distributed systems and Kubernetes, and strong

Qualifications

  • 7+ years of professional experience in software engineering or related field.
  • Deep knowledge of programming languages, software architecture, and design principles.
  • Strong understanding of SDLC, testing, deployment and security practices.
  • Experience in distributed systems and SRE/DevOps.

Responsibilities

  • Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results.
  • Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery.
  • Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.
  • Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow.
  • Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives.
  • Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.

Skills

Distributed systems
Kubernetes
Automation scripting
Incident management leadership
Observability platforms

Education

Bachelor’s degree in Computer Science, Software Engineering, or related field

Tools

Dynatrace
Splunk
CI/CD tooling
Cloud-native tooling

Job description

The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams. Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime. The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks. Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management.

ESSENTIAL DUTIES AND RESPONSIBILITIES
  1. Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results.
  2. Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery.
  3. Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.
  4. Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow.
  5. Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives.
  6. Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.
  7. Evaluates emerging technologies and techniques relevant to the job area, building prototypes and solution concepts that contribute measurable input into new features, products, or capabilities.
  8. Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations, design proposals, and implementation experience.
  9. Leads large or complex initiatives within the job area, coordinating and delegating technical work that may span outside the immediate team, and ensuring cohesive, high-quality outcomes with limited supervision.
Qualifications Required
  1. Bachelor’s degree in Computer Science, Software Engineering, or related field.
  2. Minimum of 7 years of professional experience in software development.
  3. Deep knowledge of multiple programming languages, software architecture, and design principles.
  4. Deep understanding of software development lifecycle, testing, deployment, and security practices.
Preferred Qualifications
  1. Advanced degree in Computer Science or related technical discipline.
  2. Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
  3. Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps.
  4. Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management.
  5. 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
  6. Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
  7. Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
  8. Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
  9. Proven leadership in major incident management and cross-team technical coordination.
  10. Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
  11. Excellent communication skills, including executive-level situational awareness during critical incidents.
  12. Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
  13. Financial services or regulated industry experience.
  14. Experience enabling large-scale SRE transformations or modernization initiatives.
  15. Familiarity with chaos engineering, resilience assessments, and service failure modeling.
  16. Exposure to hybrid-cloud and multi-cloud operational frameworks.
  17. Experience contributing to or leading Center for Enablement functions or Communities of Practice.
Key Responsibilities

Incident & Problem Management Leadership Lead major and high-severity incident response efforts, focusing on diagnosing technical root causes therein, and driving multi-team technical resolution. Drive problem management to closure, ensuring systemic fixes replace recurring operational risks. Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks.

Reliability Engineering & Automation Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience. Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools. Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision‑making and prioritization.

Observability & Operational Excellence Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk. Define and standardize enterprise observability practices, dashboards, and KPIs. Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.

Cross‑Functional Leadership & Influence Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution. Act as a change agent to elevate operational maturity and drive transformative improvements across Wholesale. Lead workshops, maturity assessments, and enablement sessions through the SRE C4E and Communities of Practice.

Standardization & Documentation Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns. Contribute to enterprise SRE frameworks, templates, and maturity models. Promote consistent adoption of best practices across domains and lines of business.

Mentorship & Technical Development Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline. Provide thought leadership in SRE methodologies, cloud‑native operational patterns, and automated reliability engineering.

Eligibility and Legal Disclaimers

For this opportunity, Truist will not sponsor an applicant for work visa status or employment authorization, nor will we offer any immigration-related support for this position (including, but not limited to H-1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN-1 or TN-2, E-3, O-1, or future sponsorship for U.S. lawful permanent residence status.)

Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA.

Benefits
  • medical
  • dental
  • vision
  • life insurance
  • disability
  • accidental death and dismemberment
  • tax‑preferred savings accounts
  • 401k plan
  • no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during the first year of employment
  • 10 sick days (also prorated), and paid holidays
  • defined benefit pension plan
  • restricted stock units
  • deferred compensation plan
Equal Opportunity Statement and Company Culture

Truist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace. EEO is the Law E-Verify IER Right to Work About Truist Truist is a purpose-driven financial services company, formed by the historic merger of equals of BB&T and SunTrust. We serve clients in a number of high-growth markets in the country, offering a wide range of financial services. At Truist, our purpose is to inspire and build better lives and communities. That happens through real care to make things better. To meet client needs, to empower teammates, and to lift up communities. Learn more about Truist on truist.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Socket.dev • Atlanta (GA)

On-site
USD 140,000 - 180,000
Site Reliability Engineering
Site Reliability Engineering

Habitat For Humanity Of Durham • Raleigh (NC)

On-site
USD 140,000 - 170,000
Site Reliability Engineering
Site Reliability Engineering

Fayette Chamber of Commerce • Atlanta (GA)

On-site
USD 120,000 - 170,000
Medical insurance
Dental insurance
Vision insurance
+3
Site Reliability Engineering
Site Reliability Engineering

Truist • Raleigh (NC)

On-site
USD 130,000 - 180,000
Site Reliability Engineering
Site Reliability Engineering

Truist • Charlotte (NC)

On-site
USD 140,000 - 190,000
Medical, dental, vision
401k plan
Paid time off
Site Reliability Engineering
Site Reliability Engineering

Truist • Atlanta (GA)

On-site
USD 150,000 - 190,000
Sr. Availability & Reliability Engineering Manager
Sr. Availability & Reliability Engineering Manager

Crump Life Insurance Svcs Inc • Charlotte (NC)

On-site
USD 150,000 - 190,000
Infrastructure Automation Engineering Manager
Infrastructure Automation Engineering Manager

Crump Life Insurance Svcs Inc • Atlanta (GA)

On-site
USD 125,000 - 169,000
Health insurance
401k plan
Paid time off
Cyber Incident Management Engineer
Cyber Incident Management Engineer

Crump Life Insurance Svcs Inc • Greensboro (NC)

On-site
USD 120,000 - 165,000
Medical benefits
401(k) plan
Paid time off
Senior Infrastructure Engineer - Network Observability
Senior Infrastructure Engineer - Network Observability

Crump Life Insurance Svcs Inc • Atlanta (GA)

On-site
USD 120,000 - 160,000