At Early Warning, we've powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze℠, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.
Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.
Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.
Overall Purpose
The Sr. Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services. The role partners with Software Engineering and other technology teams to ensure reliability, observability, recoverability, performance, and operational readiness are engineered into systems throughout their lifecycle.
The role independently identifies systemic reliability risks, owns broader reliability outcomes, and influences engineering decisions across multiple services.
Essential Functions
- Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed, observed, operated, and recovered.
- Use data, evidence, experimentation, and rigorous engineering analysis appropriate to the level to identify reliability risks, test assumptions, and guide technical decisions.
- Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility.
- Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.
- Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
- Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.
- Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
- Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning appropriate to the level.
- Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices.
Leveling Intent
Senior represents the transition from independently performing SRE work to owning broader reliability outcomes and influencing how others engineer reliable systems. Senior SREs increasingly act as force multipliers by mentoring others and spreading reusable engineering practices.
Level Expectations
- Begins to act as a broader force multiplier by mentoring engineers, sharing knowledge, and developing reusable solutions and practices that increase the effectiveness of the team.
- Demonstrates software engineering, systems thinking, troubleshooting, and production reliability capabilities appropriate to the level.
- Applies evidence-driven reasoning and technical rigor to distinguish observed facts from assumptions and make defensible engineering recommendations.
- Shares knowledge and contributes to sustainable engineering capability rather than creating dependency on individual expertise.
- Operates independently across complex systems and ambiguous reliability problems.
- Owns reliability outcomes across multiple services and guides less-experienced engineers.
Minimum Qualifications
- T typically 5-8 years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline.
- Experience with software development or scripting using one or more modern programming languages.
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level.
- Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures appropriate to the level.
- Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role.
Preferred Qualifications
- Hands-on experience with AWS is preferred, or comparable experience with another major cloud platform such as Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
- Experience developing,