Stand out for this role — generate a tailored resume and cover letter in about a minute.
Jobtailor is seeking a senior Site Reliability Engineer to tackle reliability, scalability, and efficiency challenges across SRE and development teams. You will build and run large-scale, distributed fault-tolerant systems that power the Genesis platform, optimize existing infrastructure, and cut toil through automation to improve uptime.
The role demands strong Python/Go skills, deep distributed systems design, and leadership across teams.
Take on ambiguous reliability, scalability, and efficiency challenges and drive solutions across SRE and development teams
Build and run large-scale, massively distributed, fault-tolerant systems supporting the Genesis platform
Optimize existing systems and build infrastructure
Eliminate toil through automation to improve uptime and rate of change
Cultivate a culture of reliability throughout the organization
Guide technical decisions balancing system health with product priorities
Ensure long-term health, maintainability, and reliability of services
Perform capacity planning and performance analysis
Proactively prevent incidents
Work across teams to build robust, reusable solutions
Demonstrates strong software engineering skills in Python and Go, with extensive experience in designing and troubleshooting distributed systems. Proven ability to lead complex technical projects and optimize large-scale, fault-tolerant systems while fostering a culture of reliability and collaboration.