Staff Site Reliability Engineer (Technical Lead) - Platform Engineering
Location: Chicago, IL (Hybrid)
Overview
A leading technology organization is seeking a Staff Site Reliability Engineer to provide technical leadership across a global Platform Engineering environment. In this role you will drive cloud modernization, reliability engineering, platform automation, and AI-enabled operational excellence while serving as a trusted technical advisor to engineering leadership.
Responsibilities
- Define and execute the technical roadmap for platform reliability, cloud infrastructure, and automation.
- Lead enterprise-scale cloud migration and modernization initiatives, with a focus on Google Cloud Platform.
- Serve as the technical authority for large-scale infrastructure and platform architecture decisions.
- Drive adoption of AI and agentic technologies to improve automation, incident response, and operational efficiency.
- Establish reliability standards including SLIs, SLOs, and error budgets.
- Mentor engineers and provide architectural guidance across multiple teams.
- Champion Infrastructure as Code, Kubernetes, GitOps, and cloud-native best practices.
Qualifications
- 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, Systems Engineering, or Software Engineering.
- Experience as a Technical Lead, Staff Engineer, Principal Engineer, or recognized SME.
- Expert-level Python development experience, including automation and distributed systems.
- Strong hands-on experience with Google Cloud Platform (GCP) and Kubernetes.
- Proven success leading cloud migration, platform transformation, or large-scale engineering initiatives.
- Experience leveraging Generative AI and agentic tools to improve engineering productivity and operational workflows.
- Deep understanding of distributed systems, scalability, observability, and high-availability architectures.
- Experience with Terraform, CI/CD, and modern platform engineering practices.
- Strong communication skills with the ability to influence technical and business stakeholders.
Preferred
- Experience supporting globally distributed, mission-critical platforms at enterprise scale.
- Expertise with Kafka or large-scale event-driven architectures.
- Cloud and Kubernetes certifications.
- Background in highly regulated, high-availability, or performance-sensitive environments.
This is a hybrid role based out of the firms Chicago office requiring 2 days of onsite work per week. At this time the team is unable to provide visa sponsorship.