Sr. Site Reliability Engineer

State of Wisconsin Investment Board

Madison (WI)

On-site

USD 140,000 - 180,000

Full time

3 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

State of Wisconsin Investment Board seeks a Senior Site Reliability Engineer to lead the build and transformation of its cloud-native stack. You will oversee SDLC aspects across software, data, and infrastructure in a 24x6, high-availability environment.

You will partner with cloud-native teams to deliver high-velocity software, implement reusable components, and mentor engineers while staying current with industry trends.

Qualifications

  • Enablement mindset – you win when teams win.
  • Excellent verbal and written communication skills.
  • Bachelor's degree or equivalent experience.
  • 8+ years of professional SRE experience.
  • Strong background designing and delivering complex architectures.
  • Hands-on cloud production experience (AWS preferred).
  • IaC experience with Terraform/OpenTofu.
  • Experience building/running CI/CD infra (GitLab preferred).
  • Experience with Kubernetes/EKS/ECS orchestration.
  • Deep information security understanding.
  • Experience with monitoring tools: Sentry, Prometheus, Datadog.
  • Interest in AI technologies.
  • Git proficiency; strong Agile experience.
  • Experience with cloud-native microservices and event-driven design.
  • Ability to work under pressure with multiple priorities.

Responsibilities

  • Function as subject matter expert for reliable, high-performance apps across multi-region/multi-cloud.
  • Collaborate with dev teams to design, document, and maintain highly available systems.
  • Design and manage centralized monitoring for rapid feedback to teams.
  • Partner with app teams to create workflow automation and tooling.
  • Facilitate AI evaluation/adoption within the software delivery platform.
  • Create an environment of continuous experimentation and learning.
  • Evolve cloud-focused architecture for flexibility and ease of use.
  • Track tech trends and propose improvements as appropriate.
  • Mentor less experienced teammates and share knowledge.
  • Provide tier 2/3 escalations and maintain runbooks and docs.
  • Establish KPIs and SLAs; drive continuous improvement.
  • Lead vendor relationships and manage third-party support.
  • Document support processes, incidents, and resolutions.

Skills

Cloud architecture
SRE leadership
AWS
Terraform/OpenTofu
GitLab CI/CD
Kubernetes / EKS / ECS
Monitoring: Prometheus / Datadog / Sny
AI in software delivery
Security concepts
Git version control

Education

Bachelor's Degree in Computer Science or related field

Tools

Terraform/OpenTofu
GitLab CI/CD
Kubernetes
Sentry
Prometheus
Datadog

Job description

We are seeking a highly skilled and experienced Senior Site Reliability Engineer to oversee the build and transformation of SWIB’s cloud native technology stack. This role will serve in a critical capacity to ensure all aspects of the Software Development Lifecycle (SDLC) are built from the ground up with modern tools and techniques. This will include all aspects of the SDLC across software, data and infrastructure. The ideal candidate will have a strong background in financial services, exceptional leadership skills, and the ability to manage platforms that require continuity on a 24x6 basis effectively.

This role will serve as a thought leader within the technology organization, helping to drive change and transformation across all teams.

The Senior Site Reliability Engineer will partner with cloud-native application development teams with direct business alignment – empowering them to deliver quality software at high velocity, rapidly receive actionable feedback, and cultivating an environment of continuous experimentation. Additionally, the Senior SRE will create reusable components and workflows that deliver superior engineering experience.

Key Responsibilities:
  • Function as subject matter expert in the areas of delivering and operating reliable, performant, robust, and secure applications, with an emphasis on multi-region and multi-cloud patterns.
  • Work with the development teams to design, document, create and maintain highly available systems.
  • Design, implement, and manage centralized monitoring solutions that provide expedient actionable feedback to the development teams.
  • Partner with the application development teams to create flows, processes, automation, and tooling.
  • Facilitate the evaluation, adoption, and integration of AI into the software delivery platform
  • Create an environment of continuous experimentation and learning.
  • Contribute to evolution of our architecture (cloud-focused) to increase its flexibility and ease of use.
  • Follow technology trends/tools and recommend improvements to our technology when appropriate.
  • Mentor new or less senior members of the team.
  • Share experience, knowledge, and ideas to the team to improve processes and productivity.
  • Provide tier 2 and 3 escalations for related issues and questions.
  • Establish and monitor key performance indicators (KPIs) and service level agreements (SLAs) to ensure the support team meets or exceeds performance expectations.
  • Conduct regular performance reviews and provide ongoing training and development opportunities for the support team.
  • Drive continuous improvement initiatives to enhance support processes, reduce incidents, and improve overall application reliability and user satisfaction.
  • Manage vendor relationships and ensure third-party support services align with organizational needs and standards.
  • Maintain comprehensive documentation of support processes, incidents, and resolutions.
  • Stay current with industry trends and emerging technologies to ensure the SRE function remains cutting-edge and effective.
Qualifications:
  • Enablement mindset and attitude – you win when teams win
  • Excellent verbal and written communication skills
  • Bachelor’s Degree in Computer Science, or a related field, or equivalent work experience
  • 8+ years of professional Site Reliability Engineering experience (or equivalent demonstrated impact)
  • Strong background in designing, implementing, and delivering complex technical architectures
  • Hands-on experience with the Cloud in a production environment (AWS preferred)
  • Solid experience implementing Infrastructure-as-Code (Terraform or OpenTofu preferred)
  • Hands-on experience building and running CICD infrastructure (GitLab preferred)
  • Experience implementing and operating container orchestration platforms such as Kubernetes, EKS, Elastic Container Service (ECS)
  • Deep understanding of information security concepts
  • Experience implementing and operating monitoring tools such as Sentry, Prometheus, and Datadog
  • Experience with or strong interest in AI technologies
  • Familiarity with version control systems such as Git
  • Working experience with agile methodologies
  • Experience with cloud-performant microservices and event driven architectures is a plus
  • Ability to work under pressure and manage multiple priorities in a fast-paced environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Charles Schwab Corporation • Southlake (TX)

On-site
USD 140,000 - 180,000
Junior Site Relaibilty Engineer
Junior Site Relaibilty Engineer

Charles Schwab Corporation • Austin (TX)

On-site
USD 90,000 - 120,000
Bonus opportunities
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000