Senior Site Reliability Engineer

Axiom Pursuits

San Francisco (CA)

On-site

USD 150,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology firm is seeking a Lead Site Reliability Engineer to design and implement automated infrastructure and manage Kubernetes workloads. The role involves refining CI/CD pipelines and leading incident response efforts, requiring expertise in Terraform, Prometheus, and Grafana. With over 12 years of industry experience, the ideal candidate will ensure architectural scalability and reliability, emphasizing security and system resilience in a cloud-native environment. This is a pivotal role in delivering highly available and performant systems.

Qualifications

  • Minimum of 12 years of industry experience in software and systems engineering.
  • Proven ability in designing automated infrastructure.
  • Experience with monitoring and observability solutions.

Responsibilities

  • Design and implement automated infrastructure.
  • Manage containerized workloads within Kubernetes.
  • Refine CI/CD pipelines for seamless code deployment.
  • Lead incident response and post-mortem analyses.

Skills

Terraform
Kubernetes
CI/CD pipelines
Prometheus
Grafana
Automation
Security mindset
System resilience

Job description

US Corp. is seeking a Lead Site Reliability Engineer to spearhead our mission of delivering highly available and performant systems. With an average of over 12 years of industry experience, the successful candidate will bridge the gap between software development and systems engineering. You will be responsible for designing and implementing automated infrastructure using Terraform, managing containerized workloads within Kubernetes, and refining our CI/CD pipelines to ensure seamless code deployment. This role requires a deep dive into system internals, identifying bottlenecks, and implementing robust monitoring and observability solutions using Prometheus and Grafana. As a technical leader, you will define and maintain SLIs, SLOs, and SLAs while leading incident response and post-mortem analyses to prevent recurrence. You will work closely with product teams to ensure architectural scalability and reliability from the ground up. The ideal candidate is an automation enthusiast with a proactive mindset toward security, scalability, and system resilience in a cloud-native environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Remote Senior Site Reliability Lead - AWS, Kubernetes, CI/CD
Remote Senior Site Reliability Lead - AWS, Kubernetes, CI/CD

Empower • United States

On-site
USD 114,000 - 166,000
401(k) with company matching
Tuition reimbursement
Paid volunteer time
+1
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • California (MO)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Govcio LLC • Arlington (TX)

Hybrid
USD 230,000 - 250,000