Sr. Site Reliability Engineer

Jobgether

United States

Remote

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
Flexible vacation policy
Work-from-home support
Learning and development programs

Job summary

A leading technology firm is seeking a Sr. Site Reliability Engineer in the United States. The ideal candidate will enhance system reliability and stability and should possess over 8 years of relevant experience in site reliability engineering. The position covers cloud platforms, container orchestration, and system optimization, alongside competitive salary and comprehensive health benefits. Individuals who enjoy mentoring and problem-solving in collaborative environments are encouraged to apply, ensuring the firm's systems are efficient and dependable.

Qualifications

  • 8+ years of experience in site reliability, systems engineering, or operations.
  • Expert-level Linux administration and advanced troubleshooting skills.
  • Deep understanding of distributed systems and microservices architecture.

Responsibilities

  • Own and enhance the availability and performance of production services.
  • Lead complex reliability projects ensuring high-quality ownership.
  • Define and enforce service health standards including SLIs and SLOs.

Skills

Problem-solving
Collaboration
Communication

Education

Bachelor’s degree in Computer Science or related field

Tools

Kubernetes
Docker
Python
Go
AWS
Terraform

Job description

This position is posted by Jobgether on behalf of a partner company. We are looking for a Sr. Site Reliability Engineer in the United States. The role offers a unique opportunity to ensure the stability, scalability, and reliability of critical systems in a fast‑paced, cloud‑focused environment. The Sr. Site Reliability Engineer will work across engineering, product, and operations teams to embed reliability practices into daily workflows, automate processes, and proactively prevent system issues. This position requires a balance of hands‑on technical expertise and strategic thinking to drive infrastructure improvements, optimize operational efficiency, and maintain high service availability. It provides exposure to modern cloud platforms, containerized environments, and large‑scale distributed systems while giving you a chance to influence reliability standards and incident response practices. Ideal candidates are problem‑solvers who enjoy mentoring others, designing resilient systems, and improving operational processes.

Accountabilities
  • Own and enhance the availability, durability, and performance of production services across all environments
  • Lead complex reliability projects from problem identification to resolution, ensuring high‑quality technical ownership
  • Define and enforce service health standards, including SLIs, SLOs, and error budgets
  • Lead critical incident response and post‑incident reviews, translating insights into long‑term architectural improvements
  • Design and implement scalable automation, monitoring, logging, and alerting solutions to reduce manual effort
  • Build and maintain infrastructure‑as‑code, CI/CD pipelines, and operational tools to improve efficiency
  • Collaborate with engineering, product, and operations teams to embed reliability practices and guide resilient system design
  • Develop operational playbooks, runbooks, and documentation to support continuous improvement and knowledge sharing
Requirements
  • Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent experience
  • 8+ years of progressive experience in site reliability, systems engineering, or operations
  • Expert‑level Linux administration, advanced troubleshooting, and system security skills
  • Deep understanding of distributed systems, container orchestration (Kubernetes/Docker), and microservices architecture
  • Proficiency in scripting/programming languages such as Python, Go, or Bash
  • Experience with monitoring, logging, and alerting frameworks (Prometheus, Grafana, ELK, Catchpoint)
  • Strong familiarity with cloud platforms (AWS, GCP, or Azure) and Hashicorp tools (Terraform, Vault, Nomad)
  • Excellent problem‑solving, collaboration, and communication skills, with a proactive approach to continuous improvement
  • Preferred: ITIL/OSS experience, SaaS or hyper‑scale distributed system experience, and a history of mentoring teams on reliability best practices
Benefits
  • Competitive salary in the range of $150,000-$200,000 USD, based on experience and location
  • Comprehensive healthcare coverage, including dental and vision for family members
  • 401(k) plan with company matching and potential RSU grants
  • Flexible vacation policy and parental leave
  • Work‑from‑home support including equipment stipend
  • Learning and development programs to grow technical expertise and career trajectory
  • Culture that promotes work‑life balance and collaborative problem‑solving
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Oaks (PA)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare coverage
401(k) matching
Tuition reimbursement
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000