Site Reliability Engineer (SRE)

Methodic

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Methodic is seeking an experienced Site Reliability Engineer to join our San Francisco on-site team. You will own monitoring, alerting, and self-healing workflows, ensuring reliability and scalability for our platform.

You will implement capacity planning, blue-green or canary deployments, and collaborate with software engineers to embed resilience by design. Expect on-call rotations and ongoing automation to reduce toil and improve performance.

Qualifications

  • 5+ years in a Site Reliability Engineering or similar role.
  • Strong coding/scripting abilities (Python, Go, or other) for building automation and tooling.
  • Deep knowledge of systems monitoring and observability — experience with tools like Grafana, Datadog, Prometheus, etc., and the ability to interpret system metrics to spot problems.
  • Understanding of high-availability design and distributed systems principles (load balancing, consensus, graceful degradation, etc.).
  • Experience with incident management and a track record of improving systems based on lessons learned.
  • Performance tuning experience — ability to use profiling and stress testing tools to find and fix bottlenecks.
  • A collaborative mindset, capable of working with development teams to ensure reliability is built in from the start, not just after issues occur.

Responsibilities

  • Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers.
  • Perform regular capacity planning and load testing to ensure the platform can scale ahead of demand without performance degradation (validating that we can sustain extremely high transaction rates per tenant).
  • Improve deployment processes with strategies like blue-green or canary deployments to minimize risk and downtime during releases.
  • Collaborate with software engineers to design resilient architectures – for example, build redundancy and failover capabilities into critical services (so that even if one component fails, the system remains operational).
  • Participate in on-call rotations to respond to and resolve production incidents; lead blameless post-mortems to identify root causes and implement corrective actions.
  • Automate routine operational tasks (from simple scripts to more complex tooling) to reduce manual work and error potential.
  • Document reliability-related procedures and best practices, ensuring knowledge is shared and systems are well-understood by the team.

Skills

Python
Go
Scripting
Observability
Collaboration
On-call

Tools

Grafana
Datadog
Prometheus

Job description

Open role

Site Reliability Engineer (SRE)

San Francisco, CA (On-site)

Responsibilities
  • Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers.
  • Perform regular capacity planning and load testing to ensure the platform can scale ahead of demand without performance degradation (validating that we can sustain extremely high transaction rates per tenant).
  • Improve deployment processes with strategies like blue-green or canary deployments to minimize risk and downtime during releases.
  • Collaborate with software engineers to design resilient architectures – for example, build redundancy and failover capabilities into critical services (so that even if one component fails, the system remains operational).
  • Participate in on-call rotations to respond to and resolve production incidents; lead blameless post-mortems to identify root causes and implement corrective actions.
  • Automate routine operational tasks (from simple scripts to more complex tooling) to reduce manual work and error potential.
  • Document reliability-related procedures and best practices, ensuring knowledge is shared and systems are well-understood by the team.
Requirements
  • 5+ years in a Site Reliability Engineering or similar role.
  • Strong coding/scripting abilities (Python, Go, or other) for building automation and tooling.
  • Deep knowledge of systems monitoring and observability — experience with tools like Grafana, Datadog, Prometheus, etc., and the ability to interpret system metrics to spot problems.
  • Understanding of high-availability design and distributed systems principles (load balancing, consensus, graceful degradation, etc.).
  • Experience with incident management and a track record of improving systems based on lessons learned.
  • Performance tuning experience — ability to use profiling and stress testing tools to find and fix bottlenecks.
  • A collaborative mindset, capable of working with development teams to ensure reliability is built in from the start, not just after issues occur.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Washington

On-site
USD 120,000 - 180,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Shya Workforce Solutions • Town of Florida (NY)

On-site
USD 100,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000