Senior SRE: Build Resilient, Scalable Systems

Methodic

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Methodic is seeking an experienced Site Reliability Engineer to join our San Francisco on-site team. You will own monitoring, alerting, and self-healing workflows, ensuring reliability and scalability for our platform.

You will implement capacity planning, blue-green or canary deployments, and collaborate with software engineers to embed resilience by design. Expect on-call rotations and ongoing automation to reduce toil and improve performance.

Qualifications

  • 5+ years in a Site Reliability Engineering or similar role.
  • Strong coding/scripting abilities (Python, Go, or other) for building automation and tooling.
  • Deep knowledge of systems monitoring and observability — experience with tools like Grafana, Datadog, Prometheus, etc., and the ability to interpret system metrics to spot problems.
  • Understanding of high-availability design and distributed systems principles (load balancing, consensus, graceful degradation, etc.).
  • Experience with incident management and a track record of improving systems based on lessons learned.
  • Performance tuning experience — ability to use profiling and stress testing tools to find and fix bottlenecks.
  • A collaborative mindset, capable of working with development teams to ensure reliability is built in from the start, not just after issues occur.

Responsibilities

  • Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers.
  • Perform regular capacity planning and load testing to ensure the platform can scale ahead of demand without performance degradation (validating that we can sustain extremely high transaction rates per tenant).
  • Improve deployment processes with strategies like blue-green or canary deployments to minimize risk and downtime during releases.
  • Collaborate with software engineers to design resilient architectures – for example, build redundancy and failover capabilities into critical services (so that even if one component fails, the system remains operational).
  • Participate in on-call rotations to respond to and resolve production incidents; lead blameless post-mortems to identify root causes and implement corrective actions.
  • Automate routine operational tasks (from simple scripts to more complex tooling) to reduce manual work and error potential.
  • Document reliability-related procedures and best practices, ensuring knowledge is shared and systems are well-understood by the team.

Skills

Python
Go
Scripting
Observability
Collaboration
On-call

Tools

Grafana
Datadog
Prometheus

Job description

Methodic is seeking an experienced Site Reliability Engineer to join our San Francisco on-site team. You will own monitoring, alerting, and self-healing workflows, ensuring reliability and scalability for our platform.

You will implement capacity planning, blue-green or canary deployments, and collaborate with software engineers to embed resilience by design. Expect on-call rotations and ongoing automation to reduce toil and improve performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior SRE & Software Engineer — Scalable Infra
Senior SRE & Software Engineer — Scalable Infra

Harvey • San Francisco (CA)

On-site
USD 200,000 - 260,000
Senior Site Reliability Engineer: Observability & Resiliency
Senior Site Reliability Engineer: Observability & Resiliency

Early Warning • San Francisco (CA)

Hybrid
USD 139,000 - 174,000
Discretionary incentive plan
Comprehensive benefits package
Staff SRE: Scale, Resilience & Observability
Staff SRE: Scale, Resilience & Observability

Early Warning Services LLC • San Francisco (CA)

On-site
USD 130,000 - 160,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Remote Senior SRE: Reliability, Security & Scale
Remote Senior SRE: Reliability, Security & Scale

WellSaid Labs, Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive salary
Stock options
Medical, dental, and vision insurance
+4
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Hybrid SRE Engineer for Scalable Reliability & Equity
Hybrid SRE Engineer for Scalable Reliability & Equity

EarnIn • Mountain View (CA)

Hybrid
USD 139,000 - 232,000
Senior SRE — Flexible, AI-Driven Reliability
Senior SRE — Flexible, AI-Driven Reliability

Salesforce, Inc. • San Francisco (CA)

Hybrid
USD 148,000 - 224,000
Senior SRE & Platform Engineer — Observability & Automation
Senior SRE & Platform Engineer — Observability & Automation

Techunting • United States

On-site
USD 120,000 - 150,000