Site Reliability Engineer – Resilience & Observability

Neurealm

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Neurealm in Sunnyvale, CA seeks a reliability engineer to lead incident response, design scalable systems, and drive RCA across teams. You will implement observability from the ground up and help prevent outages before they affect customers.

From reviewing overnight alerts to automating monitoring, you'll collaborate with software engineers to enforce resilient-code practices, review changes pre-deployment, and maintain high SLIs/SLOs with thorough documentation.

Qualifications

  • Architect scalable and reliable systems with emphasis on fault tolerance.
  • Lead incident command, drive RCA, and coordinate cross-functional teams.
  • Define observability, SLOs/SLIs, and alerting strategies.
  • Understand Linux kernel internals for performance tuning.
  • Deliver high-quality code in Python, Go, or similar languages.
  • Apply IaC and CI/CD practices for on-prem or cloud infra.
  • Competent with TCP/IP networking and distributed storage concepts.
  • Experience with storage paradigms (object/block/file) and related APIs.

Responsibilities

  • Runs incident response drills, post-mortems, and RCA to prevent recurrence.
  • Starts the day reviewing overnight alerts and system metrics, triaging anomalies.
  • Participates in stand-ups on projects, incidents, and daily priorities.
  • Automates routine processes, analyzes logs, and builds monitoring tools.
  • Collaborates with engineers on resilient-code practices and pre-deployment reviews.
  • Maintains high SLIs/SLOs and documents work for customer-centric insights.

Skills

Architecture patterns
Incident command
Observability design
Linux kernel internals
Python/Go coding
IaC & CI/CD
Networking (TCP/IP)
Distributed storage

Tools

Ansible
Terraform
Kubernetes
GitLab CI
AWX
Prometheus
ELK/Logging

Job description

Neurealm in Sunnyvale, CA seeks a reliability engineer to lead incident response, design scalable systems, and drive RCA across teams. You will implement observability from the ground up and help prevent outages before they affect customers.

From reviewing overnight alerts to automating monitoring, you'll collaborate with software engineers to enforce resilient-code practices, review changes pre-deployment, and maintain high SLIs/SLOs with thorough documentation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer: Observability & Resiliency
Senior Site Reliability Engineer: Observability & Resiliency

Early Warning • San Francisco (CA)

Hybrid
USD 139,000 - 174,000
Discretionary incentive plan
Comprehensive benefits package
Hybrid SRE Engineer for Scalable Reliability & Equity
Hybrid SRE Engineer for Scalable Reliability & Equity

EarnIn • Mountain View (CA)

Hybrid
USD 139,000 - 232,000
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Competitive salary
Equity
Wellness benefits
+2
Senior SRE: Build Resilient, Scalable Systems
Senior SRE: Build Resilient, Scalable Systems

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer - Scale & Observability
Senior Site Reliability Engineer - Scale & Observability

Early Warning Services LLC • United States

Hybrid
USD 118,000 - 183,000
Healthcare Coverage
401(k) Retirement Plan with company-mn
Paid Time Off
+1
Site Reliability Engineering Leader
Site Reliability Engineering Leader

Oracle • Reston (VA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer — Build Resilient Cloud Platforms
Senior Site Reliability Engineer — Build Resilient Cloud Platforms

Ridgeline • San Ramon (CA)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Educational reimbursements
Wellness reimbursements
+1
Senior Site Reliability Engineer: Scalable Infra & Observability
Senior Site Reliability Engineer: Scalable Infra & Observability

Early Warning • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Matching
Paid Time Off
+1
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer: Build Resilient Systems
Senior Site Reliability Engineer: Build Resilient Systems

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 130,000 - 180,000