Site Reliability Engineer: Resilience & Observability

Castelion

Los Angeles (CA)

On-site

USD 140,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Castelion is hiring a Site Reliability Engineer to own the reliability, performance, observability, and health of critical engineering systems used for software and hardware delivery. You will collaborate across DevOps, Cloud, Software, Security, Test, and IT to diagnose failures and drive corrective actions to resolution.

You will understand and improve existing systems rather than replacing them, applying new technology where it solves reliability and scalability issues.

Qualifications

  • Strong Linux systems expertise including CPU, memory, storage, networking, processes, sockets, and services.
  • Experience building and operating observability, monitoring, alerting, and incident response systems.

Responsibilities

  • Establish reliability, availability, latency, capacity, and recovery metrics with meaningful health indicators and alerts.
  • Lead deep technical investigations and incident response across app, Linux, networking, storage, Kubernetes, cloud, and more; drive to resolution.

Skills

Linux proficiency
Observability
Networking fundamentals
Incident response
Root cause analysis
Cross-team collaboration
On-call experience

Education

Bachelor’s/Master’s/PhD in CS/EE

Tools

Kubernetes

Job description

Castelion is hiring a Site Reliability Engineer to own the reliability, performance, observability, and health of critical engineering systems used for software and hardware delivery. You will collaborate across DevOps, Cloud, Software, Security, Test, and IT to diagnose failures and drive corrective actions to resolution.

You will understand and improve existing systems rather than replacing them, applying new technology where it solves reliability and scalability issues.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: Reliability & Observability
Site Reliability Engineer: Reliability & Observability

Castelion • Allen (TX)

On-site
USD 140,000 - 190,000
Generous benefits package
Site Reliability Engineer
Site Reliability Engineer

Castelion • Allen (TX)

On-site
USD 140,000 - 190,000
Generous benefits package
Site Reliability Engineer
Site Reliability Engineer

Castelion • Los Angeles (CA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer: Cloud, CI/CD & Observability
Senior Site Reliability Engineer: Cloud, CI/CD & Observability

Castleton Commodities International, LLC • United States

On-site
USD 160,000 - 260,000
Medical & Dental
Pension Plan
Tuition assistance
+2
Senior SRE: Cloud Reliability, IaC & DR Leader
Senior SRE: Cloud Reliability, IaC & DR Leader

Castleton Commodities International • Stamford (CT)

On-site
USD 150,000 - 190,000
Competitive compensation
Comprehensive benefits
Tuition assistance
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Senior Site Reliability Engineer: Scalable Infra & Observability
Senior Site Reliability Engineer: Scalable Infra & Observability

Early Warning • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Matching
Paid Time Off
+1
Site Reliability Engineer: Cloud Platform & Resilience
Site Reliability Engineer: Cloud Platform & Resilience

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer – Observability & Cloud
Senior Site Reliability Engineer – Observability & Cloud

Cosm Inc. • El Segundo (CA), Northern (KY)

Hybrid
USD 110,000 - 145,000