Site Reliability Engineer: Reliability & Observability

Castelion

Allen (TX)

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Generous benefits package

Job summary

Castelion seeks a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.

This role is the missing reliability piece of an existing high-performing engineering organization.

Qualifications

  • 5+ years in Site Reliability Engineering or related field.
  • Strong Linux systems expertise and performance concepts such as IOPS, throughput, latency, queue depth, and connection concurrency.
  • Experience building observability, monitoring, alerting, and incident response systems.
  • Strong networking and application fundamentals (TCP, TLS, HTTP, DNS, proxies, load balancers).
  • Ability to collaborate across engineering teams and drive root cause analysis and corrective actions.

Responsibilities

  • Establish reliability, availability, latency, capacity, and recovery expectations with meaningful health indicators.
  • Lead technical investigations and incident response across multi-layer systems; collect evidence and drive resolution.
  • Build and improve monitoring/diagnostic systems to detect problems early.
  • Analyze system performance across compute, memory, storage, network; identify limits and address them.
  • Drive root cause analysis and postmortems with preventive actions implemented.
  • Partner with DevOps, Cloud, Software, Security, Test, and IT to resolve cross-team issues.
  • Operate and incrementally improve systems built by other engineers.
  • Participate in on-call rotation for critical services.

Skills

Linux systems
Observability & monitoring
Networking fundamentals
Incident response & RCAs
Cross-functional collaboration

Education

Bachelor's degree in CS/CE
Master's degree in CS/CE
PhD in a related field

Job description

Castelion seeks a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.

This role is the missing reliability piece of an existing high-performing engineering organization.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: Resilience & Observability
Site Reliability Engineer: Resilience & Observability

Castelion • Los Angeles (CA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Castelion • Allen (TX)

On-site
USD 140,000 - 190,000
Generous benefits package
Site Reliability Engineer
Site Reliability Engineer

Castelion • Los Angeles (CA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer: Cloud, CI/CD & Observability
Senior Site Reliability Engineer: Cloud, CI/CD & Observability

Castleton Commodities International, LLC • United States

On-site
USD 160,000 - 260,000
Medical & Dental
Pension Plan
Tuition assistance
+2
Associate Build Reliability Engineer — Early Impact
Associate Build Reliability Engineer — Early Impact

Castelion • Torrance (CA)

On-site
USD 90,000 - 120,000
Employee Equity
Paid Time Off
Medical
+5
Senior Site Reliability Engineer - Observability
Senior Site Reliability Engineer - Observability

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 140,000 - 190,000
Senior SRE: Scale Reliability & Observability
Senior SRE: Scale Reliability & Observability

Megaport • Abbeyville (CO)

On-site
USD 130,000 - 190,000
Contractor (PJ)
Paid Time Off
Competitive Compensation
+4
Senior Site Reliability Engineer: Scalable Infra & Observability
Senior Site Reliability Engineer: Scalable Infra & Observability

Early Warning • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Matching
Paid Time Off
+1
Senior Site Reliability Engineer – Observability & Cloud
Senior Site Reliability Engineer – Observability & Cloud

Cosm Inc. • El Segundo (CA), Northern (KY)

Hybrid
USD 110,000 - 145,000
Impactful Build Reliability Engineer - Equity & Quality
Impactful Build Reliability Engineer - Equity & Quality

Castelion Corporation • Torrance (CA)

On-site
USD 150,000 - 230,000
Employee equity
Paid time off
Medical coverage
+4