Site Reliability Engineer

Castelion

Allen (TX)

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Generous benefits package

Job summary

Castelion seeks a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.

This role is the missing reliability piece of an existing high-performing engineering organization.

Qualifications

  • 5+ years in Site Reliability Engineering or related field.
  • Strong Linux systems expertise and performance concepts such as IOPS, throughput, latency, queue depth, and connection concurrency.
  • Experience building observability, monitoring, alerting, and incident response systems.
  • Strong networking and application fundamentals (TCP, TLS, HTTP, DNS, proxies, load balancers).
  • Ability to collaborate across engineering teams and drive root cause analysis and corrective actions.

Responsibilities

  • Establish reliability, availability, latency, capacity, and recovery expectations with meaningful health indicators.
  • Lead technical investigations and incident response across multi-layer systems; collect evidence and drive resolution.
  • Build and improve monitoring/diagnostic systems to detect problems early.
  • Analyze system performance across compute, memory, storage, network; identify limits and address them.
  • Drive root cause analysis and postmortems with preventive actions implemented.
  • Partner with DevOps, Cloud, Software, Security, Test, and IT to resolve cross-team issues.
  • Operate and incrementally improve systems built by other engineers.
  • Participate in on-call rotation for critical services.

Skills

Linux systems
Observability & monitoring
Networking fundamentals
Incident response & RCAs
Cross-functional collaboration

Education

Bachelor's degree in CS/CE
Master's degree in CS/CE
PhD in a related field

Job description

Why Castelion, Why Now

Castelion is moving incredibly fast to develop and deliver advanced defense systems at a time when execution matters more than ever. We believe focus, ownership, and excellence are decisive advantages - and we're building a world-class team to turn bold ideas into real capability.

This is a rare opportunity to join at an early stage, where you'll have significant ownership, collaborate with exceptional teammates, and make a direct, measurable impact on our mission and the future of the company - regardless of your function.

Site Reliability Engineer

We are seeking a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.

This role is the missing reliability piece of an existing high-performing engineering organization. You will work across DevOps, Cloud, Software, Security, Test, and IT to identify reliability risks, diagnose failures that cross system boundaries, and drive corrective actions to resolution. You will be expected to understand and improve existing systems rather than defaulting to replacement, using new technology when it solves a demonstrated reliability, scalability, or operational problem.

Responsibilities
  • Establish meaningful reliability, availability, latency, capacity, and recovery expectations for critical engineering services, with measurable health indicators and useful alerts.
  • Lead deep technical investigations and incident response across application, Linux, networking, storage, Kubernetes, cloud, and other system boundaries; collect evidence, separate symptoms from root causes, and drive incidents through resolution.
  • Build and improve monitoring and diagnostic systems that detect problems before users report them and provide engineers with the information needed to quickly understand and resolve failures.
  • Analyze system performance and capacity across compute, memory, storage, networking, connections, and other constrained resources; identify operating limits and address issues through the simplest effective solution, whether optimization, additional capacity, scaling, caching, configuration changes, or architectural improvements.
  • Drive evidence-backed root cause analysis and postmortem actions for significant incidents, ensuring corrective and preventive actions are implemented and verified to reduce recurring failures.
  • Partner with DevOps, Cloud, Software, Security, Test, and IT to resolve reliability problems that cross team boundaries, providing technical leadership without attempting to own every component involved.
  • Understand, operate, and incrementally improve systems built by other engineers, balancing reliability and operational value against existing architecture, constraints, and engineering practices.
  • Participate in the on-call rotation for critical engineering services, providing first-response triage, escalation, and follow-up for recurring reliability issues.
Basic Qualifications
  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a related technical field.
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, Systems Engineering, Infrastructure Engineering, or a related discipline supporting production or mission-critical systems.
  • Demonstrated experience debugging complex production failures across multiple system layers and driving investigations from the first symptom to an evidence-backed root cause and lasting corrective action.
  • Strong Linux systems expertise, including CPU, memory, storage, networking, processes, sockets, and system services, with a strong understanding of performance and capacity concepts such as IOPS, throughput, latency, queue depth, and connection concurrency.
  • Experience building and operating observability, monitoring, alerting, and incident response systems, with the ability to distinguish between mitigation, workaround, corrective action, and preventive action.
  • Strong networking and application fundamentals, including TCP, TLS, HTTP, DNS, reverse proxies, load balancers, connection states, and timeouts; able to investigate application runtime behavior such as threads, connection pools, file descriptors, memory, or garbage collection when the evidence points there.
  • Demonstrated ability to work effectively within existing systems and across engineering organizations, asking why a system was designed a certain way and improving it based on measurable reliability and operational needs rather than defaulting to rewrites or replacement.

Castelion offers a generous benefits package. Please refer to the bottom of our Careers page for more details.

Other Duties

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice.

Additional Eligibility Requirements

This position may require access to classified information or restricted U.S. Government sites, systems, or information, as determined by the Company and/or applicable U.S. Government requirements. If the position is so designated, your employment in the role may be contingent upon your ability to obtain and maintain the required U.S. Government security clearance or other government authorization, and to satisfy any citizenship or other eligibility requirements imposed by applicable law, regulation, executive order, or government contract requirements. You will be notified if and when such requirements apply.

Affirmative Action/EEO Statement

Castelion is an Equal Opportunity Employer. We are committed to providing equal employment opportunities to all applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.

Castelion is committed to providing reasonable accommodations to qualified individuals with disabilities throughout the application and hiring process. If you require a reasonable accommodation to complete an application, participate in the interview process, or otherwise participate in the hiring process, please contact hr@castelion.com. Requests for accommodation will be considered on an individual basis and handled in accordance with applicable law.

Castelion is committed to fostering a workplace where employment decisions are based on qualifications, business needs, and the ability to perform the essential functions of the role, with or without reasonable accommodation.

EAR/ITAR Requirements

This position requires access to export-controlled information, and as such, employment (or hiring of a contractor) is contingent upon the candidate’s ability to access all applicable export-controlled information without additional export licensing being required by the Bureau of Industry and Security and/or the Directorate of Defense Trade Controls.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Castelion • Los Angeles (CA)

On-site
USD 140,000 - 180,000
Principal Build Reliability Engineer
Principal Build Reliability Engineer

Next Matter • Torrance (CA)

On-site
USD 120,000 - 160,000
Employee Equity
PTO 4 weeks
Holidays 10
+9
Avionics Technician Specialist
Avionics Technician Specialist

Castelion • Los Angeles (CA)

On-site
USD 65,000 - 90,000
Generous benefits package
Embedded Software Engineer (All Levels)
Embedded Software Engineer (All Levels)

Castelion • Torrance (CA)

On-site
USD 110,000 - 150,000
Avionics Reliability Engineer
Avionics Reliability Engineer

Castelion • Torrance (CA)

On-site
USD 120,000 - 180,000
EHS Technician
EHS Technician

Castelion • Torrance (CA)

On-site
USD 65,000 - 90,000
Benefits package
Lead R&D Technician
Lead R&D Technician

Next Matter • Torrance (CA)

On-site
USD 70,000 - 90,000
Employee Equity
Four weeks paid time off
Company paid holidays
+7
Technical Recruiter
Technical Recruiter

Castelion • Torrance (CA)

On-site
USD 70,000 - 110,000
Build Reliability Engineer
Build Reliability Engineer

Castelion • Torrance (CA)

On-site
USD 110,000 - 160,000
Stock incentives
Medical, Vision, Dental insurance
4 weeks PTO
Facilities Technician
Facilities Technician

Castelion • Torrance (CA)

On-site
USD 52,000 - 70,000
Benefits package