Lead Site Reliability Engineer - HPC Data Facility (Remote)

Phase2 Technology

Newport News (VA)

On-site

USD 118,000 - 187,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, Dental, Vision Plans
401(k) Plan with company contribution
Flexible Work Arrangements

Job summary

Jefferson Lab is seeking a Lead Site Reliability Engineer for the High Performance Data Facility (HPDF). You will supervise a small SRE team, design and operate monitoring and observability stacks, and own incident response and disaster recovery planning for a new, distributed research facility.

You will work with scientists and engineers, define SLOs/SLIs, and influence technology choices with a focus on reliability, scalability, and automation in a research environment.

Qualifications

  • Deep Linux systems expertise across app, OS, storage, and network layers.
  • Expertise in monitoring/observability stacks and SLOs/SLIs.
  • Strong scripting/automation in Python, Go, or shell.
  • Leadership of technical staff and performance feedback.
  • Incident response and operational process in production.
  • Resilience design: fault isolation, redundancy, graceful degradation, recovery.
  • Clear communications with scientific users and colleagues.
  • Familiarity with public cloud and large-scale HPC infra.
  • Ability to evaluate vendor/OSS solutions for reliability.

Responsibilities

  • Lead design/operation of monitoring, logging, alerting, and diagnostics for HPDF systems.
  • Supervise and mentor a small SRE team; assign, review, and plan work.
  • Establish on-call, incident management, change control, and maintenance records.
  • Design resilience, disaster recovery, and failure-domain isolation.
  • Define and report SLOs/SLIs with stakeholders; ensure service reliability.
  • Serve as incident commander; drive postmortems and preventive actions.
  • Drive reliability via automation and process optimization.
  • Collaborate on HPDF tech selection and vendor evaluations.

Skills

Linux systems
Observability
Technical leadership
Incident response
Resilience design
Cloud familiarity
Networking at scale
Cost estimation
AI automation

Education

Bachelor's Degree Computer Science or Related Field
Master's Degree Computer Science or Related Field

Tools

Prometheus
Grafana
ELK/OpenTelemetry
AWS/Azure/GCP
Terraform/Ansible
Kubernetes

Job description

Jefferson Lab is seeking a Lead Site Reliability Engineer for the High Performance Data Facility (HPDF). You will supervise a small SRE team, design and operate monitoring and observability stacks, and own incident response and disaster recovery planning for a new, distributed research facility.

You will work with scientists and engineers, define SLOs/SLIs, and influence technology choices with a focus on reliability, scalability, and automation in a research environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer, HPC Data Facility
Lead Site Reliability Engineer, HPC Data Facility

Jefferson Lab • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, and Vision Care Plans
401(k) Plan – 9% Lab Contribution; 100
Flexible Work Arrangements
+1
Remote-Eligible Data Center Infrastructure Lead
Remote-Eligible Data Center Infrastructure Lead

Jefferson Lab • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, Vision Plans
401(k) Plan – Lab Contribution
Flexible Work Arrangements
+1
High-Density Data Center Infrastructure Manager
High-Density Data Center Infrastructure Manager

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, and Vision
Flexible Spending Accounts
Paid Time-off, vacation, holidays, and
+5
Site Reliability Engineer III
Site Reliability Engineer III

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision Plans
401(k) Plan with company contribution
Flexible Work Arrangements
Remote Project Director, DOE Research Computing
Remote Project Director, DOE Research Computing

Jefferson Lab • Newport News (VA)

Hybrid
USD 169,000 - 267,000
Medical insurance
Dental insurance
Vision care
+5
Assoc Lab Director, Data & Computational Science (Remote)
Assoc Lab Director, Data & Computational Science (Remote)

Jefferson Lab • Newport News (VA)

Hybrid
USD 150,000 - 200,000
Medical, Dental, and Vision Care Plans
Flexible Spending Accounts
Paid Time-off and Leave Programs
+3
Site Reliability Engineer III
Site Reliability Engineer III

Jefferson Lab • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, and Vision Care Plans
401(k) Plan – 9% Lab Contribution; 100
Flexible Work Arrangements
+1
High-Reliability Safety Systems Engineer — Remote Options
High-Reliability Safety Systems Engineer — Remote Options

Phase2 Technology • Newport News (VA)

On-site
USD 76,000 - 102,000
Medical, Dental, and Vision Plans
401(k) Plan - Lab Contribution; 100% v
Flexible Work Arrangements
HPDF Building Infrastructure Manager
HPDF Building Infrastructure Manager

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, and Vision
Flexible Spending Accounts
Paid Time-off, vacation, holidays, and
+5
HPC Site Reliability Engineer — Onsite 24/7
HPC Site Reliability Engineer — Onsite 24/7

Bay Systems Consulting Inc. • Berkeley (CA)

On-site
USD 83,000 - 166,000