Site Reliability Engineer III

Jefferson Lab

Newport News (VA)

On-site

USD 118,000 - 187,000

Full time

30 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, Dental, and Vision Care Plans
401(k) Plan – 9% Lab Contribution; 100
Flexible Work Arrangements
Remote & Alternate Work Schedules

Job summary

Jefferson Lab seeks a Lead Site Reliability Engineer for the High Performance Data Facility (HPDF) team in Virginia. You will supervise a small SRE group, design and build reliable systems, and own incident response and monitoring for a greenfield facility.

You’ll define SLOs/SLIs, manage on-call and change processes, and collaborate with architecture and Berkeley Lab to influence technology choices. 10+ years in SRE/DevOps and strong automation skills are required.

Qualifications

  • 10+ years in SRE/DevOps or related operations engineering.
  • Experience leading or supervising engineers for multiple projects.
  • Strong expertise in monitoring, observability stacks, and incident management.

Responsibilities

  • Lead design, implementation, and operation of monitoring and diagnostic tooling for HPDF systems.
  • Supervise and mentor a small team of site reliability engineers.
  • Establish on-call, escalation, change management, and maintenance processes.
  • Design resilience, disaster recovery, and SLO/SLI reporting for a distributed facility.
  • Act as incident commander and drive postmortems and preventive improvements.
  • Drive reliability improvements through automation using Python, Go, or shell.

Skills

Site Reliability Engineering
Technical Leadership
Python
Go
Shell scripting
Incident Response

Education

Bachelor's Degree in Computer Science or related field
Master's Degree in Computer Science or related field

Tools

Prometheus
Grafana
ELK
OpenTelemetry
Kubernetes
Ansible
Terraform
Puppet
Public cloud (AWS, Azure, GCP)

Job description

At Jefferson Lab, you’ll champion cutting-edge science and operational excellence while shaping the future of discovery. Join us and make your mark – where excellence meets purpose, and great mindstruly matter.

The good-faith pay range for this role is $118,400 - $186,850 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What Your Job Will Be Like

As Lead Site Reliability Engineer on the High Performance Data Facility (HPDF) team, you will play a critical role in establishing and running the reliability practice for the facility's first systems on its path to operations. This is a technical role with manager responsibilities: you will supervise and develop a small team of site reliability engineers, and you will also design and build systems yourself, hands on in the code, the monitoring stack, and the incident response. You will design how the facility stays available and recovers, define and report on the service level objectives that measure how well it serves its users, serve as incident commander for significant incidents, and work day to day with staff at both Jefferson Lab and Berkeley Lab. HPDF is still in design, so there is room in this role to grow into influencing the technology choices the facility is built on. The users you support are research physicists and computational scientists, and helping them succeed is a core measure of this role.

In This Job You Will
  • Lead the design, implementation, and operation of monitoring, logging, alerting, and diagnostic tooling for HPDF compute, storage, network, and facility systems, contributing directly to that work as well as directing it.
  • Supervise, mentor, and develop a small team of site reliability engineers: assign and review work, set expectations, give regular feedback, support technical growth, and plan and estimate the multi-person efforts assigned to the team.
  • Establish and maintain the facility's operational framework, including on-call and escalation structure, incident management, change management, and scheduled maintenance, and keep operational records, runbooks, and documentation current.
  • Design the facility's resilience model, including failure domain isolation, redundancy, graceful degradation, and disaster recovery objectives for a geographically distributed facility, and validate that design through testing.
  • Define, implement, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in collaboration with the architecture team and scientific stakeholders, and hold facility operations to them.
  • Serve as incident commander for significant incidents, own the postmortem process, and drive root cause prevention back into the design and operation of the systems.
  • Drive reliability improvement through automation, process optimization, and the elimination of manual operations, using Python, Go, or shell and standard software development practices.
  • Partner with the architecture team on HPDF technology selection from a reliability standpoint, lead evaluations of vendor and open source technologies, and represent HPDF site reliability engineering in the Berkeley Lab partnership.
Additional Responsibilities
  • Participate in an on-call rotation as the facility moves toward operations.
Lead - Supervisory - Management
  • Supervises a team of site reliability engineers
  • Assigns and reviews work, sets performance expectations, provides regular feedback, conducts performance discussions, and supports the technical development of the team.
  • Participates in hiring for the group. Does not hold fiscal or budget authority.
Experience
  • Required: 10 or more years experience in Site Reliability Engineering, DevOps, systems engineering, or operations engineering, including at least two years leading or supervising engineers. Technical leadership of engineering teams or projects qualifies.
  • Preferred: Supporting scientific computing, HPC, or research environments.
  • Preferred: Establishing operational practice in a new or greenfield facility.
  • Preferred: High availability or around the clock operations.
  • Preferred: Serving as the reliability or availability authority during the design phase of a large system or facility, before it entered operations.
  • Preferred: Evaluating vendor compute, storage, and network solutions against reliability requirements, including acceptance criteria and benchmarking.
  • Preferred: Experience with containers and Kubernetes
  • Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet).
  • Preferred: Experience with storage systems, data movement, or large scale data infrastructure.
  • Preferred: Experience with IT service management practice and tooling (for example ServiceNow, ITIL).
  • Preferred: Experience with HPC infrastructure and environments.
  • Preferred: Supporting formal project milestone or gate reviews, such as DOE critical decision reviews, and defining KPPs or acceptance criteria.
Education
  • Required: Bachelor's Degree Computer Science or Related Field
  • Preferred: Master's Degree Computer Science or Related Field
Experience And Education Exchange

Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.

Knowledge, Skills, And Abilities
  • Deep Linux systems expertise, with the ability to troubleshoot across the application, operating system, storage, and network layers and to guide others in doing so.
  • Expertise in designing and operating monitoring and observability stacks (for example Prometheus, Grafana, ELK, OpenTelemetry) and in defining and reporting on SLOs and SLIs.
  • Strong scripting and automation skills (Python, Go, or shell) with standard software development practices, together with the judgment to decide what is worth automating.
  • Demonstrated ability to lead and mentor technical staff: setting expectations, assigning work, giving constructive feedback, and addressing performance issues in a constructive manner.
  • Ability to establish and run incident response and operational process in a production environment.
  • Ability to design for resilience, including failure domain isolation, redundancy, graceful degradation, and recovery objectives, and to validate the design through testing and analysis.
  • Clear written and verbal communication, including the ability to present options, argue persuasively for proposals, and work productively with a community of scientific users and with colleagues across both laboratories.
  • Familiarity with public cloud environments (AWS, Azure, GCP).
  • Networking at scale: IPv4/IPv6, DNS, firewalls and access control lists, high speed interconnects, and data transfer protocols.
  • Ability to review system and vendor designs from a reliability standpoint and to argue a technical position persuasively with architects, vendors, and scientific stakeholders.
  • Load testing, performance analysis, and capacity modeling to validate design assumptions and identify bottlenecks in the data path.
  • Ability to estimate cost and effort for multi-person projects and to plan staffing accordingly.
  • Practical experience developing and deploying AI assisted or autonomous automation for operational work, and routine use of AI tools in day to day engineering.
About Jefferson Lab

Join a community with a common purpose of solving the most challenging scientific and engineering problems of our time. The Jefferson Lab campus is located in southeastern Virginia amidst a vibrant and growing technology community.

A career at Jefferson Lab is more than a job. You will be part of “big science” and work alongside top scientists and engineers from around the world unlocking the secrets of our visible universe. Managed by SURATech, LLC, Thomas Jefferson National Accelerator Facility is entering an exciting period of mission growth and is seeking new team members ready to apply their skills and passion to have an impact. You could call it work, or you could call it a mission. We call it a challenge. We do things that will change the world.

Total Rewards at Jefferson Lab
Benefits
  • Medical, Dental, and Vision Care Plans
  • Flexible Spending Accounts
  • Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and sick leave)
  • 401(k) Plan – 9% Lab Contribution; 100% vested
  • Flexible Work Arrangements

(Remote & Alternate Work Schedules available)

  • Tuition Assistance, Training and Professional Development Programs
  • Live near the waterways of the Chesapeake Bay region with access to nearby beaches, mountains, and all major metropolitan centers on the East Coast

SURATech, LLC manages and operates the Thomas Jefferson National Accelerator Facility (Jefferson Lab). SURATech is an Equal Opportunity Employer.

SURATech is committed to providing reasonable accommodation for people with disabilities (unless doing so will result in an undue hardship). If you need a reasonable accommodation for any part of the employment process, please send an e-mail to recruiting@jlab.org or contact Human Resources by calling (757) 269-7100 and selecting option 1 between 8 am – 5 pm EST to provide the nature of your request.

Employment with SURATech is conditional upon DOE approval if at any time during your employment you are participating in a Foreign Government Talent Recruitment Program or Affiliated activity. Generally, such programs/activities include any foreign-state-sponsored attempt to acquire U.S.-funded scientific research through programs run or funded by the government that target scientists, engineers, students, academics, researchers, and entrepreneurs of all nationalities working or educated in the United States. This includes positions or appointments, both domestic and foreign, titled academic, professional, or institutional appointments whether or not remuneration is received and whether full-time, part-time or voluntary.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer III
Site Reliability Engineer III

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision Plans
401(k) Plan with company contribution
Flexible Work Arrangements
High-Performance Computing (HPC) Systems Engineer
High-Performance Computing (HPC) Systems Engineer

Jefferson Lab • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, Vision plans
Flexible spending accounts
Paid time off and parental leave
+3
HPDF Building Infrastructure Manager
HPDF Building Infrastructure Manager

Jefferson Lab • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, Vision Plans
401(k) Plan – Lab Contribution
Flexible Work Arrangements
+1
High-Performance Computing (HPC) Systems Engineer
High-Performance Computing (HPC) Systems Engineer

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, Vision
Paid Time Off
401(k) Plan
+2
DCS Data Scientist II - Data Steward
DCS Data Scientist II - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, and Vision Plans
Flexible Spending Accounts
Paid Time-off and Leave Programs
+2
DCS Data Scientist III - Data Steward
DCS Data Scientist III - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical plan
Dental plan
Vision plan
+5
Storage Architect
Storage Architect

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision
Flexible Work Arrangements
401(k) Plan
Facilities Mechanical Engineer II
Facilities Mechanical Engineer II

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 132,000
Medical, Dental, Vision
Flexible Spending Accounts
Paid Time-off & Leave
+3
HPDF Building Infrastructure Manager
HPDF Building Infrastructure Manager

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, and Vision
Flexible Spending Accounts
Paid Time-off, vacation, holidays, and
+5
Safety Systems Engineer
Safety Systems Engineer

Phase2 Technology • Newport News (VA)

On-site
USD 76,000 - 102,000
Medical, Dental, and Vision Plans
401(k) Plan - Lab Contribution; 100% v
Flexible Work Arrangements