Site Reliability Engineer III

Phase2 Technology

Newport News (VA)

On-site

USD 118,000 - 187,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision Plans
401(k) Plan with company contribution
Flexible Work Arrangements

Job summary

Jefferson Lab is seeking a Lead Site Reliability Engineer for the High Performance Data Facility (HPDF). You will supervise a small SRE team, design and operate monitoring and observability stacks, and own incident response and disaster recovery planning for a new, distributed research facility.

You will work with scientists and engineers, define SLOs/SLIs, and influence technology choices with a focus on reliability, scalability, and automation in a research environment.

Qualifications

  • Deep Linux systems expertise across app, OS, storage, and network layers.
  • Expertise in monitoring/observability stacks and SLOs/SLIs.
  • Strong scripting/automation in Python, Go, or shell.
  • Leadership of technical staff and performance feedback.
  • Incident response and operational process in production.
  • Resilience design: fault isolation, redundancy, graceful degradation, recovery.
  • Clear communications with scientific users and colleagues.
  • Familiarity with public cloud and large-scale HPC infra.
  • Ability to evaluate vendor/OSS solutions for reliability.

Responsibilities

  • Lead design/operation of monitoring, logging, alerting, and diagnostics for HPDF systems.
  • Supervise and mentor a small SRE team; assign, review, and plan work.
  • Establish on-call, incident management, change control, and maintenance records.
  • Design resilience, disaster recovery, and failure-domain isolation.
  • Define and report SLOs/SLIs with stakeholders; ensure service reliability.
  • Serve as incident commander; drive postmortems and preventive actions.
  • Drive reliability via automation and process optimization.
  • Collaborate on HPDF tech selection and vendor evaluations.

Skills

Linux systems
Observability
Technical leadership
Incident response
Resilience design
Cloud familiarity
Networking at scale
Cost estimation
AI automation

Education

Bachelor's Degree Computer Science or Related Field
Master's Degree Computer Science or Related Field

Tools

Prometheus
Grafana
ELK/OpenTelemetry
AWS/Azure/GCP
Terraform/Ansible
Kubernetes

Job description

At Jefferson Lab,you'llchampioncutting-edgescience and operational excellence while shaping the future of discovery. Join us and make your mark - where excellence meets purpose, andgreat mindstrulymatter.

The good-faith pay range for this role is $118,400 - $186,850 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What your job will be like:

As Lead Site Reliability Engineer on the High Performance Data Facility (HPDF) team, you will play a critical role in establishing and running the reliability practice for the facility's first systems on its path to operations. This is a technical role with manager responsibilities: you will supervise and develop a small team of site reliability engineers, and you will also design and build systems yourself, hands on in the code, the monitoring stack, and the incident response. You will design how the facility stays available and recovers, define and report on the service level objectives that measure how well it serves its users, serve as incident commander for significant incidents, and work day to day with staff at both Jefferson Lab and Berkeley Lab. HPDF is still in design, so there is room in this role to grow into influencing the technology choices the facility is built on. The users you support are research physicists and computational scientists, and helping them succeed is a core measure of this role.

In this job you will:
  • Lead the design, implementation, and operation of monitoring, logging, alerting, and diagnostic tooling for HPDF compute, storage, network, and facility systems, contributing directly to that work as well as directing it.
  • Supervise, mentor, and develop a small team of site reliability engineers: assign and review work, set expectations, give regular feedback, support technical growth, and plan and estimate the multi-person efforts assigned to the team.
  • Establish and maintain the facility's operational framework, including on-call and escalation structure, incident management, change management, and scheduled maintenance, and keep operational records, runbooks, and documentation current.
  • Design the facility's resilience model, including failure domain isolation, redundancy, graceful degradation, and disaster recovery objectives for a geographically distributed facility, and validate that design through testing.
  • Define, implement, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in collaboration with the architecture team and scientific stakeholders, and hold facility operations to them.
  • Serve as incident commander for significant incidents, own the postmortem process, and drive root cause prevention back into the design and operation of the systems.
  • Drive reliability improvement through automation, process optimization, and the elimination of manual operations, using Python, Go, or shell and standard software development practices.
  • Partner with the architecture team on HPDF technology selection from a reliability standpoint, lead evaluations of vendor and open source technologies, and represent HPDF site reliability engineering in the Berkeley Lab partnership.
Additional Responsibilities
  • Participate in an on-call rotation as the facility moves toward operations.
Lead - Supervisory - Management
  • Supervises a team of site reliability engineers
  • Assigns and reviews work, sets performance expectations, provides regular feedback, conducts performance discussions, and supports the technical development of the team.
  • Participates in hiring for the group. Does not hold fiscal or budget authority.
Experience
  • Required: 10 or more years experience in Site Reliability Engineering, DevOps, systems engineering, or operations engineering, including at least two years leading or supervising engineers. Technical leadership of engineering teams or projects qualifies.
  • Preferred: Supporting scientific computing, HPC, or research environments.
  • Preferred: Establishing operational practice in a new or greenfield facility.
  • Preferred: High availability or around the clock operations.
  • Preferred: Serving as the reliability or availability authority during the design phase of a large system or facility, before it entered operations.
  • Preferred: Evaluating vendor compute, storage, and network solutions against reliability requirements, including acceptance criteria and benchmarking.
  • Preferred: Experience with containers and Kubernetes
  • Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet).
  • Preferred: Experience with storage systems, data movement, or large scale data infrastructure.
  • Preferred: Experience with IT service management practice and tooling (for example ServiceNow, ITIL).
  • Preferred: Experience with HPC infrastructure and environments.
  • Preferred: Supporting formal project milestone or gate reviews, such as DOE critical decision reviews, and defining KPPs or acceptance criteria.
Education
  • Required: Bachelor's Degree Computer Science or Related Field
  • Preferred: Master's Degree Computer Science or Related Field
Experience and Education Exchange

Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.

Knowledge, Skills, and Abilities
  • Deep Linux systems expertise, with the ability to troubleshoot across the application, operating system, storage, and network layers and to guide others in doing so.
  • Expertise in designing and operating monitoring and observability stacks (for example Prometheus, Grafana, ELK, OpenTelemetry) and in defining and reporting on SLOs and SLIs.
  • Strong scripting and automation skills (Python, Go, or shell) with standard software development practices, together with the judgment to decide what is worth automating.
  • Demonstrated ability to lead and mentor technical staff: setting expectations, assigning work, giving constructive feedback, and addressing performance issues in a constructive manner.
  • Ability to establish and run incident response and operational process in a production environment.
  • Ability to design for resilience, including failure domain isolation, redundancy, graceful degradation, and recovery objectives, and to validate the design through testing and analysis.
  • Clear written and verbal communication, including the ability to present options, argue persuasively for proposals, and work productively with a community of scientific users and with colleagues across both laboratories.
  • Familiarity with public cloud environments (AWS, Azure, GCP).
  • Networking at scale: IPv4/IPv6, DNS, firewalls and access control lists, high speed interconnects, and data transfer protocols.
  • Ability to review system and vendor designs from a reliability standpoint and to argue a technical position persuasively with architects, vendors, and scientific stakeholders.
  • Load testing, performance analysis, and capacity modeling to validate design assumptions and identify bottlenecks in the data path.
  • Ability to estimate cost and effort for multi-person projects and to plan staffing accordingly.
  • Practical experience developing and deploying AI assisted or autonomous automation for operational work, and routine use of AI tools in day to day engineering.

About Jefferson Lab

Join a community with a common purpose of solving the most challenging scientific and engineering problems of our time. The Jefferson Lab campusis located insoutheasternVirginiaamidst a vibrant and growing technology community.

A career at Jefferson Lab is more than a job. You will be part of "big science" and work alongside top scientists and engineers from around the world unlocking the secrets of our visible universe. Managed by SURATech, LLC, Thomas Jefferson National Accelerator Facility is entering an exciting period of mission growth and is seeking new team members ready to apply their skills and passion to have an impact. You could call it work, or you could call it a mission. We call it a challenge. We do things that will change the world.

Total Rewards at Jefferson Lab

  • * Medical, Dental, and Vision Care Plans * Flexible Spending Accounts
  • * Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and sick leave)
  • * 401(k) Plan - 9% Lab Contribution; 100% vested * Flexible Work Arrangements
  • (Remote & Alternate Work Schedules available)
  • * Tuition Assistance, Training and Professional Development Programs
  • * Live near the waterways of the Chesapeake Bay region with access to nearby beaches,
  • mountains, and all major metropolitan centers on the East Coast

SURATech, LLC manages and operates the Thomas Jefferson National Accelerator Facility (Jefferson Lab). SURATech is an Equal Opportunity Employer.

Employment with SURATech is conditional upon DOE approval if at any time during your employment you are participating in a Foreign Government Talent Recruitment Program or Affiliated activity. Generally, such programs/activities include any foreign-state-sponsored attempt to acquire U.S.-funded scientific research through programs run or funded by the government that target scientists, engineers, students, academics, researchers, and entrepreneurs of all nationalities working or educated in the United States. This includes positions or appointments, both domestic and foreign, titled academic, professional, or institutional appointments whether or not remuneration is received and whether full-time, part-time or voluntary.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, Vision plans
Flexible work arrangements
Paid time off
High-Performance Computing (HPC) Systems Engineer
High-Performance Computing (HPC) Systems Engineer

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, Vision
Paid Time Off
401(k) Plan
+2
Scientific Software & DevOps Engineer
Scientific Software & DevOps Engineer

Jefferson Lab • Newport News (VA)

Hybrid
USD 92,000 - 145,000
Flexible Work Arrangements
Tuition Assistance
Paid Time Off
+2
HPDF Building Infrastructure Manager
HPDF Building Infrastructure Manager

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, and Vision
Flexible Spending Accounts
Paid Time-off, vacation, holidays, and
+5
DCS Data Scientist II - Data Steward
DCS Data Scientist II - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, and Vision Plans
Flexible Spending Accounts
Paid Time-off and Leave Programs
+2
Storage Architect
Storage Architect

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision
Flexible Work Arrangements
401(k) Plan
DCS Data Scientist III - Data Steward
DCS Data Scientist III - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical plan
Dental plan
Vision plan
+5
HVAC Technician II
HVAC Technician II

Jefferson Lab • Newport News (VA)

On-site
USD 58,000 - 84,000
Medical, Dental, and Vision Plans
Flexible Work Arrangements
Paid Time-off and Leave Programs
+3
Manager - PMO Project Controls
Manager - PMO Project Controls

Phase2 Technology • Newport News (VA)

Remote
USD 151,000 - 220,000
Medical, Dental, and Vision Plans
401(k) Plan with Lab contribution
Flexible Work Arrangements
+1
EI&C Technician I
EI&C Technician I

Phase2 Technology • Newport News (VA)

On-site
USD 45,000 - 71,000
Medical, Dental, and Vision Care Plans
Paid Time-off and Leave Programs
401(k) Plan - 9% Lab Contribution; 100