Site Reliability Engineer II

Phase2 Technology

Newport News (VA)

On-site

USD 92,000 - 145,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision plans
Flexible work arrangements
Paid time off

Job summary

Jefferson Lab in Newport News, VA, seeks a Site Reliability Engineer for the High Performance Data Facility (HPDF) team. You will build monitoring, logging, and automation for HPDF systems, participate in incident response, and define service level objectives with researchers and engineers.

You will work with Jefferson Lab and Berkeley Lab staff to support scientific users, develop tooling in Python/Go, and ensure reliability across HPC environments.

Qualifications

  • Required: 3+ years related experience in Site Reliability Engineering, DevOps, systems engineering, or software engineering with operational responsibility.
  • Preferred: 3+ years experience supporting scientific computing, HPC, or research environments; experience operating systems in support of users outside the engineer's own team
  • Preferred: Experience with containers and Kubernetes
  • Preferred: Practical experience applying AI assisted or autonomous automation to operational tasks, and routine use of AI tools in day to day engineering work.
  • Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet)

Responsibilities

  • Build and operate monitoring, logging, and alerting for HPDF systems using observability tooling (Prometheus, Grafana, OpenTelemetry, ELK).
  • Develop automation and tooling in Python, Go, or shell to reduce manual operations.
  • Respond to incidents and author runbooks, postmortems, and operational documentation for improvements.
  • Implement and report on SLOs and SLIs defined with architectural and scientific stakeholders.
  • Support researchers by translating their needs into reliability requirements.
  • Collaborate with Jefferson Lab and Berkeley Lab staff to share tooling and practices.
  • Conduct testing and performance analysis to validate reliability and identify bottlenecks.
  • Contribute to evaluations of vendor/open source technologies against reliability, performance and security needs.

Skills

Linux
Scripting (Python/Go)
Incident response
Public cloud (AWS/Azure/GCP)
Observability
Autonomous automation

Education

Bachelor's Degree Computer Science or Related Field
Master's Degree Computer Science or Related Field

Tools

Kubernetes
Ansible
Terraform
Puppet
OpenTelemetry
Prometheus
Grafana
ELK

Job description

At Jefferson Lab,you'llchampioncutting-edgescience and operational excellence while shaping the future of discovery. Join us and make your mark - where excellence meets purpose, andgreat mindstruly matter.

The good-faith pay range for this role is $91,800 - $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What your job will be like:

As a Site Reliability Engineer on the High Performance Data Facility (HPDF) team, you will help build and operate the facility's first systems on its path to operations. You will create the monitoring, alerting, and automation that the full facility will eventually run on, participate in incident response, and help define and report on the service level objectives that measure how well the facility serves its users. You will work as part of a small site reliability engineering team, with day to day direction from the team's lead, and alongside staff at both Jefferson Lab and Berkeley Lab. The users you support are research physicists and computational scientists, and helping them succeed is a core measure of this role.

In this job you will:
  • Build and operate monitoring, logging, and alerting for HPDF systems using modern observability tooling (for example Prometheus, Grafana, OpenTelemetry, ELK), within guidelines set with the lead site reliability engineer and senior staff.
  • Develop automation and tooling in Python, Go, or shell that eliminates manual operations and reduces operational risk, following standard software development practices including version control, code review, and testing.
  • Respond to incidents and author the runbooks, postmortems, and operational documentation that convert each incident into a lasting improvement to facility reliability.
  • Implement and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) defined with the architecture team and scientific stakeholders.
  • Support expert scientific users by helping research physicists and computational scientists understand system behavior and by translating their needs into reliability requirements.
  • Collaborate with colleagues at both Jefferson Lab and Berkeley Lab, sharing tooling, reviews, and operational practice across the HPDF partnership.
  • Conduct testing and performance analysis to validate reliability and resilience decisions and to identify bottlenecks.
Additional Responsibilities
  • Contribute to evaluations of vendor and open source technologies against reliability, performance, and security requirements.
  • Participate in an on-call rotation as the facility moves toward operations.
Experience
  • Required: 3 or more years related experience in Site Reliability Engineering, DevOps, systems engineering, or software engineering with operational responsibility.
  • Preferred: 3 or more years experience supporting scientific computing, HPC, or research environments; experience operating systems in support of users outside the engineer's own team
  • Preferred: Experience with containers and Kubernetes
  • Preferred: Practical experience applying AI assisted or autonomous automation to operational tasks, and routine use of AI tools in day to day engineering work.
  • Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet).
Education
  • Required: Bachelor's Degree Computer Science or Related Field
  • Preferred: Master's Degree Computer Science or Related Field
Experience and Education Exchange

Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.

Knowledge, Skills, and Abilities
  • Solid Linux systems skills and command line fluency, with the ability to troubleshoot across the application, operating system, and network layers.
  • Scripting and automation ability in Python, Go, or shell, including familiarity with standard software development practices.
  • Working knowledge of monitoring and observability tools (for example Prometheus, Grafana, ELK, OpenTelemetry) and strong motivation to deepen that expertise in a new facility environment.
  • Clear written and verbal communication, and the ability to work productively with a community of scientific users and with colleagues across both laboratories.
  • A self starter's approach to learning new technologies and to identifying and solving problems within an assigned scope.
  • Familiarity with public cloud environments (AWS, Azure, GCP).
  • Networking fundamentals (IPv4/IPv6, DNS, firewalls, access control lists) and security conscious operational habits.
  • Exposure to HPC or scientific computing environments.
About Jefferson Lab

Join a community with a common purpose of solving the most challenging scientific and engineering problems of our time. The Jefferson Lab campusis located insoutheasternVirginiaamidst a vibrant and growing technology community.

A career at Jefferson Lab is more than a job. You will be part of "big science" and work alongside top scientists and engineers from around the world unlocking the secrets of our visible universe. Managed by SURATech, LLC, Thomas Jefferson National Accelerator Facility is entering an exciting period of mission growth and is seeking new team members ready to apply their skills and passion to have an impact. You could call it work, or you could call it a mission. We call it a challenge. We do things that will change the world.

Total Rewards at Jefferson Lab

At Jefferson Lab, we believe that a comprehensive employee benefits program is an important and meaningful part of the compensation employees receive. Our benefits program includes, but is not limited to:

  • Medical, Dental, and Vision Care Plans
  • Flexible Spending Accounts
  • Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and sick leave)
  • 401(k) Plan - 9% Lab Contribution; 100% vested
  • Flexible Work Arrangements
  • (Remote & Alternate Work Schedules available)
  • Tuition Assistance, Training and Professional Development Programs
  • Live near the waterways of the Chesapeake Bay region with access to nearby beaches,
  • mountains, and all major metropolitan centers on the East Coast

SURATech, LLC manages and operates the Thomas Jefferson National Accelerator Facility (Jefferson Lab). SURATech is an Equal Opportunity Employer.

SURATech is committed to providing reasonable accommodation for people with disabilities (unless doing so will result in an undue hardship). If you need a reasonable accommodation for any part of the employment process, please send an e-mail to recruiting@jlab.org or contact Human Resources by calling (757) 269-7100 and selecting option 1 between 8 am - 5 pm EST to provide the nature of your request.

Employment with SURATech is conditional upon DOE approval if at any time during your employment you are participating in a Foreign Government Talent Recruitment Program or Affiliated activity. Generally, such programs/activities include any foreign-state-sponsored attempt to acquire U.S.-funded scientific research through programs run or funded by the government that target scientists, engineers, students, academics, researchers, and entrepreneurs of all nationalities working or educated in the United States. This includes positions or appointments, both domestic and foreign, titled academic, professional, or institutional appointments whether or not remuneration is received and whether full-time, part-time or voluntary.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer III
Site Reliability Engineer III

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision Plans
401(k) Plan with company contribution
Flexible Work Arrangements
HPDF Building Infrastructure Manager
HPDF Building Infrastructure Manager

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, and Vision
Flexible Spending Accounts
Paid Time-off, vacation, holidays, and
+5
High-Performance Computing (HPC) Systems Engineer
High-Performance Computing (HPC) Systems Engineer

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, Vision
Paid Time Off
401(k) Plan
+2
HVAC Technician II
HVAC Technician II

Jefferson Lab • Newport News (VA)

On-site
USD 58,000 - 84,000
Medical, Dental, and Vision Plans
Flexible Work Arrangements
Paid Time-off and Leave Programs
+3
Scientific Software & DevOps Engineer
Scientific Software & DevOps Engineer

Jefferson Lab • Newport News (VA)

Hybrid
USD 92,000 - 145,000
Flexible Work Arrangements
Tuition Assistance
Paid Time Off
+2
DCS Data Scientist II - Data Steward
DCS Data Scientist II - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, and Vision Plans
Flexible Spending Accounts
Paid Time-off and Leave Programs
+2
DCS Data Scientist III - Data Steward
DCS Data Scientist III - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical plan
Dental plan
Vision plan
+5
DC Power Technician II
DC Power Technician II

Phase2 Technology • Newport News (VA)

On-site
USD 58,000 - 84,000
Medical insurance
Dental insurance
Vision care
+2
Storage Architect
Storage Architect

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision
Flexible Work Arrangements
401(k) Plan
EI&C Technician I
EI&C Technician I

Phase2 Technology • Newport News (VA)

On-site
USD 45,000 - 71,000
Medical, Dental, and Vision Care Plans
Paid Time-off and Leave Programs
401(k) Plan - 9% Lab Contribution; 100