High-Performance Computing (HPC) Systems Engineer

Phase2 Technology

Newport News (VA)

On-site

USD 92,000 - 145,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision
Paid Time Off
401(k) Plan
Flexible Work Arrangements
Tuition Assistance

Job summary

Jefferson Lab is seeking an HPC Systems Engineer to architect and maintain petabyte-scale computing infrastructure. You will blend Linux system administration with modern automation to support state-of-the-art research clusters.

The role emphasizes high availability, low-latency networking, and containerized services using Kubernetes and Docker/Podman.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering or related field.
  • 3+ years of Linux system administration in enterprise environments.
  • Experience with HPC infrastructure and container orchestration.

Responsibilities

  • Architect and operate petabyte-scale storage and compute infrastructure.
  • Maintain high availability and optimize performance across clusters.
  • Develop automation and IaC for deployment and on-call support.

Skills

Kubernetes
Docker/Podman
Linux admin
Ansible/Puppet/Foreman
High Performance Computing (HPC)

Education

Bachelor's in CS/CE or related
Master's in CS/Data Science (preferred)

Tools

Lustre
Ceph/CephFS
NFS
Slurm

Job description

At Jefferson Lab,you'llchampioncutting-edgescience and operational excellence while shaping the future of discovery. Join us and make your mark - where excellence meets purpose, andgreat mindstrulymatter.

The good-faith pay range for this role is $91,800 - $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What your job will be like:

We are seeking a High-Performance Computing (HPC) Systems Engineer to architect, deploy, and maintain the large-scale physical hardware, distributed filesystems, and low-latency networking infrastructure powering our scientific computing ecosystem. This role focuses on the bare-metal and system-level foundations of a petabyte-scale environment, ensuring high availability, peak storage performance, and reliable data movement for experimental nuclear physics workloads. The ideal candidate will blend deep Linux systems administration expertise with modern infrastructure-as-code automation to support state-of-the-art research computing clusters.

In this job you will:
  • Collaboratively design and implement software infrastructure supporting of High Performance and High Throughput computing using best-in-class containerization and orchestration tools.
  • Deploy, maintain, and operate physical infrastructure for scientific computing services.
  • Create highly available services through consideration of the entire stack from hardware and networking though the user application.
  • Engage with users to meet service requirements for performance and availability.
  • Develop software to address gaps in existing tools.
Additional Responsibilities
  • Audit and recommend architectural changes to improve performance, security, and availability.
  • Contribute to upstream development efforts of community tools.
  • Work with students and interns on research projects.
Experience
  • Required: 3 or more years Experience performing enterprise linux System Administration tasks including installation, configuration, and support of COTS/GOTS/FOSS software, file, network, and large-scale storage systems.
  • Required: Experience operating in a production environment with high availability requirements.
  • Required: Experience with automating test procedures
  • Preferred: Experience with NP/HEP HPC infrastructure like slurm, Rucio, Globus, XrootD
Education
  • Required: Bachelor's Degree in Computer Science, Computer Engineering, or related degree with significant Computer Science coursework
  • Preferred: Master's Degree Computer Science, Data Science, or related discipline
Experience and Education Exchange

Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.

Knowledge, Skills, and Abilities
  • Ability to architect, provision, tune, and maintain petabyte-scale, high-performance parallel and distributed storage systems, including Lustre, Ceph, CephFS, and NFS.
  • Expert Kubernetes, Docker/Podman, and containerization administration.
  • Expert Knowledge of Linux system administration including installation and configuration management (ie. Ansible/Puppet/Foreman)
  • Ability to develop, debug, and test applications based upon design and performance requirements.
  • Ability to communicate clearly in writing (e.g., email, presentations, drawings) and explain work to their supervisor and others in the group.
  • Ability to work effectively with peers and participate in participate in troubleshooting, including off-hours during outages and as part of an occasional on-call rotation.
About Jefferson Lab

Join a community with a common purpose of solving the most challenging scientific and engineering problems of our time. The Jefferson Lab campusis located insoutheasternVirginiaamidst a vibrant and growing technology community.

A career at Jefferson Lab is more than a job. You will be part of "big science" and work alongside top scientists and engineers from around the world unlocking the secrets of our visible universe. Managed by SURATech, LLC, Thomas Jefferson National Accelerator Facility is entering an exciting period of mission growth and is seeking new team members ready to apply their skills and passion to have an impact. You could call it work, or you could call it a mission. We call it a challenge. We do things that will change the world.

Total Rewards at Jefferson Lab
  • * Medical, Dental, and Vision Care Plans * Flexible Spending Accounts
  • * Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and sick leave)
  • * 401(k) Plan - 9% Lab Contribution; 100% vested * Flexible Work Arrangements
  • (Remote & Alternate Work Schedules available)
  • * Tuition Assistance, Training and Professional Development Programs
  • * Live near the waterways of the Chesapeake Bay region with access to nearby beaches,
  • mountains, and all major metropolitan centers on the East Coast

SURATech, LLC manages and operates the Thomas Jefferson National Accelerator Facility (Jefferson Lab). SURATech is an Equal Opportunity Employer.

SURATech is committed to providing reasonable accommodation for people with disabilities (unless doing so will result in an undue hardship). If you need a reasonable accommodation for any part of the employment process, please send an e‑mail to recruiting@jlab.org or contact Human Resources by calling (757) 269-7100 and selecting option 1 between 8 am - 5 pm EST to provide the nature of your request.

Employment with SURATech is conditional upon DOE approval if at any time during your employment you are participating in a Foreign Government Talent Recruitment Program or Affiliated activity. Generally, such programs/activities include any foreign-state-sponsored attempt to acquire U.S.-funded scientific research through programs run or funded by the government that target scientists, engineers, students, academics, researchers, and entrepreneurs of all nationalities working or educated in the United States. This includes positions or appointments, both domestic and foreign, titled academic, professional, or institutional appointments whether or not remuneration is received and whether full-time, part-time or voluntary.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Scientific Software & DevOps Engineer
Scientific Software & DevOps Engineer

Jefferson Lab • Newport News (VA)

Hybrid
USD 92,000 - 145,000
Flexible Work Arrangements
Tuition Assistance
Paid Time Off
+2
Storage Architect
Storage Architect

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision
Flexible Work Arrangements
401(k) Plan
Site Reliability Engineer III
Site Reliability Engineer III

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical, Dental, Vision Plans
401(k) Plan with company contribution
Flexible Work Arrangements
Site Reliability Engineer II
Site Reliability Engineer II

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, Vision plans
Flexible work arrangements
Paid time off
DCS Data Scientist II - Data Steward
DCS Data Scientist II - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 92,000 - 145,000
Medical, Dental, and Vision Plans
Flexible Spending Accounts
Paid Time-off and Leave Programs
+2
DCS Data Scientist III - Data Steward
DCS Data Scientist III - Data Steward

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 187,000
Medical plan
Dental plan
Vision plan
+5
HPDF Building Infrastructure Manager
HPDF Building Infrastructure Manager

Phase2 Technology • Newport News (VA)

On-site
USD 118,000 - 170,000
Medical, Dental, and Vision
Flexible Spending Accounts
Paid Time-off, vacation, holidays, and
+5
HVAC Technician II
HVAC Technician II

Jefferson Lab • Newport News (VA)

On-site
USD 58,000 - 84,000
Medical, Dental, and Vision Plans
Flexible Work Arrangements
Paid Time-off and Leave Programs
+3
EI&C Technician I
EI&C Technician I

Phase2 Technology • Newport News (VA)

On-site
USD 45,000 - 71,000
Medical, Dental, and Vision Care Plans
Paid Time-off and Leave Programs
401(k) Plan - 9% Lab Contribution; 100
Hall A Mechanical Technician II
Hall A Mechanical Technician II

Jefferson Lab • Newport News (VA)

On-site
USD 58,000 - 84,000
Medical, Dental, and Vision Care Plans
401(k) Plan – 9% Lab Contribution
Flexible Work Arrangements