Senior HPC Architect

Oak Ridge National Laboratory

Oak Ridge (TN)

On-site

USD 150,000 - 210,000

Full time

33 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical plan
Retirement plan
Relocation assistance
Wellness programs
On-site amenities

Job summary

Oak Ridge National Laboratory is seeking a Senior HPC Architect to lead the architecture, design, and evolution of high-performance computing platforms supporting open research and classified computing missions.

You will drive end-to-end technical direction, develop reference architectures, guide system design and technology selection, and ensure operational excellence through automation and observability. Strong leadership and collaboration with cybersecurity are essential.

Qualifications

  • BS degree in Computer Science, Engineering, or related field.
  • 8+ years of experience in HPC engineering with system architecture, cluster operations, parallel computing, and performance optimization.
  • Experience in high-security or regulated environments.
  • Strong HPC cluster management and scheduling experience.
  • Experience with HPC performance monitoring/benchmarking tools (Grafana, Nagios, Ganglia).
  • Ability to write clear technical docs and communicate with engineers and non-engineers.

Responsibilities

  • Lead the design and deployment of HPC systems meeting performance, reliability, and security requirements for research and/or classified environments.
  • Define and maintain reference architectures for compute, network, storage, and platform services.
  • Produce and maintain architecture diagrams, configuration standards, and operational procedures.
  • Guide system and accelerator architecture decisions with GPU/accelerator impact on datacenter and AI/HPC workloads.

Skills

HPC architecture
Cluster operations
Performance optimization
Linux at scale
Security/regulatory compliance
Documentation & communication
Mentorship

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

Ansible
Puppet
Chef
Kickstart
Satellite
Docker
Kubernetes
Kubeflow
Grafana
Nagios
Ganglia

Job description

Overview

As a Senior HPC Architect at Oak Ridge National Laboratory (ORNL), you will lead the architecture, design, and evolution of high-performance computing platforms supporting both open research and classified computing missions. This role is responsible for translating mission and scientific requirements into scalable, reliable, and secure HPC architectures - spanning compute (CPU/GPU), high-speed interconnects, storage and parallel file systems, cluster management, and user-facing platform services.

In practice, the Senior HPC Architect drives end-to-end technical direction for HPC environments: developing reference architectures, guiding system design decisions and technology selection, defining performance and capacity models, and ensuring operational excellence through standardization, automation, and observability. The position partners closely with cybersecurity and compliance stakeholders to deliver secure-by-design infrastructure aligned with regulatory requirements and authorization processes (e.g., ATO, NIST controls, STIG implementation).

This role also provides technical leadership across teams—mentoring engineers, leading architecture reviews, and coordinating delivery across infrastructure, security, and scientific programs. Success requires deep expertise in modern HPC system architecture, GPU-centric platforms, cluster scheduling, Linux at scale, and rigorous performance engineering, with the ability to operate effectively in regulated environments.

Major Duties/Responsibilities
  • Lead the design and deployment of HPC systems to meet performance, reliability, and security requirements for research and/or classified computing environments.
  • Define and maintain reference architectures for compute, network, storage, and platform services, including lifecycle planning and roadmap development.
  • Produce and maintain technical documentation: architecture diagrams, configuration standards, operational procedures, and engineering decision records.
  • Guide system and accelerator architecture decisions with a strong grasp of how GPU/accelerator architecture impacts datacenter and AI/HPC workloads (including LLM-adjacent workloads where applicable).
Automation, Infrastructure Platforms & Enablement
  • Identify automation targets and lead adoption of infrastructure-as-code and configuration management (e.g., Ansible, Puppet, Chef, Kickstart, Satellite).
  • Build standardized, repeatable deployment workflows; contribute to internal platform tools and codebases where appropriate (e.g., Python-based automation).
  • Where applicable, architect and support container and platform capabilities (e.g., Docker/Kubernetes) to enable reproducible scientific workflows and service deployment.
Leadership & Collaboration
  • Lead HPC-related projects from planning through implementation and steady-state operations; manage technical risks, dependencies, and milestones.
  • Partner with scientists, researchers, and mission stakeholders to translate workflow requirements into platform capabilities.
  • Mentor junior engineers; create knowledge-sharing practices, documentation, and engineering standards.
Basic Qualifications
  • BS degree in Computer Science, Engineering, or a related field and a minimum of 8+ years of relevant experience (or equivalent combination of education and experience).
  • 8+ years of experience in HPC engineering with demonstrated strength in system architecture, cluster operations, parallel computing environments, and performance optimization.
  • Demonstrated experience working in high-security and/or regulated environments
  • Strong experience with HPC cluster management and scheduling
  • Experience with HPC performance monitoring and benchmarking using tools such as Grafana, Nagios, Ganglia (or equivalent).
  • Ability to lead technical initiatives, write clear technical documentation, and communicate effectively with both engineering and non-engineering stakeholders.
Preferred Qualifications
  • Familiarity with parallel file systems / advanced storage
  • Experience with containerization and HPC-adjacent platforms (e.g., Docker, Kubernetes, Kubeflow) in a way that complements scheduler-based HPC usage.
  • Experience with virtualization platforms (e.g., VMware) in support of HPC infrastructure services.
  • Strong Infrastructure-as-Code background (e.g., Ansible, plus cloud/IaC such as Terraform/Packer where relevant).
  • Experience supporting scientific software development and deployment and research user workflows.
  • Preferred: experience with geospatial data workflows, including large geospatial/raster/vector datasets, spatial ETL pipelines, and performance considerations for geospatial analytics at scale (e.g., tiling/partitioning strategies, I/O patterns, reproducibility, and access controls for sensitive geospatial data).
  • Strong leadership, mentoring, and cross-team coordination skills; ability to manage multiple priorities in fast-paced, high-consequence environments.
Special Requirements
  • Q clearance with SCI : This position requires the ability to obtain and maintain a Secret Compartmented Information (SCI) clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse (WSAP) testing designated position. WSAP positions require passing a pre-placement drug test and participation in an ongoing random drug testing program. In addition, due the SCI, you may also be subject to random polygraph testing.
About ORNL

As a U.S. Department of Energy (DOE) Office of Science national laboratory, ORNL has an impressive 80-year legacy of addressing the nation’s most pressing challenges. Our team is made up of over 7,000 dedicated and innovative individuals! Our goal is to create an environment where a variety of perspectives and backgrounds are valued, ensuring ORNL is known as a top choice for employment. These principles are essential for supporting our broader mission to drive scientific breakthroughs and translate them into solutions for energy, environmental, and security challenges facing the nation.

ORNL offers competitive pay and benefits programs to attract and retain individuals who demonstrate exceptional work behaviors. The laboratory provides a range of employee benefits, including medical and retirement plans and flexible work hours, to support the well-being of you and your family. Employee amenities such as on-site fitness, banking, and cafeteria facilities are also available for added convenience.

Other benefits include the following: Prescription Drug Plan, Dental Plan, Vision Plan, 401(k) Retirement Plan, Contributory Pension Plan, Life Insurance, Disability Benefits, Generous Vacation and Holidays, Parental Leave, Legal Insurance with Identity Theft Protection, Employee Assistance Plan, Flexible Spending Accounts, Health Savings Accounts, Wellness Programs, Educational Assistance, Relocation Assistance, and Employee Discounts.

ORNL is an equal opportunity employer. All qualified applicants, including individuals with disabilities and protected veterans, are encouraged to apply. UT‑Battelle is an E-Verify employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Software Engineer
HPC Software Engineer

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 90,000 - 120,000
Medical and retirement plans
Flexible work hours
On-site fitness and amenities
+1
Classified HPC Systems Engineer (Hybrid Eligible)
Classified HPC Systems Engineer (Hybrid Eligible)

UT-Battelle • Oak Ridge (TN)

Hybrid
USD 120,000 - 180,000
Hybrid/onsite working arrangement
Relocation assistance
Hybrid work environment
Senior High Performance Computing Engineer - Classified Environment
Senior High Performance Computing Engineer - Classified Environment

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 140,000 - 190,000
Medical benefits
Flexible working hours
On-site amenities
Senior Research Scientist Advanced Computing Systems
Senior Research Scientist Advanced Computing Systems

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 170,000 - 210,000
On-site fitness
Cafeteria facilities
Relocation Assistance
Classified HPC Systems Engineer (Hybrid Eligible)
Classified HPC Systems Engineer (Hybrid Eligible)

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 110,000 - 150,000
401(k) Retirement Plan
Medical Plan
Relocation Assistance
+2
HPC Linux Systems Engineer
HPC Linux Systems Engineer

UT-Battelle • Oak Ridge (TN)

On-site
USD 120,000 - 180,000
Senior HPC Engineer - Classified Environment
Senior HPC Engineer - Classified Environment

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 140,000 - 190,000
Relocation Assistance
Wellness Programs
Education Assistance
Senior HPC Linux Systems Engineer, Classified Environment
Senior HPC Linux Systems Engineer, Classified Environment

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 100,000 - 130,000
Access to professional development
Mission-driven environment
High quality of life in East Tennessee
Group Leader, IC Classified Compute (IC3)
Group Leader, IC Classified Compute (IC3)

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 140,000 - 190,000
Senior HPC Engineer, Classified Computing
Senior HPC Engineer, Classified Computing

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 90,000 - 120,000
Medical, dental, vision insurance
401(k) and pension plan
Generous vacation and holidays