Sr. HPC Systems Engineer (IT@JH Research Computing) - #Staff

Johns Hopkins University

Baltimore (MD)

On-site

USD 86,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Johns Hopkins University IT@JH Research Computing seeks a Sr. HPC Systems Engineer to design, build, and maintain advanced high‑performance computing environments.

The role focuses on reliable operation, configuration, and optimization of HPC and AI systems, including multi‑node CPU/GPU clusters, InfiniBand networks, and large‑scale storage. You will work with faculty, researchers, and students on ticketed support and project deployments, while implementing secure, reproducible platforms and

Qualifications

  • Bachelor’s degree.
  • Six years of related experience.
  • Additional education may substitute for required experience beyond a high school diploma/graduation equivalent, to the extent permitted by the JHU equivalency formula.

Responsibilities

  • Support and administer production systems used by researchers.
  • Provide technical leadership/project management for system configuration, implementation, management, and user support for both new and existing systems.
  • Research and recommend new functionality for HPC management and administration tools by exploring system-wide impacts.
  • Expertise with architecting, operating, and debugging large scale HPC network and storage infrastructure, including MPI, NCCL, RDMA, Infiniband, and parallel file systems
  • Work with scientific support specialists and assigns tasks and provides oversight to HPC engineering team for researchers using diverse applications
  • Analyze results of server monitoring and implement changes to improve performance, processing, and utilization
  • Propose, maintain, and enforce policies, practices and security procedures
  • Provide break/fix support, setup/installation support, escalation support, and solutions support
  • Collaborate with stakeholders on all aspects of projects
  • Other duties as assigned.

Skills

Six years related experience
HPC systems administration
Linux administration
Automation scripting

Education

Bachelor’s degree

Tools

Slurm
Ansible
Puppet
Salt
GPFS
Lustre
BeeGFS
WekaFS
Infiniband
Docker
Kubernetes

Job description

IT@JH Research Computing is seeking a Sr. HPC Systems Engineer who will design, build, and maintain advanced high-performance computing environments supporting Johns Hopkins University’s research mission. This position focuses on the reliable operation, configuration, and optimization of HPC and AI systems, including multi-node CPU and GPU clusters, high-speed InfiniBand and Ethernet networks, and large-scale parallel and object storage. The engineer implements and automates secure, efficient, and reproducible computing platforms used by faculty, researchers, and students across diverse scientific disciplines. Assignments include both ticket-based support and project-based deployments. The role operates with moderate independence, collaborating closely with the IT Architect, Research Computing, and reporting to the IT Manager for Research Computing to ensure scalable, sustainable, and high-performance systems that enable cutting‑edge scientific discovery.

Specific Duties & Responsibilities
  • Support and administer production systems used by researchers and Research Centers.
  • Provide technical leadership/project management for system configuration, implementation, management, and user support for both new and existing systems.
  • Research and recommend new functionality for HPC management and administration tools by exploring system-wide impacts, working with functional users to define current and future processes.
  • Expertise with architecting, operating, and debugging large scale HPC network and storage infrastructure, including MPI, NCCL, RDMA, Infiniband, and parallel file systems
  • Works with scientific support specialists and assigns tasks and provides oversight as appropriate to HPC engineering team to support scientific researchers who use a broad spectrum of applications from diverse fields.
  • Analyze results of server monitoring and implement changes to improve performance, processing, and utilization.
  • Propose, maintain, and enforce policies, practices and security procedures.
  • Provide break/fix support, setup/installation support, escalation support, and solutions support.
  • Collaborate closely with a variety of stakeholders, both internal and external, on all aspects of projects.
  • Other duties as assigned.

In Addition to the Duties Described Above

  • Deploy, configure, and maintain large-scale Linux-based HPC clusters comprising CPU and GPU nodes, high-speed interconnects, and parallel file systems.
  • Implement and optimize workload schedulers (Slurm) and job submission policies to maximize system throughput and fair-share usage.
  • Administer and monitor distributed storage systems (GPFS, Lustre, WekaFS, Ceph, MinIO) to ensure reliability and performance across multi-petabyte environments.
  • Maintain high-speed fabric and network infrastructure (Infiniband, Ethernet) to support low-latency data transfer and MPI workloads.
  • Support research groups in deploying, testing, and optimizing scientific applications and AI/ML workflows on shared computing resources.
  • Develop and maintain automation and monitoring frameworks for system provisioning, metrics collection, and alerting (Prometheus, Grafana, ELK).
  • Participate in capacity planning, hardware lifecycle management, and evaluation of new technologies in collaboration with architects and management.
  • Ensure security and compliance through configuration hardening, patch management, and integration with campus identity and access control systems.
  • Document system designs, procedures, and troubleshooting guides to support knowledge transfer and team continuity.
  • Contribute to a collaborative engineering culture that emphasizes service quality, innovation, and continuous improvement in research computing operations.
Minimum Qualifications
  • Bachelor’s degree.
  • Six years of related experience.
  • Additional education may substitute for required experience and additional related experience may substitute for required education beyond a high school diploma/graduation equivalent, to the extent permitted by the JHU equivalency formula.
Preferred Qualifications
  • Eight plus years of experience in high-performance computing systems administration or engineering, including experience with cluster management, workload scheduling (e.g., Slurm), and distributed or parallel storage.
  • Deep proficiency in Linux systems administration, configuration management (Ansible, Puppet, or Salt), performance monitoring, and tuning for HPC workloads.
  • Experience with high-speed interconnects (Infiniband, 100/400 Gb Ethernet) and parallel file systems (e.g., GPFS, Lustre, BeeGFS, or WekaFS).
  • Working knowledge of containerization and orchestration (Singularity, Docker, Kubernetes for HPC).
  • Ability to automate deployments and routine operations through scripting (Bash, Python).
  • Familiarity with data-center operations, GPU acceleration, and research software environments (e.g., CUDA, MPI, AI/ML frameworks).
  • Strong analytical and troubleshooting skills, with proven ability to support complex research workloads in multi-user, multi-tenant environments.
  • Experience collaborating with faculty and research groups to translate scientific requirements into practical and performant computing solutions.

Classified Title: Sr. HPC Systems Engineer

Role/Level/Range: ATP/04/PF

Starting Salary Range: $85,500 - $149,800 Annually (Commensurate w/exp.)

Employee group: Full Time

Schedule: Mon-Fri, 8:30am-5pm

FLSA Status: Exempt

Location: Johns Hopkins Bayview

Department name: IT@JH Research Computing

Personnel area: University Administration

Equal Opportunity Employer

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,500 - 149,800
HPC Scientific Software Engineer (IT@JH Research Computing)
HPC Scientific Software Engineer (IT@JH Research Computing)

The Johns Hopkins University • Baltimore (MD)

On-site
USD 80,000 - 120,000
Senior HPC Systems Engineer - AI & Research Clusters
Senior HPC Systems Engineer - AI & Research Clusters

Johns Hopkins University • Baltimore (MD)

On-site
USD 86,000 - 150,000
Lead Systems Engineer (HPC)
Lead Systems Engineer (HPC)

Princeton University • Princeton (NJ)

On-site
USD 135,000 - 150,000
Comprehensive benefits program
Support Technician, HPC
Support Technician, HPC

Columbia University Irving Medical Center • New York (NY)

On-site
USD 90,000 - 100,000
HPC Engineer - Privacy Research Program
HPC Engineer - Privacy Research Program

Tufts University • Boston (MA)

On-site
USD 89,000 - 134,000
Senior HPC & AI Software Engineer
Senior HPC & AI Software Engineer

The Johns Hopkins University • Baltimore (MD)

On-site
USD 80,000 - 120,000
Lead HPC Systems Engineer for AI and Research Clusters
Lead HPC Systems Engineer for AI and Research Clusters

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,500 - 149,800
HPC Engineer - Privacy Research Program
HPC Engineer - Privacy Research Program

Tufts University • Massachusetts

On-site
USD 89,000 - 134,000
Sr. Systems Administrator - KSAS (Krieger Information Technology)
Sr. Systems Administrator - KSAS (Krieger Information Technology)

Johns Hopkins University • Baltimore (MD)

On-site
USD 64,000 - 113,000