Sr. HPC System Administrator

uchicago

Chicago (IL)

Hybrid

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

HPC workstation access

Job summary

The University of Chicago Research Computing Center is hiring a Senior HPC System Administrator to design automated, scalable infrastructure and manage HPC clusters. You will install, configure, and maintain Linux and Windows servers, monitor systems, and ensure security and backups.

The role includes procurement and management of HPC hardware and software, in collaboration with researchers and vendors. This hybrid position requires three days onsite and offers opportunities to work with

Qualifications

  • Education: degree in a related field; advanced degrees preferred.
  • 5–7 years of work experience in a related HPC/system administration role.
  • Experience with HPC clusters used for scientific research is desirable.

Responsibilities

  • Install, configure, and maintain large clusters/servers and software.
  • Operate day-to-day systems admin tasks, monitoring, and storage performance.
  • Configure scheduling/queuing systems and manage backups.
  • Diagnose problems, coordinate with vendors, and assist users with access.
  • Automate system tasks using scripting and programming.
  • Build and deploy open-source and vendor software; document procedures.

Skills

Linux
HPC
SLURM
Networking
Scripting
Security
Automation
MPI/OpenMP
Infiniband
Performance monitoring

Education

Bachelor's degree
Master's preferred

Tools

Job scheduling tools (SLURM, Moab, TORQUE, PBS)
Distributed file systems
XCAT/ROCKS deployment
Network storage subsystems
Automation tools (Ansible, Puppet)

Job description

Department Provost Research Computing Center

Provost Research Computing Center

About the Department

The University of Chicago Research Computing Center (RCC), a unit in the Office of Research, provides high-end research computing resources to researchers at the University of Chicago. It is dedicated to enabling research by providing access to centrally managed High‑Performance Computing (HPC), storage, and visualization resources. These resources include hardware, software, high-level scientific and technical user support, and the education and training required to help researchers make full use of modern HPC technology and local and national supercomputing resources. The Office of Research oversees the conduct of sponsored research, research program development, and contract management functions.

Job Summary

The job uses specialized knowledge and breadth of expertise to design automated, scalable, and rapidly deployable solutions to infrastructure development and server configuration. Leads installation, configuration, and maintenance of operating systems. Uses best practices and systems knowledge to monitor and alert systems, utility software, and firewalls. Guides maintenance for production servers as well as Windows and Linux servers.

The University of Chicago is seeking a highly qualified Senior HPC System Administrator to join the system and operation team that builds and manages RCC HPC systems and facility operations. The individual in this position will be involved in the procurement and management of HPC hardware and software.

This is a hybrid position requiring 3 days onsite.

Responsibilities
  • Installing, configuring, and maintaining large computer clusters/servers and software.
  • Day‑to‑day operations of the systems including systems administration, monitoring and storage performance up to and including network components. Management of the system's network switch, parallel file system and HPC software stack and tools.
  • Configuration of the scheduling and queuing system.
  • Diagnosing and resolving system operational problems quickly and effectively. Coordinating with vendors to resolve hardware and software problems. Assist users with access and other help desk ticket requests or issues.
  • Use scripting/programming skills to enable system‑level automation, problem detection, security maintenance and patch management.
  • Building and deploying open‑source software and software from vendors/partners.
  • Providing reliable and efficient backups/restores for all managed systems.
  • Documenting system administration procedures for routine and complex tasks.
  • Maintaining and monitoring the security of the HPC systems and servers.
  • Plans and installs necessary patches and upgrades for servers and their associated storage, network, communications, and peripheral sub‑systems. Installs and maintains an appropriate level of intrusion detection, monitoring, and auditing software as required.
  • Tracks compliance and maintains documentation for hardware, software, and service inventories for management reports.
  • Performs other related work as needed.
Minimum Qualifications
  • Education: Minimum requirements include a college or university degree in related field.
  • Work Experience: Minimum requirements include knowledge and skills developed through 5-7 years of work experience in a related job discipline.
  • Certifications: ---
Preferred Qualifications
  • Education: Master's degree in Computer Science or closely related field.
  • Experience: Full time Linux system administration experience in a large distributed computing environment.
  • Previous experience in providing support for Linux HPC cluster used for scientific research.
Technical Skills or Knowledge
  • Experience with installing, configuring, and maintaining job management tools (such as SLURM, Moab, TORQUE, PBS, etc.).
  • Experience configuring, installing and troubleshooting MPI and OpenMP.
  • Experience with operating system deployment tools (e.g. XCAT, ROCKS).
  • Experience configuring, administering and supporting network storage subsystems (e.g. IBM, NetAppl DataDirect Network, LSI, etc.).
  • Hands‑on experience of at least one distributed file system (Spectrum Scale‑GPFS, Lustre, BeeGFS, Gluster, IMRIX, PVFS, etc.).
  • Direct experience working with Infiniband (must at least be able to demonstrate a working knowledge of Infiniband concepts, OFED layers, sub‑net managers).
  • Experience configuring, installing, tuning and maintaining scientific application software on large‑scale systems.
  • Experience supporting HPC compilers and libraries.
  • Experience with systems automation tools such as Ansible or Puppet.
  • Experience configuring, installing, maintaining and/or using performance monitoring and optimization tools.
Preferred Competencies
  • Ability to work well with faculty and researchers.
  • Ability to identify and gain expertise in appropriate new technologies and/or software tools.
  • Ability to function as part of an interactive team while demonstrating self‑initiative to achieve project's goals and Research Com
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. HPC System Administrator
Sr. HPC System Administrator

The University Of Chicago • Austin (TX)

Hybrid
USD 100,000 - 125,000
Senior HPC Systems Engineer — Clusters & Automation
Senior HPC Systems Engineer — Clusters & Automation

The University Of Chicago • Austin (TX)

Hybrid
USD 100,000 - 125,000
HPC Data Center Linux IT Specialist
HPC Data Center Linux IT Specialist

CGG Services (U.S.) Inc. • Houston (TX)

On-site
USD 95,000 - 140,000
Senior HPC Systems Administrator - Hybrid
Senior HPC Systems Administrator - Hybrid

uchicago • Chicago (IL)

Hybrid
USD 120,000 - 180,000
HPC workstation access
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000
HPC Systems Engineer
HPC Systems Engineer

Radix Trading Experienced Job Board • New York (NY), Chicago (IL)

On-site
USD 120,000 - 150,000
Systems Administrator III - HPC & Networking
Systems Administrator III - HPC & Networking

University of Chicago • Chicago (IL)

On-site
USD 100,000 - 110,000
Senior HPC Systems Administrator
Senior HPC Systems Administrator

Omega Enterprise Solutions, LLC • Corridor North (MD)

On-site
USD 90,000 - 120,000
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

The Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000
Support Technician, HPC
Support Technician, HPC

Columbia University Irving Medical Center • New York (NY)

On-site
USD 90,000 - 100,000