HPC DevOps Engineer

University of North Carolina at Chapel Hill

Chapel Hill (NC)

Hybrid

USD 105,000 - 110,371

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Retail discounts
Savings on child care
Campus events discounts
Comprehensive benefits package
Paid leave
Health and retirement plans

Job summary

The University of North Carolina at Chapel Hill seeks an HPC DevOps Engineer to design, deploy, operate and maintain HPC/HTC services in Research Computing. The role emphasizes Linux systems, automation, storage, networking and observability in a high-performance academic research environment.

The successful candidate will lead complex projects, engage with faculty and researchers, and contribute to a world-class computing infrastructure, with hybrid work possibilities and a strong focus on

Qualifications

  • Master’s degree with 1–2 years’ experience or Bachelor’s with 2–4 years’ experience.
  • Proficiency with Linux operating systems in HPC environments.
  • Experience with automation/configuration-management tools (Cobbler, SALT, Ansible).
  • Familiarity with enterprise storage, monitoring, and logging.
  • Strong communication and ability to lead complex projects.
  • Ability to lift up to 50 pounds and work in data-center spaces.

Responsibilities

  • Design, deploy, operate and maintain HPC/HTC services.
  • Lead architecture and automation for observability platforms.
  • Coordinate with vendors and internal stakeholders.
  • Ensure security and data-handling compliance.

Skills

Linux administration
Automation tooling
Monitoring & logging
RHEL repo management
Hardware maintenance
Communication
Team collaboration

Education

Master’s degree
Bachelor’s degree

Tools

Cobbler
SALT
Ansible

Job description

Department: Research Computing-615000

Career Area : Information Technology

Posting Open Date: 07/23/2026

Application Deadline: 08/27/2026

Open Until Filled: No

Position Type: Permanent Staff (EHRA NF)

Working Title: HPC DevOps Engineer

Appointment Type: EHRA Non-Faculty

Position Number: 20025114

Vacancy ID: NF0009910

Full Time/Part Time: Full-Time Permanent

FTE: 1

Hours per week: 40

Position Location: North Carolina, US

Hiring Range: $105,000 - $110,371

Proposed Start Date: 08/31/2026

Be a Tar Heel!: A global higher education leader in innovative teaching, research and public service, the University of North Carolina at Chapel Hill consistently ranks as one of the nation’s top public universities. Known for its beautiful campus, world-class medical care, commitment to the arts and top athletic programs, Carolina is an ideal place to teach, work and learn.One of the best college towns and best places to live in the United States, Chapel Hill has diverse social, cultural, recreation and professional opportunities that span the campus and community.University employees can choose from a wide range of professional training opportunities for career growth, skill development and lifelong learning and enjoy exclusive perks for numerous retail, restaurant and performing arts discounts, savings on local child care centers and special rates on select campus events. UNC-Chapel Hill offers full-time employees a comprehensive benefits package, paid leave, and a variety of health, life and retirement plans and additional programs that support a healthy work/life balance.

Primary Purpose of Organizational Unit: ITS Research Computing (RC) aims to provide a world-class computing infrastructure as well as other technology tools and capabilities to support the research needs of University faculty and staff. Its goal is to provide a state-of-the-art environment to support the highest level of multidisciplinary research and help UNC-Chapel Hill become the premier research university in the United States.

Position Summary: This position may be eligible for a hybrid work arrangement that may include a partially remote work location, consistent with System Office policy. UNC Chapel Hill employees are generally required to reside within a reasonable commuting distance of their assigned duty station.

The HPC DevOps Engineer will have a broad role within Research Computing at UNC Chapel Hill, designing, deploying, operating and maintaining services and solutions in support of leading -edge academic research, primarily related to High Performance Computing (HPC) as well as High Throughput Computing (HTC).

Responsibilities of the HPC DevOps Engineer include leading and contributing to a variety of areas within Research Computing related to the design, build, and operation of HTC/HPC services along with ancillary systems including high-speed storage and networking, containers, databases, monitoring, orchestration services, etc.

The ideal candidate for this position should possess a comprehensive understanding of Linux systems administration in HPC environments, broad technical capabilities, and enjoy applying these skills in academic research to further the mission of the University.

The successful candidate should possess excellent leadership and communication skills to effectively lead complex projects. They should have the ability to engage directly with faculty and researchers, as well as leverage research computing communities of practice beyond Carolina.

ITS Research Computing aims to provide world-class computing and data infrastructure as well as other services and capabilities in support of research needs for faculty, staff, students, and collaborators. Our goal is to provide state-of-the-art environments and services supporting the highest level of multidisciplinary research.

Minimum Education and Experience Requirements: Master’s and 1-2 years’ experience; or Bachelors and 2-4 years’ experience; or will accept a combination of related education and experience in substitution.

Required Qualifications, Competencies, and Experience:

  • Proficiency with Linux operating systems, including installation, configuration, troubleshooting, and lifecycle management in an enterprise or research-computing context. Experience with automation and configuration-management tools such as Cobbler, SALT, or comparable platforms used to deploy and manage systems at scale. Familiarity with enterprise storage platforms, including routine configuration tasks, capacity monitoring, health assessment, and basic hardware maintenance such as drive replacement. Ability to interpret and act on monitoring, logging, and alerting data to maintain operational continuity across diverse service lines. Understanding of RHEL repo management and practices for maintaining consistent, secure OS-layer operations.
  • Ability to lift and maneuver equipment up to 50 pounds with or without reasonable accommodations, work in hot/cold aisle conditions, and perform tasks in confined rack environments. Capacity to stand, bend, and work in data-center spaces for extended periods while performing installation, cabling, and hardware maintenance tasks.
  • Strong diagnostic and problem-solving abilities, including the capacity to assess hardware and system issues under time constraints. Demonstrated adherence to change-management and operational best practices in complex technical environments. Ability to coordinate effectively with vendors, service providers, and internal stakeholders to support hardware and storage operations.
  • Clear written and verbal communication skills, including the ability to convey technical information to both technical and non-technical audiences. Ability to work collaboratively within a team where some responsibilities are shared and others are independently owned.
  • Understanding of operational practices required to support environments involving controlled information and regulated data, including attention to security, access control, and data-handling expectations.
  • The position requires expertise with Linux operating systems and the ability to engineer, maintain, and troubleshoot large-scale OS deployments for HPC and research-computing environments. The role demands sustained focus, careful change-management discipline, and the ability to diagnose complex system-level issues in a fast-moving research environment.
  • Duties include contributing to the architecture and automation that underpin observability platforms, ensuring reliable telemetry collection and correlation across diverse service lines, and interpreting operational signals to maintain service health, security posture, and compliance expectations. The role requires disciplined incident response, careful analysis of system behavior, and the ability to translate monitoring insights into stable, well-governed operations.
  • The role requires strong organizational awareness, clear communication with service owners and vendors, and steady operational judgment to maintain reliable, well-governed storage services.

Preferred Qualifications, Competencies, and Experience:

Experience managing InfiniBand HPC clusters. Experience managing multi petabyte enterprise storage platforms, including vendor specific administration tools, performance tuning concepts, and lifecycle planning. Hands on experience with data center hardware ecosystems such as multi node server deployments, high density rack environments, and hardware lifecycle management. Familiarity with infrastructure automation, including developing or extending provisioning and configuration management code in tools such as Cobbler, SALT, Ansible, or similar platforms. Experience maintaining RHEL based environments at scale, including custom repository management, patch orchestration, and compliance aligned OS baselines. Background working with observability platforms (e.g., Prometheus, Grafana, ELK/Opensearch, Splunk, Alertmanager)and contributing to monitoring or logging architecture.

Special Physical/Mental Requirements: The position requires the ability to lift and maneuver hardware components up to 50 pounds with or without reasonable accommodations, work in confined rack environments, and remain on one’s feet for extended periods. It also requires sustained concentration, situational awareness, and the ability to perform precise technical tasks under time pressure in an active enterprise data center environment.

Campus Security Authority Responsibilities: Not Applicable.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC DevOps Engineer
HPC DevOps Engineer

University of North Carolina - Chapel Hill • Elkins Hills (NC)

Hybrid
USD 105,000 - 111,000
Hybrid HPC DevOps Engineer for Research Computing
Hybrid HPC DevOps Engineer for Research Computing

University of North Carolina at Chapel Hill • Chapel Hill (NC)

Hybrid
USD 105,000 - 110,371
Retail discounts
Savings on child care
Campus events discounts
+3
Hybrid HPC DevOps Engineer: Scale HPC with Automation
Hybrid HPC DevOps Engineer: Scale HPC with Automation

University of North Carolina - Chapel Hill • Elkins Hills (NC)

Hybrid
USD 105,000 - 111,000
Senior Research Administration Specialist
Senior Research Administration Specialist

University of North Carolina at Chapel Hill • Chapel Hill (NC)

Hybrid
USD 62,000 - 85,000
Sr. HRIM Reporting Analyst
Sr. HRIM Reporting Analyst

Inside Higher Ed • Chapel Hill (NC)

Remote
USD 60,000 - 80,000
Comprehensive benefits package
Professional training opportunities
Discounts on retail and childcare
Technology Support Technician - Journey
Technology Support Technician - Journey

The University of North Carolina at Chapel Hill • Chapel Hill (NC)

On-site
USD 55,000 - 66,000
Cloud Senior Principal Engineer
Cloud Senior Principal Engineer

The University of North Carolina System • Raleigh (NC)

Hybrid
USD 140,000 - 210,000
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

The Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000
HPC Infrastructure Platform Engineer
HPC Infrastructure Platform Engineer

Cadre5 • Knoxville (TN)

Hybrid
USD 120,000 - 190,000
Medical, dental, and vision coverage
401K match
15 days PTO
+1