Research Computing Engineer (AI Infrastructure and HPC)

Pennsylvania State University

State College (Centre County)

On-site

USD 91,000 - 137,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical coverage
Dental coverage
Vision coverage
Retirement plans
Paid time off
Tuition discount

Job summary

Pennsylvania State University’s Institute for Computational and Data Sciences seeks a Research Computing Engineer to design, operate, and optimize GPU, AI, and computing infrastructure used by researchers across the university. This role is on-site at University Park and requires U.S.

citizenship due to access requirements. The role focuses on building scalable systems for machine learning, data-intensive research, and advanced computing workloads, collaborating across GPU, HPC, storage,

Qualifications

  • Bachelor’s degree and 3+ years of relevant experience.
  • U.S. citizenship required due to access controls.
  • Experience Linux admin and scripting (Bash, Python).
  • Experience with AI tools to improve workflows.
  • Ability to work collaboratively in a technical team.

Responsibilities

  • Collaborate across compute, storage, network, and software environments.
  • Monitor, automate, and improve AI and HPC systems.
  • Design, deploy, and optimize systems using IaC and DevOps.
  • Support GPU-accelerated workloads for AI and science.
  • Develop solutions for workload requirements with researchers.
  • Ensure security, logging, documentation, and compliance.
  • Contribute to planning and readiness for new services.
  • Maintain up-to-date system documentation and user guidance.
  • Evaluate tools and workflows for AI model development at scale.

Skills

GPU nodes
Linux
Bash
Python
Team collaboration
AI tools

Education

Bachelor’s Degree

Tools

Kubernetes
Docker
OpenStack
Warewulf
xCAT
MAAS
OpenCHAMI

Job description

Position Specifics

The Institute for Computational and Data Sciences (ICDS) at Penn State seeks a Research Computing Engineer to join its technical team. This role supports Penn State’s research mission by designing, operating, automating, and optimizing the GPU, AI, and computing infrastructure used by researchers across the university, along with the high-performance computing systems that support it. The position will be filled at the Research Computing Systems Engineer - Advanced Professional level. Candidates must be U.S. citizens due to specific access requirements associated with this position.

This position is ideal for an engineer who enjoys building reliable, scalable systems for machine learning, data-intensive research, and advanced computing workloads. The successful candidate will work across GPU systems, HPC platforms, storage, networking, automation, and user-facing research workflows to enable cutting-edge research in AI, simulation, and computational science.

Work Arrangement: This is a full-time position, which will report to the Research HPC Manager and requires on-site work at University Park and is not supportive of remote work.

Responsibilities
  • Collaborate with teammates, users, and vendor support to diagnose issues and implement solutions across compute, storage, networking, and software environments
  • Monitor, maintain, automate, and improve AI and HPC systems and supporting infrastructure
  • Design, deploy, operate, troubleshoot, and optimize systems using DevOps and infrastructure-as-code practices
  • Support GPU-accelerated computing environments for AI, machine learning, and scientific workloads
  • Partner with researchers and ICDS staff to understand workload requirements and develop engineering solutions for system configuration, performance, and research workflows
  • Support security, logging, documentation, and compliance processes for the systems operated, including environments subject to federal research security
  • Contribute to planning, requirements gathering, process improvement, and operational readiness for new services and infrastructure
  • Provide timely updates to system documentation and respond to user questions with clear, actionable guidance
  • Evaluate and improve tools, platforms, and workflows that support AI model development, training, inference, and data movement at scale
Required Qualifications
  • Administration of multi-GPU nodes at scale, including driver and firmware lifecycle management, NVLink/NVSwitch topology validation, and GPU health monitoring and tuning for multi-GPU or multi-node GPU workloads
  • Ability to work effectively in a Linux environment, including command-line tools, file editing, POSIX permissions, and system configuration
  • Strong scripting ability in Bash and Python
  • Strong problem-solving skills and the ability to debug complex systems
  • Ability to work collaboratively as part of a technical team
  • Experience using AI tools or AI agents to improve programming, debugging, development, or prototyping workflows
  • Clear written and verbal communication skills
Preferred Qualifications

Experience with any of the following is helpful but not required:

  • Experience supporting AI/ML infrastructure for model training, inference, experiment workflows, and large-scale data processing
  • Experience with job schedulers such as Slurm, PBS, HTCondor, or LSF
  • Software development experience with HPC programming environments such as C/C++, Fortran, CUDA, MPI, or OpenMP
  • DevOps experience, including Git-based workflows, CI/CD, automation, and collaborative development practices
  • Experience with deployment tools such as xCAT, Warewulf, OpenCHAMI, OpenStack/Bifrost, or MAAS
  • Experience administering or supporting Kubernetes
  • Experience with virtualization/containerization technologies such as VMware, Docker, Apptainer, or Podman
  • Networking experience including EVPN, BGP, and IPv6
  • Experience with high-speed interconnects such as InfiniBand or HPE Slingshot
  • Experience with HPC or distributed storage systems such as GPFS, Lustre, Ceph, or VAST
  • Monitoring and observability experience with tools such as Grafana, Graphite, Prometheus, VictoriaMetrics, or InfluxDB
  • Experience with databases such as MySQL/MariaDB or PostgreSQL
  • Security experience including identity and access management, single sign-on, and LDAP/Active Directory
  • Familiarity with Agile project development
  • Prior experience in academic research computing, research data infrastructure, or large-scale shared computing environments
Minimum Education, Work Experience & Certifications

Bachelor’s Degree and 3+ years of relevant experience; or an equivalent combination of education and experience accepted. No certifications required.

Background Checks/Clearances

Employment with the University will require successful completion of background check(s) in accordance with University policies. Penn State does not sponsor or take over sponsorship of a staff employment Visa; applicants must be authorized to work in the U.S.

Salary & Benefits

The salary range for this position, including all possible grades, is $91,488.00 - $137,280.00.

  • comprehensive medical coverage
  • dental coverage
  • vision coverage
  • robust retirement plans
  • substantial paid time off
  • 75% tuition discount available to employees as well as eligible spouses and children
EEO Is the Law

Penn State is an equal opportunity employer and is committed to providing employment opportunities to all qualified applicants without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Computing Engineer (AI Infrastructure and HPC)
Research Computing Engineer (AI Infrastructure and HPC)

The Pennsylvania State University • State College

On-site
USD 91,000 - 137,000
75% tuition discount
Competitive benefits package
Research Computing Engineer (AI Infrastructure and HPC)
Research Computing Engineer (AI Infrastructure and HPC)

Penn State University • United States

On-site
USD 90,000 - 130,000
Research Computing Engineer (AI Infrastructure and HPC)
Research Computing Engineer (AI Infrastructure and HPC)

Penn State University • University Park (TX)

On-site
USD 81,000 - 122,000
Medical, dental, and vision coverage
Retirement plans
75% tuition discount
+1
Artificial Intelligence/Machine Learning Data Science Engineer
Artificial Intelligence/Machine Learning Data Science Engineer

The Pennsylvania State University • State College

On-site
USD 81,000 - 122,000
75% tuition discount
Excellent retirement plans
Artificial Intelligence/Machine Learning Data Science Engineer
Artificial Intelligence/Machine Learning Data Science Engineer

Pennsylvania State University • State College

On-site
USD 81,000 - 122,000
AI HPC Systems Engineer - Research Computing
AI HPC Systems Engineer - Research Computing

Penn State University • United States

On-site
USD 90,000 - 130,000
AI & HPC Systems Engineer for Research Computing
AI & HPC Systems Engineer for Research Computing

The Pennsylvania State University • State College

On-site
USD 91,000 - 137,000
75% tuition discount
Competitive benefits package
GPU-Accelerated AI & HPC Systems Engineer
GPU-Accelerated AI & HPC Systems Engineer

Penn State University • University Park (TX)

On-site
USD 81,000 - 122,000
Medical, dental, and vision coverage
Retirement plans
75% tuition discount
+1
Artificial Intelligence/Machine Learning Data Science Engineer
Artificial Intelligence/Machine Learning Data Science Engineer

Penn State University • Northern (KY)

Hybrid
USD 81,000 - 122,000
On-Site AI & HPC Systems Engineer (GPU Infra)
On-Site AI & HPC Systems Engineer (GPU Infra)

Pennsylvania State University • State College

On-site
USD 91,000 - 137,000
Medical coverage
Dental coverage
Vision coverage
+3