Research Computing Engineer (AI Infrastructure and HPC)

Penn State University

United States

On-site

USD 90,000 - 130,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Penn State University seeks a Research Computing Engineer to join our ICDS technical team. You will design, operate, automate, and optimize GPU/AI and HPC infrastructure used by researchers across the university, including storage, networking, and workflow systems.

The role develops scalable solutions for machine learning, data-intensive research, and computational science, collaborating with researchers and ICDS staff to meet workload requirements and ensure secure, compliant operations.

Qualifications

  • Proficiency using Linux and command-line tools; able to edit files and configure systems.
  • Strong scripting in Bash and Python for automation.
  • Excellent problem-solving and debugging of complex systems.
  • Ability to collaborate effectively within a technical team.
  • Experience using AI tools to improve workflows in programming, debugging or prototyping.
  • Clear written and verbal communication skills.

Responsibilities

  • Collaborate with teammates, users, and vendor support to diagnose issues and implement solutions across compute, storage, networking, and software environments.
  • Monitor, maintain, automate, and improve AI and HPC systems and supporting infrastructure.
  • Design, deploy, operate, troubleshoot, and optimize systems using DevOps and infrastructure-as-code practices.
  • Support GPU-accelerated computing environments for AI, machine learning, and scientific workloads.
  • Partner with researchers to understand workload requirements and develop practical engineering solutions.
  • Support security, logging, documentation, and compliance processes for the systems operated.
  • Contribute to planning, requirements gathering, process improvement, and operational readiness for new services.
  • Provide timely updates to system documentation and respond to user questions with guidance.
  • Evaluate and improve tools, platforms, and workflows for AI model development and data movement at scale.

Skills

Linux
Bash
Python
Problem solving
Teamwork
AI tools experience
Communication

Tools

Slurm
PBS
HTCondor
LSF
Kubernetes
Docker
OpenStack
MAAS
XCAT

Job description

POSITION SPECIFICS

The Institute for Computational and Data Sciences (ICDS) at Penn State seeks a Resea rch Computing Engineer to join our technical team. This role supports Penn State's resea rch mission by designing, operating , automating, and optimizing the GPU and AI and computing infrastructure used by resea rchers across the university , along with the high-performance computing systems that support it.

This position will be filled at the Research Computing Systems Engineer - Intermediate Professional level.

Candidates must be U.S. citizens due to specific access requirements associated with this position.

This position is ideal for an engineer who enjoys building reliable, scalable systems for machine learning, data-intensive research, and advanced computing workloads. The successful candidate will work across GPU systems, HPC platforms, storage, networking, automation, and user-facing research workflows to enable cutting-edge research in AI, simulation, and computational science.

Work Arrangement : This is a full-time position, which will report to the Research HPC Manager and requires on-site work at University Park and is not supportive of remote work.

Responsibilities
  • Collaborate with teammates, users, and vendor support to diagnose issues and implement solutions across compute , storage, networking, and software environments.
  • Monitor, maintain , automate, and improve AI and HPC systems and supporting infrastructure.
  • Design, deploy, operate , troubleshoot, and optimize systems using DevOps and infrastructure-as-code practices.
  • Support GPU-accelerated computing environments for AI, machine learning, and scientific workloads.
  • Partner with researchers and ICDS staff to understand workload requirements and develop practical engineering solutions for system configuration, performance, and research workflows.
  • Support security, logging, documentation, and compliance process for the systems we operate , including environments subject to federal research security.
  • Contribute to planning, requirements gathering, process improvement, and operational readiness for new services and infrastructure.
  • Provide timely updates to system documentation and respond to user questions with clear, actionable guidance.
  • Evaluate and improve tools, platforms, and workflows that support AI model development, training, inference, and data movement at scale.
Required qualifications and skills include the following
  • Ability to work effectively in a Linux environment, including command-line tools, file editing, POSIX permissions, and system configuration.
  • Strong scripting ability in Bash and Python.
  • Strong problem-solving skills and the ability to debug complex systems.
  • Ability to work collaboratively as part of a technical team.
  • Experience using AI tools or AI agents to improve programming, debugging, development, or prototyping workflows.
  • Clear written and verbal communication skills.
Preferred qualifications

Experience with one or more of the following is helpful but not required.

AI and HPC Workloads
  • Administration of multi-GPU nodes at scale , including driver and firmware lifecycle management, NVLink / NVSwitch topology validation, GPU health monitoring and tuning for multi-GPU or multi-node GPU workloads.
  • Experience supporting AI/ML infrastructure, including environments used for model training, inference, experiment workflows, and large-scale data processing.
  • Experience with job schedulers such as Slurm , PBS, HTCondor , or LSF.
  • Software development experience and familiarity with HPC programming environments such as C/C++, Fortran, CUDA, MPI, or OpenMP.
Infrastructure and Automation
  • DevOps experience, including Git-based workflows, CI/CD, automation, and collaborative development practices.
  • Experience with system deployment tools such as xCAT , Warewulf , OpenCHAMI , OpenStack/Bifrost, or MAAS.
  • Experience administering or supporting Kubernetes.
  • Experience with virtualization and containerization technologies such as VMware, Docker, Apptainer , or Podman .
Networking and Storage
  • Networking experience, including EV
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Computing Engineer (AI Infrastructure and HPC)
Research Computing Engineer (AI Infrastructure and HPC)

Pennsylvania State University • State College

On-site
USD 91,000 - 137,000
Medical coverage
Dental coverage
Vision coverage
+3
Research Computing Engineer (AI Infrastructure and HPC)
Research Computing Engineer (AI Infrastructure and HPC)

The Pennsylvania State University • State College

On-site
USD 91,000 - 137,000
75% tuition discount
Competitive benefits package
AI HPC Systems Engineer - Research Computing
AI HPC Systems Engineer - Research Computing

Penn State University • United States

On-site
USD 90,000 - 130,000
AI & HPC Systems Engineer for Research Computing
AI & HPC Systems Engineer for Research Computing

The Pennsylvania State University • State College

On-site
USD 91,000 - 137,000
75% tuition discount
Competitive benefits package
GPU-Accelerated AI & HPC Systems Engineer
GPU-Accelerated AI & HPC Systems Engineer

Penn State University • University Park (TX)

On-site
USD 81,000 - 122,000
Medical, dental, and vision coverage
Retirement plans
75% tuition discount
+1
Research Computing Engineer (AI Infrastructure and HPC)
Research Computing Engineer (AI Infrastructure and HPC)

Penn State University • University Park (TX)

On-site
USD 81,000 - 122,000
Medical, dental, and vision coverage
Retirement plans
75% tuition discount
+1
On-Site AI & HPC Systems Engineer (GPU Infra)
On-Site AI & HPC Systems Engineer (GPU Infra)

Pennsylvania State University • State College

On-site
USD 91,000 - 137,000
Medical coverage
Dental coverage
Vision coverage
+3
Artificial Intelligence/Machine Learning Data Science Engineer
Artificial Intelligence/Machine Learning Data Science Engineer

Penn State University • United States

On-site
USD 120,000 - 160,000
Artificial Intelligence/Machine Learning Data Science Engineer
Artificial Intelligence/Machine Learning Data Science Engineer

Pennsylvania State University • State College

On-site
USD 81,000 - 122,000
Research Computing Software Engineer - HPC, AI & Cloud
Research Computing Software Engineer - HPC, AI & Cloud

The Applied Research Laboratory at Penn State University • University Park (TX)

On-site
USD 92,000 - 175,000
75% tuition discount
Comprehensive benefits