Lead HPC and Systems Engineer

EPAM Systems

United States

On-site

USD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EPAM Systems Inc. seeks a Lead HPC and Systems Engineer to drive modernization of enterprise infrastructure. You will enable Windows Workspaces and Linux HPC, supporting simulation workloads with Ansys, Abaqus, MATLAB, and more.

AWS PCS, PBS, and front-end portals will be central to scaling compute across a brand-new cloud account. You will architect multi-AZ, storage-heavy environments using FSx, EFS, NetApp ONTAP, and GPU visualization.

Qualifications

  • Proven HPC engineer with Windows Workspaces and Linux HPC experience.
  • Hands-on with ANSYS, Abaqus, LS-DYNA, CREO, NX, MATLAB workloads.
  • Strong AWS services background and parallel computing orchestration (PCS).
  • Expert PBS scheduler (v20.0.1) and HPC front-end portals (NICE EnginFrame, EF Portal, Open OnDemand).
  • Experience configuring FSx for Lustre/Windows, EFS, NetApp ONTAP, and GPU visualization.

Responsibilities

  • Modernize and migrate legacy HPC environments to cloud-native architectures.
  • Plan and deploy scalable HPC/workspace environments in a new AWS account.
  • Implement compute orchestration and remote desktop solutions (WorkSpaces, PCS, GPU).
  • Administer PBS schedulers and enable front-end portals for job submission and monitoring.
  • Manage high-performance storage (FSx, EFS, NetApp ONTAP) and ensure high availability.
  • Design multi-AZ elastic deployments for enterprise workloads.

Skills

HPC engineering
Windows Workspaces
Linux HPC
AWS PCS
PBS Scheduler
Front-end portals
NICE EnginFrame/Open OnDemand
FSx/Lustre & EFS
NetApp ONTAP
Terraform/CloudFormation/Ansible
Docker/Singularity
Python/Bash/PowerShell
Job orchestration

Tools

PBS Scheduler
NICE EnginFrame
EF Portal
Open OnDemand
AWS PCS
Terraform
CloudFormation
Ansible
Docker
Singularity/Apptainer
Nextflow/Snakemake

Job description

We are looking for a Lead HPC and Systems Engineer . If you are looking to give your career a real boost with a global leader in digital transformation, EPAM is the perfect choice. You will drive the automation and modernization of enterprise infrastructure, supporting Windows Workspaces and Linux HPC environments. High-Performance Computing (HPC) serves as the core engine behind simulation-driven product development, directly impacting time-to-market and engineering productivity while modernizing infrastructure to alleviate existing pain points and deploying into a new AWS account with primary workloads including Ansys, Simulia, ABAQUS, LS-DYNA, and MATLAB.

Responsibilities
  • Modernize & Migrate Environments: Spearhead the transition from legacy setups to modern, cloud-native HPC architectures, alleviating operational pain points
  • New Account Deployment: Plan, execute, and deploy scalable HPC and virtual workspace environments in a brand-new cloud account
  • Compute Orchestration & Workspaces: Implement scalable compute orchestration using AWS Parallel Computing Service (PCS) and deliver high-performance remote desktop solutions (Amazon WorkSpaces, NICE DCV on PCS, and GPU visualization capabilities)
  • Job Scheduling & Management: Administer and optimize PBS schedulers (v20.0.1) and implement modern front-end portals such as EF Portal or Open OnDemand for seamless job submission and monitoring
  • Storage Architecture: Manage, optimize, and scale high-performance storage solutions including FSx for Lustre, FSx for Windows, EFS, and NetApp ONTAP
  • Reliability & Elasticity: Design and maintain multi-AZ elastic deployments to ensure high availability and robust performance for engineering workloads
Requirements
  • Core Expertise: Proven professional experience as an HPC Engineer with deep expertise in both Windows Workspaces and Linux HPC environments
  • Workload Experience: Hands-on experience supporting and troubleshooting engineering simulation workloads including ANSYS, Abaqus, LS-DYNA, CREO, NX, and MATLAB
  • Cloud & Orchestration: Strong background in AWS services and scalable compute orchestration - specifically Parallel Computing Service (PCS)
  • Schedulers & Portals: Expert knowledge of the PBS scheduler (v20.0.1) and HPC front-end management portals (e.g., NICE EnginFrame, EF Portal, Open OnDemand)
  • Storage & Desktops: Extensive experience configuring and managing FSx for Lustre, FSx for Windows, EFS, and NetApp ONTAP, alongside high-performance remote desktop solutions and GPU visualization capabilities
  • Architecture Design: Demonstrated ability to architect and deploy multi-AZ elastic deployments supporting enterprise R&D environments
  • Nice to have Infrastructure as Code (IaC): Experience with Terraform, AWS CloudFormation, or Ansible for automating HPC cluster deployments and workspace provisioning
  • Containerization & Workflow Management: Familiarity with containers (Docker, Singularity/Apptainer) and workflow orchestration tools (e.g., Nextflow, Snakemake) in HPC environments
  • Cloud Certifications: Relevant AWS certifications (e.g., AWS Certified Solutions Architect - Professional or AWS Certified DevOps Engineer)
  • Scripting & Automation: Proficiency in Python, Bash, or PowerShell for automating administrative tasks, job monitoring, and telemetry collection
  • Cost Optimization: Experience implementing cloud cost management strategies, auto-scaling policies, and spot instance utilization for large-scale simulation workloads
  • Security & Compliance: Experience setting up secure HPC environments adhering to enterprise compliance standards, IAM policies, and encrypted storage configurations
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead HPC & Systems Engineer – Cloud-Based Simulations
Lead HPC & Systems Engineer – Cloud-Based Simulations

EPAM Systems • United States

Remote
USD 150,000 - 190,000
ZR_2817_JOB
ZR_2817_JOB

Zohorecruit • Jacksonville (FL)

Remote
USD 150,000 - 210,000
HPC on AWS Lead /Specialist/ SME- REMOTE
HPC on AWS Lead /Specialist/ SME- REMOTE

Simple Solutions • Jacksonville (FL)

Remote
USD 140,000 - 210,000
Senior AWS EDA/HPC Architect & SME - Remote
Senior AWS EDA/HPC Architect & SME - Remote

Simple Solutions • Jacksonville (FL)

Hybrid
USD 160,000 - 230,000
Heavy AWS + HPC
Heavy AWS + HPC

Zeal Solutions Inc • Charlotte (NC)

Remote
USD 120,000 - 180,000
HPC (High-Performance Computing) Consultant @ Remote
HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLC • United States

Remote
USD 120,000 - 180,000
Lead Platform Engineer/Architect - HPC, Kubernetes
Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 150,000 - 230,000
Lead HPC/Cloud Systems Administrator
Lead HPC/Cloud Systems Administrator

RedLine Performance Solutions, LLC. • Silver Spring (MD), Northern (KY)

Hybrid
USD 120,000 - 150,000
Remote work
Full benefits package
401k match
+1
HPC Customer Solutions Engineer
HPC Customer Solutions Engineer

GTN Technical Staffing • United States

On-site
USD 120,000 - 180,000
HPC Engineer
HPC Engineer

Tata Consultancy Services • Indianapolis (IN)

On-site
USD 75,000 - 80,000
Discretionary annual incentive
Comprehensive medical coverage (M/D/V,
Family support leaves
+6