We are looking for a Lead HPC and Systems Engineer . If you are looking to give your career a real boost with a global leader in digital transformation, EPAM is the perfect choice. You will drive the automation and modernization of enterprise infrastructure, supporting Windows Workspaces and Linux HPC environments. High-Performance Computing (HPC) serves as the core engine behind simulation-driven product development, directly impacting time-to-market and engineering productivity while modernizing infrastructure to alleviate existing pain points and deploying into a new AWS account with primary workloads including Ansys, Simulia, ABAQUS, LS-DYNA, and MATLAB.
Responsibilities
- Modernize & Migrate Environments: Spearhead the transition from legacy setups to modern, cloud-native HPC architectures, alleviating operational pain points
- New Account Deployment: Plan, execute, and deploy scalable HPC and virtual workspace environments in a brand-new cloud account
- Compute Orchestration & Workspaces: Implement scalable compute orchestration using AWS Parallel Computing Service (PCS) and deliver high-performance remote desktop solutions (Amazon WorkSpaces, NICE DCV on PCS, and GPU visualization capabilities)
- Job Scheduling & Management: Administer and optimize PBS schedulers (v20.0.1) and implement modern front-end portals such as EF Portal or Open OnDemand for seamless job submission and monitoring
- Storage Architecture: Manage, optimize, and scale high-performance storage solutions including FSx for Lustre, FSx for Windows, EFS, and NetApp ONTAP
- Reliability & Elasticity: Design and maintain multi-AZ elastic deployments to ensure high availability and robust performance for engineering workloads
Requirements
- Core Expertise: Proven professional experience as an HPC Engineer with deep expertise in both Windows Workspaces and Linux HPC environments
- Workload Experience: Hands-on experience supporting and troubleshooting engineering simulation workloads including ANSYS, Abaqus, LS-DYNA, CREO, NX, and MATLAB
- Cloud & Orchestration: Strong background in AWS services and scalable compute orchestration - specifically Parallel Computing Service (PCS)
- Schedulers & Portals: Expert knowledge of the PBS scheduler (v20.0.1) and HPC front-end management portals (e.g., NICE EnginFrame, EF Portal, Open OnDemand)
- Storage & Desktops: Extensive experience configuring and managing FSx for Lustre, FSx for Windows, EFS, and NetApp ONTAP, alongside high-performance remote desktop solutions and GPU visualization capabilities
- Architecture Design: Demonstrated ability to architect and deploy multi-AZ elastic deployments supporting enterprise R&D environments
- Nice to have Infrastructure as Code (IaC): Experience with Terraform, AWS CloudFormation, or Ansible for automating HPC cluster deployments and workspace provisioning
- Containerization & Workflow Management: Familiarity with containers (Docker, Singularity/Apptainer) and workflow orchestration tools (e.g., Nextflow, Snakemake) in HPC environments
- Cloud Certifications: Relevant AWS certifications (e.g., AWS Certified Solutions Architect - Professional or AWS Certified DevOps Engineer)
- Scripting & Automation: Proficiency in Python, Bash, or PowerShell for automating administrative tasks, job monitoring, and telemetry collection
- Cost Optimization: Experience implementing cloud cost management strategies, auto-scaling policies, and spot instance utilization for large-scale simulation workloads
- Security & Compliance: Experience setting up secure HPC environments adhering to enterprise compliance standards, IAM policies, and encrypted storage configurations