Job Title: HPC Lead & Operations Engineer
Experience: 8–10 Years
Role Overview
We are looking for an experienced HPC Lead & Operations Engineer to architect, operate, and support high-performance, cloud-based HPC platforms. You will be responsible for designing scalable infrastructure, ensuring operational excellence, and supporting scientific and AI-driven workloads across multi cloud environments.
This role requires a strong mix of HPC engineering, cloud operations, AI skill sets and user support, acting as a bridge between infrastructure and research teams.
Key Responsibilities
- Design, deploy, and manage cloud-based HPC environments, primarily on AWS
- Manage day‑to‑day cluster operations including monitoring, incident handling, troubleshooting, and patching
- Provide L2‑L3 application support for scientific and computational workloads such as
- Build and maintain CI/CD pipelines and Infrastructure-as-Code (IaC) for automated and repeatable deployments
- Administer and optimize job schedulers (SLURM/PBS) for efficient resource utilization
- Drive cost optimization, capacity planning, and auto‑scaling strategies
- Support AI/ML workloads running on HPC or hybrid infrastructure
- Mentor team members and maintain clear operational documentation and runbooks
Must-Have Skills & Experience
- 8–10 years of experience in HPC administration and cloud infrastructure
- Strong multi‑cloud experience across AWS and Google Cloud Platform (GCP) with expertise in HPC and AI/ML workloads
- AWS: EC2, ParallelCluster, FSx, EFS, S3
- GCP: Compute Engine, Filestore, Cloud Storage, HPC Toolkit (or equivalent)
- Hands‑on experience supporting AI/ML workloads or frameworks (e.g., TensorFlow, PyTorch, distributed training environments) and Automations/Innovations
- Proficiency in Python and Bash scripting
- Experience with
- Schedulers: SLURM, PBS
- Parallel file systems: Lustre, GPFS, or equivalent
- Containers: Docker, Singularity/Apptainer
- Expertise in Infrastructure-as-Code & automation tools
- Terraform, CloudFormation, Packer, Ansible, Git
- Proven experience in L2/L3 support for scientific or HPC applications
- Familiarity with monitoring and observability tools
- CloudWatch, Prometheus, Grafana, or similar
Nice-to-Have
- Background in life sciences, bioinformatics, or drug discovery
- Hands‑on experience with Schrödinger Suite
- Experience with Kubernetes (EKS/GKE) for HPC, Terraform AI workloads
- Exposure to GPU computing and distributed training environments
- Certifications such as AWS Solutions Architect, Google Cloud Professional Architect, or RHCE
- Afternoon shift (2PM to 11PM IST) with flexibility for global collaboration