Lead HPC Engineer

Clovertex

Hyderabad

On-site

INR 2,000,000 - 3,000,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Clovertex is seeking an experienced HPC Lead & Operations Engineer based in Hyderabad, India. The role involves architecting, operating, and supporting high-performance, cloud-based HPC platforms across multi-cloud environments. The ideal candidate will have a strong background in HPC administration, cloud infrastructure, and a passion for supporting scientific workloads.

Applicants should have 8-10 years of experience in managing HPC systems, particularly with AWS and GCP, and possess strong scripting skills. This role includes responsibilities in infrastructure design, cost optimization, and team mentorship.

Qualifications

  • 8–10 years of experience in HPC administration and managing cloud infrastructure.
  • Expertise in multi-cloud environments, particularly AWS and GCP for HPC.
  • Hands-on experience with AI/ML workloads and frameworks like TensorFlow or PyTorch.

Responsibilities

  • Design and manage cloud-based HPC environments on AWS.
  • Oversee daily operations, incident handling, and troubleshooting.
  • Create and manage CI/CD pipelines for automated deployments.

Skills

HPC administration
Cloud infrastructure
AWS
Google Cloud Platform (GCP)
AI/ML frameworks
Python
Bash scripting
Terraform
Infrastructure-as-Code
Monitoring tools

Tools

Docker
SLURM
PBS
Grafana
Prometheus

Job description

Job Title: HPC Lead & Operations Engineer
Experience: 8–10 Years
Role Overview

We are looking for an experienced HPC Lead & Operations Engineer to architect, operate, and support high-performance, cloud-based HPC platforms. You will be responsible for designing scalable infrastructure, ensuring operational excellence, and supporting scientific and AI-driven workloads across multi cloud environments.

This role requires a strong mix of HPC engineering, cloud operations, AI skill sets and user support, acting as a bridge between infrastructure and research teams.

Key Responsibilities
  • Design, deploy, and manage cloud-based HPC environments, primarily on AWS
  • Manage day‑to‑day cluster operations including monitoring, incident handling, troubleshooting, and patching
  • Provide L2‑L3 application support for scientific and computational workloads such as
  • Build and maintain CI/CD pipelines and Infrastructure-as-Code (IaC) for automated and repeatable deployments
  • Administer and optimize job schedulers (SLURM/PBS) for efficient resource utilization
  • Drive cost optimization, capacity planning, and auto‑scaling strategies
  • Support AI/ML workloads running on HPC or hybrid infrastructure
  • Mentor team members and maintain clear operational documentation and runbooks
Must-Have Skills & Experience
  • 8–10 years of experience in HPC administration and cloud infrastructure
  • Strong multi‑cloud experience across AWS and Google Cloud Platform (GCP) with expertise in HPC and AI/ML workloads
    • AWS: EC2, ParallelCluster, FSx, EFS, S3
    • GCP: Compute Engine, Filestore, Cloud Storage, HPC Toolkit (or equivalent)
  • Hands‑on experience supporting AI/ML workloads or frameworks (e.g., TensorFlow, PyTorch, distributed training environments) and Automations/Innovations
  • Proficiency in Python and Bash scripting
  • Experience with
    • Schedulers: SLURM, PBS
    • Parallel file systems: Lustre, GPFS, or equivalent
    • Containers: Docker, Singularity/Apptainer
  • Expertise in Infrastructure-as-Code & automation tools
  • Terraform, CloudFormation, Packer, Ansible, Git
  • Proven experience in L2/L3 support for scientific or HPC applications
  • Familiarity with monitoring and observability tools
    • CloudWatch, Prometheus, Grafana, or similar
Nice-to-Have
  • Background in life sciences, bioinformatics, or drug discovery
  • Hands‑on experience with Schrödinger Suite
  • Experience with Kubernetes (EKS/GKE) for HPC, Terraform AI workloads
  • Exposure to GPU computing and distributed training environments
  • Certifications such as AWS Solutions Architect, Google Cloud Professional Architect, or RHCE
  • Afternoon shift (2PM to 11PM IST) with flexibility for global collaboration
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud HPC Infrastructure Engineer
Cloud HPC Infrastructure Engineer

Amgen Inc • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Sr. HPC ENGINEER
Sr. HPC ENGINEER

Cognizant • Hyderabad

On-site
INR 800,000 - 1,200,000
HPC Engineer
HPC Engineer

Clovertex • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Lead Engineer - R01571349
Lead Engineer - R01571349

Brillio • Bengaluru

On-site
INR 2,400,000 - 4,200,000
HPC Engineer
HPC Engineer

Whiteblue • Chennai

On-site
INR 1,500,000 - 2,500,000
Senior HPC Engineer SME/Architect
Senior HPC Engineer SME/Architect

Tata Consultancy Services • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,800,000 - 3,000,000
Lead Engineer - R01571349
Lead Engineer - R01571349

Brillio 2 • Bengaluru

On-site
INR 2,400,000 - 4,200,000
HPC Cloud Engineer
HPC Cloud Engineer

5 Star Recruitment • Chennai District

On-site
INR 1,500,000 - 3,500,000
Specialist Software Engineer
Specialist Software Engineer

Amgen SA • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Lead Engineer – R01571349
Lead Engineer – R01571349

Brillio • Bengaluru

On-site
INR 1,800,000 - 2,400,000