Lead HPC Engineer

Clovertex

Hyderabad

On-site

INR 2,000,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Clovertex is seeking an experienced HPC Lead & Operations Engineer based in Hyderabad, India. The role involves architecting, operating, and supporting high-performance, cloud-based HPC platforms across multi-cloud environments. The ideal candidate will have a strong background in HPC administration, cloud infrastructure, and a passion for supporting scientific workloads.

Applicants should have 8-10 years of experience in managing HPC systems, particularly with AWS and GCP, and possess strong scripting skills. This role includes responsibilities in infrastructure design, cost optimization, and team mentorship.

Qualifications

  • 8–10 years of experience in HPC administration and managing cloud infrastructure.
  • Expertise in multi-cloud environments, particularly AWS and GCP for HPC.
  • Hands-on experience with AI/ML workloads and frameworks like TensorFlow or PyTorch.

Responsibilities

  • Design and manage cloud-based HPC environments on AWS.
  • Oversee daily operations, incident handling, and troubleshooting.
  • Create and manage CI/CD pipelines for automated deployments.

Skills

HPC administration
Cloud infrastructure
AWS
Google Cloud Platform (GCP)
AI/ML frameworks
Python
Bash scripting
Terraform
Infrastructure-as-Code
Monitoring tools

Tools

Docker
SLURM
PBS
Grafana
Prometheus

Job description

Job Title: HPC Lead & Operations Engineer
Experience: 8–10 Years
Role Overview

We are looking for an experienced HPC Lead & Operations Engineer to architect, operate, and support high-performance, cloud-based HPC platforms. You will be responsible for designing scalable infrastructure, ensuring operational excellence, and supporting scientific and AI-driven workloads across multi cloud environments.

This role requires a strong mix of HPC engineering, cloud operations, AI skill sets and user support, acting as a bridge between infrastructure and research teams.

Key Responsibilities
  • Design, deploy, and manage cloud-based HPC environments, primarily on AWS
  • Manage day‑to‑day cluster operations including monitoring, incident handling, troubleshooting, and patching
  • Provide L2‑L3 application support for scientific and computational workloads such as
  • Build and maintain CI/CD pipelines and Infrastructure-as-Code (IaC) for automated and repeatable deployments
  • Administer and optimize job schedulers (SLURM/PBS) for efficient resource utilization
  • Drive cost optimization, capacity planning, and auto‑scaling strategies
  • Support AI/ML workloads running on HPC or hybrid infrastructure
  • Mentor team members and maintain clear operational documentation and runbooks
Must-Have Skills & Experience
  • 8–10 years of experience in HPC administration and cloud infrastructure
  • Strong multi‑cloud experience across AWS and Google Cloud Platform (GCP) with expertise in HPC and AI/ML workloads
    • AWS: EC2, ParallelCluster, FSx, EFS, S3
    • GCP: Compute Engine, Filestore, Cloud Storage, HPC Toolkit (or equivalent)
  • Hands‑on experience supporting AI/ML workloads or frameworks (e.g., TensorFlow, PyTorch, distributed training environments) and Automations/Innovations
  • Proficiency in Python and Bash scripting
  • Experience with
    • Schedulers: SLURM, PBS
    • Parallel file systems: Lustre, GPFS, or equivalent
    • Containers: Docker, Singularity/Apptainer
  • Expertise in Infrastructure-as-Code & automation tools
  • Terraform, CloudFormation, Packer, Ansible, Git
  • Proven experience in L2/L3 support for scientific or HPC applications
  • Familiarity with monitoring and observability tools
    • CloudWatch, Prometheus, Grafana, or similar
Nice-to-Have
  • Background in life sciences, bioinformatics, or drug discovery
  • Hands‑on experience with Schrödinger Suite
  • Experience with Kubernetes (EKS/GKE) for HPC, Terraform AI workloads
  • Exposure to GPU computing and distributed training environments
  • Certifications such as AWS Solutions Architect, Google Cloud Professional Architect, or RHCE
  • Afternoon shift (2PM to 11PM IST) with flexibility for global collaboration
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Engineer
Senior HPC Engineer

Netweb Technologies India Ltd. • Faridabad District

On-site
INR 1,500,000 - 2,100,000
HPC Engineer
HPC Engineer

Clovertex • Hyderabad

On-site
INR 1,200,000 - 2,400,000
HPC Engineer
HPC Engineer

Whiteblue • Chennai

On-site
INR 1,500,000 - 2,500,000
AWS Senior HPC Engineer SME
AWS Senior HPC Engineer SME

Tata Consultancy Services • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,800,000 - 3,200,000
HPC Engineer
HPC Engineer

Yotta Data Services Private Limited • Mumbai

On-site
INR 400,000 - 700,000
Senior HPC Engineer
Senior HPC Engineer

Binaire Private Limited • New Delhi

On-site
INR 1,500,000 - 2,500,000
Opportunity to influence hardware selection
Ownership of high-performance compute infrastructure
Linux System Administrator
Linux System Administrator

SISL Global • Chennai District

On-site
INR 800,000 - 1,200,000
Lead Solutions Architect – AI Infrastructure
Lead Solutions Architect – AI Infrastructure

Cortex Consultants LLC • Bengaluru Urban

On-site
INR 1,800,000 - 2,400,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA AI • Bengaluru

On-site
INR 3,000,000 - 4,200,000