Platform Engineer (ML Infrastructure)

TANUH

Bengaluru

On-site

INR 2,500,000 - 4,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TANUH is seeking a Systems Engineer / Site Reliability Engineer to own our bare‑metal ML infrastructure in Bengaluru. You will manage a multi‑node GPU cluster, bridge HPC hardware with AI researchers, and ensure secure, optimized environments for training healthcare models.

Responsibilities include deploying OpenHPC/SLURM, Linux administration on Rocky Linux, and deploying ML frameworks for distributed training. Strong scripting and cloud integration experience is valued.

Qualifications

  • 4–6 years in Systems Engineer, SRE, or Linux admin in high-performance environments.
  • Deep Linux expertise (Rocky Linux/CentOS/RHEL), OS security & permissions.
  • Hands-on with GPU-accelerated servers, NVIDIA architectures and drivers.
  • Experience with SLURM/OpenHPC and workload management, distributed ML support.

Responsibilities

  • HPC & cluster management: deploy, configure, and maintain our GPU cluster with OpenHPC/SLURM.
  • Linux admin: secure, scalable user/group management and storage access.
  • GPU infra: install/troubleshoot NVIDIA drivers, CUDA, cuDNN, MIG configurations.
  • ML platform support: optimize PyTorch/JAX/TensorFlow for multi-node training.
  • Monitoring & optimization: track SLURM queues, diagnose bottlenecks, resolve issues.
  • Automation & scripting: Bash/Python/Ansible for provisioning and onboarding.

Skills

Linux administration
HPC cluster management
Python scripting
Bash scripting
GPU hardware & CUDA
ML framework deployment
Documentation & comms
Cloud integration (GCP)

Tools

SLURM
OpenHPC
CUDA toolkit / cuDNN
Google Cloud SDKs/CLI
Ansible

Job description

We are seeking a highly skilled Systems Engineer / Site Reliability Engineer (SRE) to take ownership of our state‑of‑the‑art, bare‑metal Machine Learning infrastructure. In this role, you will be the bridge between our high‑performance computing (HPC) hardware and our AI researchers. You will be responsible for provisioning, deploying, managing, and troubleshooting our multi‑node GPU cluster, ensuring our data scientists have a seamless, secure, and highly optimized environment to train foundational models for healthcare.

Key Responsibilities
  • HPC & Cluster Management: Deploy, configure, and maintain our AI computing cluster using OpenHPC and SLURM. Manage workload queues, allocate resources efficiently, and prevent bottlenecks.
  • Linux Systems Administration: Administer our core OS (Rocky Linux), implementing robust user and group management, directory structures, network storage permissions (NFS/XFS), and strict security policies across segregated project verticals.
  • GPU Infrastructure: Take full ownership of the NVIDIA hardware stack. Install, update, and troubleshoot NVIDIA drivers, CUDA toolkits, cuDNN, and multi‑instance GPU (MIG) configurations to maximize hardware utilization.
  • ML Platform Support: Deploy and maintain machine learning environments. Identify dependencies and optimize the installation of frameworks like PyTorch, JAX, and TensorFlow for distributed, multi‑node training.
  • Monitoring & Optimization: Periodically monitor cluster health, node performance, and SLURM job queues. Identify workload imbalances, debug failed jobs, and proactively resolve hardware/software conflicts.
  • Automation & Scripting: Design and implement automation scripts (Bash, Python, Ansible) for routine cluster management, provisioning, user onboarding, and data synchronization tasks.
  • Hybrid Cloud & Data Pipelines: Manage the integration between our on‑premise bare‑metal cluster and Google Cloud Platform (GCP). Install and maintain Google Cloud SDKs and CLIs across the environment. Configure IAM Service Accounts and automate secure, high‑throughput storage operations to synchronize massive "golden" healthcare datasets and model checkpoints between Google Cloud Storage (GCS) and our local NFS/XFS shared drives.
  • Documentation & Enablement: Write clear, comprehensive documentation, usage policies, and Standard Operating Procedures (SOPs). Act as a technical guide to help clinical researchers and data scientists use the SLURM cluster efficiently.
Qualified Candidates
  • 4‑6 years of hands‑on experience as a Systems Engineer, SRE, or Linux Administrator in a high‑performance or heavy‑compute environment.
  • Deep proficiency in Linux administration (specifically Rocky Linux, CentOS, or RHEL), including advanced file permissions, user lifecycle management, and OS security.
  • Experience in managing GPU‑accelerated servers, with a strong understanding of NVIDIA architectures, CUDA installations, and driver troubleshooting.
  • Demonstrated hands‑on experience with workload managers and job schedulers, specifically SLURM and OpenHPC.
  • Solid understanding of the Python ecosystem and experience deploying ML frameworks (PyTorch, JAX) in virtual environments (Miniforge, Anaconda).
  • Strong scripting skills for system automation (Bash, Python).
  • Cloud Integration Experience (Optional): Hands‑on experience with cloud platforms (GCP preferred). Proficient in deploying and using cloud command‑line tools, managing cloud storage buckets, configuring IAM permissions for headless machine accounts, and scripting secure hybrid data transfers.
  • Excellent written communication skills with a track record of authoring clear technical documentation for non‑systems engineers.
Benefits

You will be managing the infrastructure that directly powers life‑saving AI models. Your work will enable faster, more accurate detection of diseases like oral cancer, breast cancer, and diabetes, transforming frontline healthcare across India.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

SourcingXPress • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
DevOps Engineer
DevOps Engineer

Recrew AI • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Competitive compensation
Access to cutting-edge GPU compute infrastructure
Opportunity to work alongside AI researchers
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure

NVIDIA • Mumbai

On-site
INR 1,800,000 - 2,500,000
Competitive salary
Generous benefits package
Diversity and inclusion commitment
Lead HPC Engineer
Lead HPC Engineer

Clovertex • Hyderabad

On-site
INR 2,000,000 - 3,000,000
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure
Senior Solutions Architect, GPU Cloud GenAI – Infrastructure

NVIDIA Gruppe • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Generous benefits package
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000