Sr. Principal Infrastr Eng (HPC)

Mphasis

Bengaluru

On-site

INR 3,600,000 - 6,000,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mphasis seeks a Senior Principal Infrastructure Engineer to administer and optimize ClearML Server and related infrastructure in a demanding AI/ML environment. You will drive MLOps platform stability, HPC integration, and containerized workloads across CPU/GPU nodes.

Responsibilities include configuring agents, managing execution queues and user access, enforcing security, and automating deployment, monitoring, and troubleshooting with Python, Bash, Ansible, and Git.

Qualifications

  • Strong Linux admin skills (RHEL/SLES/Rocky/Ubuntu).
  • Hands-on ClearML or similar MLOps platform experience.
  • Familiarity with HPC schedulers (Slurm/PBS/LSF).
  • Knowledge of GPU platforms including NVIDIA drivers and CUDA.
  • Proficient in container tech: Docker, Singularity, Kubernetes.
  • Strong scripting/automation: Python, Bash, Ansible, Git.
  • Security fundamentals: RBAC, IAM, TLS, secrets mgmt, audit logging.
  • Confidential computing concepts: TEE, TPM, AMD/Intel features.

Responsibilities

  • Administer ClearML Server, including agents, queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
  • Configure ClearML Agents on CPU and GPU worker nodes; integrate with HPC platforms (Slurm, PBS Pro, Kubernetes).
  • Support GPU-based AI/ML workloads with NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
  • Maintain secure container execution using Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Implement confidential-computing controls (AMD SEV-SNP, Intel TDX, NVIDIA CC, Secure Boot, TPM).
  • Integrate authentication, RBAC, TLS, secrets management, audit controls for ClearML and confidential workloads.
  • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, platform health.
  • Automate deployment, configuration, monitoring, and troubleshooting with Python, Bash, Ansible, Git.

Skills

Linux administration
ClearML administration
HPC schedulers (Slurm/PBS/LSF)
NVIDIA drivers & CUDA
Container technologies (Docker, Singul

Tools

Docker
Kubernetes
Singularity

Job description

We are seeking a highly skilled Senior Principal Infrastructure Engineer to administer and optimize our ClearML Server and associated infrastructure. The ideal candidate will have a strong background in MLOps platforms, HPC execution environments, and containerized solutions, with a focus on supporting GPU-based AI/ML workloads. This role requires a proactive approach to managing resources, ensuring security, and automating processes to enhance operational efficiency.

Responsibilities:

  • Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
  • Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Professional, or Kubernetes.
  • Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
  • Maintain secure container execution using technologies like Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation.
  • Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads.
  • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health.
  • Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git.

Mandatory Skills:

  • Strong Linux administration skills, particularly with RHEL, SLES, Rocky Linux, or Ubuntu.
  • Hands-on experience with ClearML administration or a comparable MLOps platform.
  • Familiarity with HPC schedulers such as Slurm, PBS Professional, or LSF.
  • Knowledge of GPU platforms, including NVIDIA drivers, CUDA, and distributed training basics.
  • Proficiency in container technologies: Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Strong scripting and automation skills in Python, Bash, Ansible, and Git.
  • Understanding of security fundamentals, including RBAC, IAM, TLS, certificates, secrets management, secure boot, and audit logging.
  • Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, AMD SEV-SNP, Intel TDX, or equivalent technologies.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Principal Infrastructure Engineer
Senior Principal Infrastructure Engineer

Mphasis • Bengaluru

On-site
INR 400,000 - 700,000
HPC EXPERT – CLEARML & CONFIDENTIAL COMPUTING
HPC EXPERT – CLEARML & CONFIDENTIAL COMPUTING

Vrinda International • Bengaluru

On-site
INR 1,260,000 - 2,100,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Corporation • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Sr. HPC ENGINEER
Sr. HPC ENGINEER

Cognizant • Hyderabad

On-site
INR 800,000 - 1,200,000
Lead Engineer (HPC, GPU, CUDA)
Lead Engineer (HPC, GPU, CUDA)

AIRA Matrix • Thane

On-site
INR 1,200,000 - 2,500,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000
Principal Software Architect- High Performance Computing
Principal Software Architect- High Performance Computing

Applied Materials India • Chennai District

On-site
INR 4,000,000 - 6,000,000
Lead / Principal MLOps Engineer
Lead / Principal MLOps Engineer

Nextloop Technologies LLP • Chennai District

On-site
INR 4,000,000 - 7,000,000