Senior Principal Infrastructure Engineer

Mphasis

Bengaluru

On-site

INR 400,000 - 700,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mphasis in Bengaluru is seeking a Senior Principal Infrastructure Engineer to administer and optimize the ClearML Server and related infrastructure for GPU-enabled AI/ML workloads. Strong MLOps and HPC experience is essential.

You will configure agents, integrate with Slurm/PBS/Kubernetes, ensure security, automate deployments with Python and Ansible, and monitor platform health to keep workloads running efficiently.

Qualifications

  • Strong Linux administration experience (RHEL/SLES/Ubuntu).
  • Hands-on administration of ClearML or similar MLOps platforms.
  • Familiarity with HPC schedulers like Slurm, PBS Professional, or LSF.
  • Knowledge of GPU platforms including NVIDIA drivers, CUDA, and distributed training basics.
  • Proficiency with container tech: Apptainer/Singularity, Docker, Kubernetes.
  • Scripting and automation skills in Python, Bash, Ansible, and Git.
  • Understanding of security fundamentals: RBAC, IAM, TLS, secrets management, audit logs.
  • Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, and AMD SEV-SNP/Intel TDX.

Responsibilities

  • Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
  • Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Professional, or Kubernetes.
  • Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
  • Maintain secure container execution using Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation.
  • Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads.
  • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health.
  • Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git.

Skills

Linux administration
ClearML administration
HPC schedulers
NVIDIA CUDA
Python scripting
Bash scripting
Automation with Ansible
Git
RBAC/IAM/TLS security
Confidential computing concepts
Container technologies
GPU/AI workloads

Education

Bachelor's degree in Computer Science / IT / Engineering

Tools

ClearML
Slurm
PBS Professional
LSF
Docker
Kubernetes
Apptainer/Singularity
NVIDIA drivers
CUDA
NCCL
UCX

Job description

Job Summary

We are seeking a highly skilled Senior Principal Infrastructure Engineer to administer and optimize our ClearML Server and associated infrastructure. The ideal candidate will have a strong background in MLOps platforms, HPC execution environments, and containerized solutions, with a focus on supporting GPU-based AI/ML workloads. This role requires a proactive approach to managing resources, ensuring security, and automating processes to enhance operational efficiency.

Responsibilities
  • Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
  • Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Professional, or Kubernetes.
  • Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
  • Maintain secure container execution using technologies like Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation.
  • Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads.
  • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health.
  • Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git.
Mandatory Skills
  • Strong Linux administration skills, particularly with RHEL, SLES, Rocky Linux, or Ubuntu.
  • Hands-on experience with ClearML administration or a comparable MLOps platform.
  • Familiarity with HPC schedulers such as Slurm, PBS Professional, or LSF.
  • Knowledge of GPU platforms, including NVIDIA drivers, CUDA, and distributed training basics.
  • Proficiency in container technologies: Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
  • Strong scripting and automation skills in Python, Bash, Ansible, and Git.
  • Understanding of security fundamentals, including RBAC, IAM, TLS, certificates, secrets management, secure boot, and audit logging.
  • Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, AMD SEV-SNP, Intel TDX, or equivalent technologies.
Preferred Skills
  • Experience with the Nvidia Nemo stack on Cloud.
  • Knowledge of advanced monitoring and logging tools for infrastructure management.
  • Familiarity with cloud platforms and services related to AI/ML workloads.
Qualifications

A degree in Computer Science, Information Technology, Engineering, or a related field is preferred. Relevant certifications in cloud computing, MLOps, or infrastructure management will be considered an advantage.

About Mphasis

Mphasis applies to next-generation technology to help enterprises transform businesses globally. Customer centricity is foundational to Mphasis and is reflected in the Mphasis' Front2Back Transformation approach. Front2Back uses the exponential power of cloud and cognitive to provide hyper-personalized (C=X2C2TM=1) digital experience to clients and their end customers. Mphasis' Service Transformation approach helps 'shrink the core' through the application of digital technologies across legacy environments within an enterprise, enabling businesses to stay ahead in a changing world. Mphasis' core reference architectures and tools, speed and innovation with domain expertise and specialization are key to building strong relationships with marquee clients.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Principal Infrastr Eng (HPC)
Sr. Principal Infrastr Eng (HPC)

Mphasis • Bengaluru

On-site
INR 3,600,000 - 6,000,000
Senior Full Stack Developer - Cloud Infrastructure & DevOps
Senior Full Stack Developer - Cloud Infrastructure & DevOps

Mphasis Ltd • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Principal AI Software Engineer
Principal AI Software Engineer

Mphasis • Maharashtra

On-site
INR 4,000,000 - 7,000,000
Cloud Infrastructure and DevOps Architect Engineer
Cloud Infrastructure and DevOps Architect Engineer

Mphasis Ltd • Bengaluru

On-site
INR 2,400,000 - 4,800,000
Senior Full Stack Engineer (Java / Node.js / AWS)
Senior Full Stack Engineer (Java / Node.js / AWS)

Mphasis Ltd • Bengaluru

On-site
INR 3,200,000 - 4,600,000
Senior Tech Lead/Architect
Senior Tech Lead/Architect

Cloudxtreme • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Technical Lead
Technical Lead

Mphasis • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Senior MLOps Engineer - DSX Enablement
Senior MLOps Engineer - DSX Enablement

NVIDIA • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Module Lead - Systems
Module Lead - Systems

Mphasis • Hyderabad

On-site
INR 2,500,000 - 4,200,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000