T&T | EAD | Senior Consultant/Manager | AI Infra GPU | PAN India

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft

Bengaluru

On-site

INR 3,000,000 - 5,500,000

Full time

4 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft seeks a Senior Consultant/Manager for AI Infra GPU across PAN India. The role focuses on NVIDIA GPU infrastructure, cluster design, and hybrid cloud integration to support distributed training and low-latency inference.

You will build GPU landing zones, select GPU instances, engineer clusters with Kubernetes or HPC schedulers, and ensure observability, automation and FinOps for production AI workloads.

Qualifications

  • 2-12 years of relevant experience in infrastructure, SRE, HPC or platform engineering.
  • Hands-on experience with GPU or accelerated-computing environments.
  • Strong implementation, automation, testing, troubleshooting and technical-documentation skills.
  • Strong understanding of NVIDIA GPU architecture and systems, including DGX/HGX or equivalent certified platforms.
  • Demonstrated experience in GPU cluster design across compute, high-speed networking, storage, control plane and management plane.
  • Understanding of AI workloads including distributed training, fine-tuning, RAG, batch inference, real-time inference and HPC.
  • Proficiency in Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
  • Strong Linux, container, CUDA, NCCL, driver, firmware and GPU observability fundamentals.
  • Experience with security, resilience, capacity, performance, automation and day-2 operations for production AI infrastructure.
  • Deep expertise in at least one of AWS, Microsoft Azure or Google Cloud, with working awareness of the other major platforms.

Responsibilities

  • Design GPU landing zones covering accounts/subscriptions/projects, network topology, private connectivity, identity, encryption, policy and observability.
  • Select NVIDIA GPU instances and cluster patterns for distributed training, fine-tuning, batch inference and low-latency serving.
  • Engineer cloud GPU clusters using managed Kubernetes or HPC schedulers, placement/topology controls and high-performance network adapters.
  • Design high-throughput object, file and block storage, data ingestion, checkpointing, caching and cross-region/data-centre movement patterns.
  • Build hybrid connectivity and workload portability between private GPU clusters and public cloud.
  • Implement Terraform, image pipelines, CI/CD/GitOps, autoscaling, quota automation, reservations/capacity blocks and environment promotion.
  • Integrate cloud ML services where appropriate while retaining infrastructure controls for custom NVIDIA-based workloads.
  • Establish observability for GPU availability, utilization, tokens, latency, throughput, reliability and cost.
  • Implement AI infrastructure FinOps covering commitments, spot/preemptible usage, idle

Skills

GPU infrastructure
Kubernetes
OpenShift
Slurm
Linux
CI/CD
Terraform
NVIDIA GPUs
NCCL
Cloud platforms

Education

Any Degree

Tools

Terraform
Kubernetes/OpenShift
CI/CD pipelines
GitOps
HPC schedulers

Job description

Select how often (in days) to receive an alert:

Job Title: T&T | EAD | Senior Consultant/Manager | AI Infra GPU | PAN India

The Team

Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information that can support fact-driven decision-making and generate an insight-driven advantage. Our practice addresses the continuum of opportunities in business intelligence & visualization, data management, performance management and next-generation analytics and technologies, including big data, cloud, cognitive and machine learning.

  • Strong expertise in NVIDIA GPU infrastructure, architecture and production AI platforms.
  • Strong implementation and automation skills across cloud GPU infrastructure.
  • Experience with GPU cluster design, high-performance networking, storage and cloud-native orchestration.
  • Strong understanding of Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
  • Strong Linux, containers, CUDA ecosystem, NCCL, drivers, firmware and GPU observability fundamentals.
  • Experience with Infrastructure as Code, Terraform, CI/CD and GitOps.
  • Strong troubleshooting, benchmarking, testing and operational handover capabilities.

Experience:

  • 2-12 Years

Key Responsibilities:

  • Design GPU landing zones covering accounts/subscriptions/projects, network topology, private connectivity, identity, encryption, policy and observability.
  • Select NVIDIA GPU instances and cluster patterns for distributed training, fine-tuning, batch inference and low-latency serving.
  • Engineer cloud GPU clusters using managed Kubernetes or HPC schedulers, placement/topology controls and high-performance network adapters.
  • Design high-throughput object, file and block storage, data ingestion, checkpointing, caching and cross-region/data-centre movement patterns.
  • Build hybrid connectivity and workload portability between private GPU clusters and public cloud.
  • Implement Terraform, image pipelines, CI/CD/GitOps, autoscaling, quota automation, reservations/capacity blocks and environment promotion.
  • Integrate cloud ML services where appropriate while retaining infrastructure controls for custom NVIDIA-based workloads.
  • Establish observability for GPU availability, utilization, tokens, latency, throughput, reliability and cost.
  • Implement AI infrastructure FinOps covering commitments, spot/preemptible usage, idle

Required Qualifications & Skills:

  • 2-12 years of relevant experience in infrastructure, SRE, HPC or platform engineering.
  • Hands-on experience with GPU or accelerated-computing environments.
  • Strong implementation, automation, testing, troubleshooting and technical-documentation skills.
  • Strong understanding of NVIDIA GPU architecture and systems, including DGX/HGX or equivalent certified platforms.
  • Demonstrated experience in GPU cluster design across compute, high-speed networking, storage, control plane and management plane.
  • Understanding of AI workloads including distributed training, fine-tuning, RAG, batch inference, real-time inference and HPC.
  • Proficiency in Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
  • Strong Linux, container, CUDA, NCCL, driver, firmware and GPU observability fundamentals.
  • Experience with security, resilience, capacity, performance, automation and day-2 operations for production AI infrastructure.
  • Deep expertise in at least one of AWS, Microsoft Azure or Google Cloud, with working awareness of the other major platforms.
  • Experience with cloud GPU capacity, high-performance networking, managed Kubernetes/HPC, Infrastructure as Code

Education:

  • Any Degree, Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering or a related technical discipline is preferred.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Solutions Architect
Senior Solutions Architect

E2E Networks Limited • Delhi

On-site
INR 400,000 - 700,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA Gruppe • Gurugram District

On-site
INR 3,000,000 - 6,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000
Specialist - System Management
Specialist - System Management

Vivantify • India

On-site
INR 1,300,000 - 2,100,000
Competitive salary and benefits
Professional growth opportunities
Collaborative work environment
+1
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA • Gurugram District

On-site
INR 3,500,000 - 7,500,000
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Bengaluru

On-site
INR 1,500,000 - 2,500,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000
NVIDIA - AI Infrastructure Specialist (GDC) - 62808 (DEAI DS) IoT India
NVIDIA - AI Infrastructure Specialist (GDC) - 62808 (DEAI DS) IoT India

Hitachids • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Industry-leading benefits
Flexible work arrangements
Inclusive culture
NOC Technical Lead (L3)
NOC Technical Lead (L3)

Larsen & Toubro • Chennai District

On-site
INR 4,500,000 - 7,500,000