Senior HPC Platform Architect

NVIDIA Gruppe

Bengaluru

On-site

INR 400,000 - 900,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking an exceptional engineer to join the HPC Infrastructure team responsible for design, review, and optimization of compute infrastructure powering AI workloads and silicon design. You will be the primary architecture reviewer and performance champion for new data center clusters across sites globally.

Responsibilities include leading architecture reviews, benchmarking, and tuning at multiple layers, collaborating with vendors, and driving operational readiness for tapeout

Qualifications

  • B.E./B.Tech or M.Tech/M.S. with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale.
  • Deep understanding of data center architecture fundamentals: compute, storage, and high-speed networking.
  • Proven ability to evaluate and challenge infrastructure design decisions across rack layout and cooling constraints.
  • Experience with OS/kernel performance tuning and system parameter optimization.
  • Hands-on experience with large-scale Linux HPC cluster administration using workload managers such as LSF and Slurm.
  • Strong Linux/Unix administration and scripting skills (Python, Bash, Perl).
  • Experience running HPC performance benchmarks and profiling tools.
  • Excellent problem-solving and communication skills for leadership presentation.

Responsibilities

  • Own data center architecture reviews for new HPC clusters — evaluate compute, storage, networking, cooling decisions and challenge assumptions.
  • Serve as BDC representative in cluster build meetings across multiple programs, driving architecture alignment.
  • Analyze and validate cluster design choices, surface risk and tradeoffs to leadership.
  • Lead performance benchmarking and profiling of HPC infrastructure, run sanity benchmarks and health checks.
  • Drive infrastructure optimization at scheduler, hardware, and OS/kernel levels.
  • Collaborate with platform and operations teams on cluster health and capacity planning.
  • Partner with vendors to evaluate new hardware and networking fabrics; produce data-backed recommendations.
  • Improve observability, benchmarking frameworks, and architecture documentation across the BDC HPC platform.

Skills

HPC infrastructure
Data center architecture
Linux system administration
Scripting (Python/Bash/Perl)
LSF/Slurm scheduling
Performance benchmarking
Networking (InfiniBand/Ethernet)
Problem solving & communication

Education

BE/BTech or ME/MTech with 5+ years in HPC

Tools

LSF
Slurm

Job description

NVIDIA is looking for an exceptional engineer to grow and thrive alongside our HPC Infrastructure team that designs, evaluates, and optimizes the compute inrastructure powering NVIDIA's next-generation silicon design and AI workloads. You will be the primary architecture reviewer and performance champion for new data center clusters being built across multiple sites globally.

What you'll be doing
  • Own data center architecture reviews for new HPC clusters — evaluating compute, storage,networking, and cooling decisions and challenging assumptions before clusters are built.
  • Serve as the BDC representative in cluster build meetings across multiple simultaneous programs (GPU compute clusters, AI infrastructure, and EDA environments), driving architecture alignment independently.
  • Analyze and validate cluster design choices across storage-to-compute distance, cross-mount latency, rack layout, and multi-site topology — and surface risk and tradeoff recommendations to leadership.
  • Lead performance benchmarking and profiling of HPC cluster infrastructure, running sanity benchmarks, regression suites, and cluster health checks to identify bottlenecks early.
  • Drive infrastructure optimization at multiple layers: scheduler-level tuning (LSF/Slurm),hardware-level tuning (GPU, CPU, networking, storage), and OS/kernel-level tuning (NUMA binding, socket binding, huge page configuration, kernel image selection).
  • Collaborate with platform and operations teams on cluster health, capacity planning, and operational readiness for tapeout milestones.
  • Partner with vendors and internal teams to evaluate new hardware, storage systems, and networking fabrics; produce architecture recommendations backed by data.
  • Continuously improve infrastructure observability, benchmarking frameworks, and architecture documentation to raise the bar across the BDC HPC platform.
What we need to see
  • B.E./B.Tech or M.Tech/M.S. with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale.
  • Deep understanding of data center architecture fundamentals: compute (CPU/GPU servers), storage (parallel file systems, NVMe, tiered storage), and high-speed networking (InfiniBand,Ethernet, NVLink).
  • Proven ability to evaluate and challenge infrastructure design decisions — including rack layout, power/cooling constraints, storage-to-compute distance, and network fabric topology.
  • Experience with OS and kernel-level performance tuning: NUMA binding, socket affinity, huge page configuration, kernel image selection, and system parameter optimization.
  • Hands-on experience with large-scale Linux HPC cluster administration using workload managers such as LSF and/or Slurm, including job scheduling optimization and resource utilization analysis.
  • Strong Linux/Unix system administration skills and proficiency in scripting (Python, Bash, or Perl) for automation and analysis.
  • Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks.
  • Excellent problem-solving and communication skills — able to synthesize complex architectural tradeoffs and present clear recommendations to engineering leadership.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Corporation • India

On-site
INR 3,500,000 - 7,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA AI • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Senior HPC Engineer
Senior HPC Engineer

Binaire Private Limited • New Delhi

On-site
INR 1,500,000 - 2,500,000
Opportunity to influence hardware selection
Ownership of high-performance compute infrastructure
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
HPC Engineer
HPC Engineer

Yotta Data Services Private Limited • Mumbai

On-site
INR 400,000 - 700,000