Principal Software Architect- High Performance Computing

Applied Materials India

Chennai District

On-site

INR 4,000,000 - 6,000,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Applied Materials India seeks an experienced Software Architect to design and implement high-performance computing infrastructure across CPUs, GPUs, and FPGAs. You will analyze workloads, partition tasks to accelerators, and prototype with real code and data to validate designs.

You will optimize cost of ownership through profiling, capacity planning, and tuning, collaborating with Algo engineers, product managers, and stakeholders. Strong CUDA and distributed computing knowledge is essential.

Qualifications

  • 12 to 18 years of experience in robust, scalable infrastructure across CPUs/GPUs/FPGA.

Responsibilities

  • Design and implement robust, scalable HPC infrastructure combining CPUs, GPUs, and FPGAs.

Skills

CUDA parallel programming
Multi-threading
Distributed computing
Performance profiling
C/C++
GPU/CPU/FPGA heterogeneity
System architecture

Tools

NVIDIA Triton
MPI
UCX
OpenMP
OpenACC
OpenCL
vtune / Nsight

Job description

Our team is developing a high-performance computing solution for low-latency and high throughput image processing and deep-learning workload that enables our Chip Manufacturing process control equipment to offer differentiated value to our customers.

As an architect, you will get the opportunity to grow in the field of high-performance computing, GPU compute infra, complex system design and low-level optimizations for better cost of ownership.

Roles and Responsibility
  • As a Software Architect, you will be responsible for design and implementation of robust, scalable infrastructure solutions combining diverse processors (CPUs, GPUs, FPGAs).
  • You will analyze and partition workloads to the most appropriate compute unit, ensuring tasks like AI inference and parallel processing runs on specialized accelerators, while serial tasks run on CPUs.
  • You will work closely with cross-functional teams, including Algo engineers, product managers, and business stakeholders, to understand requirements and translate them into architectural/software designs that meet business needs.
  • You will be coding and developing quick prototypes to establish your design with real code and data.
  • You will be a subject Matter expert to unblock software engineers in the HPC domain.
  • You will be expected to profile entire cluster of nodes and each node with profilers to understand bottlenecks, optimize workflows and code and processes to improve cost of ownership.
  • Conduct performance tuning and capacity planning, monitoring GPU metrics (e.g., using NVIDIA DCGM) for reliability
  • Evaluate and recommend appropriate technologies and frameworks to meet project requirements.
  • Lead the design and implementation of complex software components and systems.
  • Ensure that software systems are scalable, reliable, and maintainable.
  • Your primary focus will be on ensuring that the software systems are scalable, reliable, maintainable and cost effective.
Our Ideal Candidate

Someone who is passionate about and has deep understanding and experience in design and development of cutting edge HPC systems and heterogenous computing infrastructure. He should have very good hands-on experience in parallel programming (CUDA) and AI inference infrastructure. He should be able to multi-task and switch contexts based on business needs.

Qualifications
  • 12 to 18 years of experience in implementing robust, scalable, and secure infrastructure solutions combining diverse processors (CPUs, GPUs, FPGAs)
  • Working experience of GPU inference server like Nvidia Triton.
  • Very good knowledge C/C++, Data structure and Algorithms and complexity analysis.
  • Experience in developing Distributed High Performance Computing software using Parallel programming frameworks like MPI, UCX etc.
  • Experience in GPU programming using CUDA, OpenMP, OpenACC, OpenCL etc.
  • In depth experience in Multi-threading, Thread Synchronization, Inter process communication, and distributed computing fundamentals.
  • Experience in Inter Process communication using Shared memory and Pipes.
  • Experience in performance profiling at application and system level (e.g. vtune, Oprofiler, perf, Nividia Nsight etc.)
  • Experience in low level code optimization techniques using Vectorization and Intrinsics, cache-aware programming, lock free data structures etc.
  • Familiarity with microservices architecture and containerization technologies (docker/singularity) and low latency Message queues.
  • Excellent problem-solving and analytical skills.
  • Strong communication and collaboration abilities.
  • Ability to mentor and coach junior team members.
Additional Qualifications
  • Experience in HPC Job-Scheduling and Cluster Management Software (SLURM, Torque, LSF etc.)
  • Good knowledge of Low-latency and high-throughput data transfer technologies (RDMA, RoCE, InfiniBand)
  • Good Knowledge of Parallel processing and DAG execution Frameworks like Intel TBB flowgraph, OpenCL/SYCL etc.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Engineer (HPC, GPU, CUDA)
Lead Engineer (HPC, GPU, CUDA)

AIRA Matrix • Thane

On-site
INR 1,200,000 - 2,500,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA Corporation • India

On-site
INR 3,500,000 - 7,000,000
GPU Cluster Architect
GPU Cluster Architect

Nebius B.V. • India

On-site
INR 4,000,000 - 6,500,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Corporation • Mumbai

On-site
INR 3,500,000 - 5,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA • Bengaluru

On-site
INR 2,500,000 - 4,000,000
GPU Architect
GPU Architect

NVIDIA Corporation • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • India

On-site
INR 3,000,000 - 6,000,000
Sr. HPC ENGINEER
Sr. HPC ENGINEER

Cognizant • Hyderabad

On-site
INR 800,000 - 1,200,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Principal Engineer – Scale-Up GPU Networking (HPC / AI)
Principal Engineer – Scale-Up GPU Networking (HPC / AI)

Hewlett Packard Enterprise Development LP • India

Hybrid
INR 5,000,000 - 8,500,000