Lead Engineer (HPC, GPU, CUDA)

AIRA Matrix

Thane

On-site

INR 1,800,000 - 3,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AIRA Matrix in India is seeking an AI/ML Architect to define architectural changes and accelerate deep learning models across GPU/CPU stacks, translating business needs into production-ready goals.

You will benchmark CV algorithms, optimize performance, and drive end-to-end workflows from data curation to deployment, collaborating with cross-functional teams to deliver scalable AI solutions.

Qualifications

  • Bachelor-level degree in CS/EE or related field.
  • Strong background in deployment of complex deep learning architectures.
  • 1+ years in ML/DNN with focus on DL fundamentals and model optimization.
  • Proficient in C++ with knowledge of data structures and algorithms.
  • Experience with DNN frameworks (Torch, Caffe, TensorFlow).
  • Proficient in Python and bash scripting.
  • Experience with Windows, Ubuntu and Centos.
  • Excellent communication and collaboration skills.

Responsibilities

  • Identify architectural changes or new approaches to accelerate deep learning models.
  • Translate business needs into product goals for AI-ML workloads and deployment.
  • Benchmark and optimize CV algorithms on heterogeneous hardware (GPU/CPU).
  • Collaborate to drive end-to-end workflow from data curation to deployment.

Skills

Deep learning architectures
DNN deployment
C++ programming
Python scripting
Performance optimization
DL fundamentals
Communication skills

Education

Bachelor's or higher in CS/EE

Tools

CMake
Make
Ninja
Clang-Tools
TensorRT
CuDNN
PyTorch
CUDA
OpenCL
OpenMP
MPI
Docker
Singularity
Kubernetes
SLURM
LSF

Job description

Responsibilities
  • We seek an expert to identify architectural changes and/or completely new approaches for accelerating our deep learning models.
  • As an architect you are responsible for converting business needs associated with AI-ML algorithms into a set of product goals covering workload scenarios, end user expectations, compute infrastructure and time of execution; this should lead to a plan for making the algorithms production ready.
  • Benchmark and optimize the Computer Vision Algorithms for performance and quality KPIs on the heterogeneous hardware stacks (GPU + CPU, etc.).
  • Collaborate with various teams to drive an end to end workflow from data curation and training to performance optimization and deployment.
Skills Required
  • Bachelors or Higher in Computer Science, Electrical Engineering, or related field.
  • A strong background in deployment of complex deep learning architectures.
  • 1+ years of relevant experience in at least a few of the following relevant areas: Machine learning (with focus on Deep Neural Networks), including understanding of DL fundamentals; Experience adapting and inferencing DNNs for various tasks; Experience developing code for one or more of the DNN training frameworks (such as Torch, Caffe or TensorFlow): Numerical analysis, Performance analysis, Model compression and Optimization & Computer architecture.
  • Strong data structures and algorithms knowhow with excellent modern C++ programming skills.
  • Good grasp over software engineering and tools like CMake, Make (or Ninja), Clang-Tools, etc.
  • Hands-on expertise with TensorRT, CuDNN, PyTorch.
  • Hands-on expertise with GPU computing (CUDA, OpenCL or OpenACC) and HPC (MPI, OpenMP).
  • Proficient in Python programming and bash scripting.
  • Proficient in Windows, Ubuntu and Centos operating systems.
  • Excellent communication and collaboration skills.
  • Self-motivated and able to find creative practical solutions to problems.
Good to Have
  • Hands-on experience with PTX-ISA for CUDA or vector intrinsics like AVX, SSE, etc.
  • In-depth understanding of container technologies like Docker, Singularity, Shifter, Charliecloud.
  • Hands-on experience with HPC cluster job schedulers such as Kubernetes, SLURM, LSF.
  • Familiarity with cloud computing architectures.
  • Hands-on experience with Software Defined Networking and HPC cluster networking.
  • Working knowledge of cluster configuration management tools such as Ansible, Puppet, Salt.
  • Understanding of fast, distributed storage systems and Linux file systems for HPC workloads.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer
AI Engineer

STCO India • Hyderabad

On-site
INR 800,000 - 1,200,000
HPC Research and Development Engineer
HPC Research and Development Engineer

Cephas Consultancy Services • Chennai District

On-site
INR 1,200,000 - 2,000,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Bengaluru

On-site
INR 4,200,000 - 6,000,000
GPU Infrastructure Engineer / HPC Engineer
GPU Infrastructure Engineer / HPC Engineer

Larsen & Toubro • Mumbai

On-site
INR 3,600,000 - 6,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior HPC Engineer
Senior HPC Engineer

Binaire Private Limited • New Delhi

On-site
INR 1,500,000 - 2,500,000
Opportunity to influence hardware selection
Ownership of high-performance compute infrastructure
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA • Bengaluru

On-site
INR 2,500,000 - 4,000,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000
Senior HPC Platform Architect
Senior HPC Platform Architect

NVIDIA AI • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000