Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA Gruppe

Gurugram District

On-site

INR 3,000,000 - 6,000,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA is seeking a Senior Solution Architect for Networking and compute infrastructure to join the Infrastructure Specialist Team. The role focuses on building AI/HPC infrastructure for customers, ensuring performance and reliability in large-scale data centres.

You will interact with customers, partners and internal teams to analyze, define and implement substantial networking projects. The ideal candidate will have 5+ years in networking and DC compute, strong Linux, automation tooling

Qualifications

  • Bachelor’s degree or higher in Computer Science, Electrical/Computer Engineering, Physics or related field with 5+ years in networking and data centre compute architecture.
  • Strong knowledge of HPC, AI & EVPN, BGP, OSPF, VXLAN.
  • Deep DC architecture understanding: compute, storage, InfiniBand, Ethernet, NVLink.
  • Experience with Linux systems and scripting (Python, Bash).

Responsibilities

  • Build AI/HPC infrastructure for new and existing customers.
  • Support operational aspects of large-scale AI clusters, focusing on performance, monitoring, logging, and alerting.
  • Engage in the full lifecycle from design to deployment and refinement.
  • Develop tooling to automate large-scale infrastructure and self-service resources.
  • Deploy monitoring solutions for servers, network and storage.
  • Troubleshoot from bare metal to application level.
  • Document standard methodologies for customers and internal teams; participate in POCs/POVs.

Skills

Networking fundamentals
TCP/IP stack
Linux administration
Automation tooling

Education

BS/MS/PhD or equivalent

Tools

Jenkins
Ansible
Puppet
Chef
Kubernetes

Job description

NVIDIA is looking for Senior Solution Architect, Networking and compute Infrastructure to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centres. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyse, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer!

What You'll Be Doing:
  • Primary responsibilities will include building AI/HPC infrastructure for new and existing customers.
  • Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting.
  • Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement.
  • Develop tooling to automate and manage of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources.
  • Deploy monitoring solutions for the servers, network and storage.
  • Perform troubleshooting bottom up from bare metal, operating system, software stack and application level.
  • Being a technical resource, develop, re-define and document standard methodologies to share with customer and internal teams Support activities and engage in POCs/POVs for future improvements.
  • Experience with system administration on Linux systems required (CentOS, RHEL, and Ubuntu preferred)
What We Need to See:
  • BS/MS/PhD or equivalent experience in Computer Science, Data Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering fields with at least 5+ years’ work or research experience in networking fundamentals, TCP/IP stack, and data centre compute architecture.
  • Advance knowledge of HPC, AI & EVPN, BGP, OSPF, VXLAN protocols.
  • Deep understanding of DC architecture fundamentals such as compute, storage (PFS) & InfiniBand, Ethernet, NVLink.
  • Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks.
  • Python programming, bash scripting experience and Advance Linux knowledge.
  • Extensive experience delivering automated network provisioning and comfortable with automation and configuration management tools including Jenkins, Ansible, Puppet. Chef, etc.
  • Possess solid working knowledge of Ethernet/InfiniBand/RDMA core principles.
  • Excellent customer-facing and communication skills (verbal and written in both languages), enabling effective engagement with customers, partners, and cross-functional teams across the India region. Listening skills in English are critical.
  • Willingness to travel
Ways To Stand Out from The Crowd:
  • Knowledge of CPU and/or GPU architecture including of Kubernetes, container related microservice technologies.
  • Background with RDMA (InfiniBand or RoCE) fabrics.
  • Linux or Networking Certifications (e.g., CCNP, CCIE) or NVIDIA-related certifications.
  • Deep Knowledge on observability stack and build experience.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA AI • Gurugram District

On-site
INR 2,500,000 - 6,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA • Gurugram District

On-site
INR 3,500,000 - 7,500,000
Senior Solutions Architect Networking And Compute Infrastructure
Senior Solutions Architect Networking And Compute Infrastructure

NVIDIA AI • Gurugram District

On-site
INR 4,000,000 - 7,000,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA Gruppe • Pune District

On-site
INR 1,200,000 - 2,000,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA Gruppe • Mumbai

On-site
INR 2,400,000 - 5,400,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA Corporation • India

On-site
INR 3,000,000 - 5,600,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA • Maharashtra

On-site
INR 1,500,000 - 1,900,000
Solution Architect, SAE
Solution Architect, SAE

NVIDIA Gruppe • Gurugram District

On-site
INR 3,500,000 - 7,000,000