Senior DevOps Engineer

NVIDIA

Pune District

On-site

INR 1,500,000 - 2,100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking an Engineering Lead in Pune to oversee the development and maintenance of Kubernetes-based infrastructure. You will work closely with skilled engineers to enhance operational capabilities and automation in data centers.

The ideal candidate will have extensive experience in deploying and managing cloud infrastructure, strong programming skills, and a degree in a relevant field. Join us in revolutionizing AI-powered compute infrastructures.

Qualifications

  • 5+ years of proven experience in relevant fields.
  • Strong background in Kubernetes with on-premises setup.
  • Proven debugging and analytical skills.

Responsibilities

  • Architect scaling operations in data centers.
  • Deploy and support end-to-end container management solutions.
  • Manage OSS CICD tools in a Kubernetes environment.

Skills

Kubernetes understanding
Python
Golang
Java
Ansible
Jenkins
SQL
NoSQL

Education

Bachelor’s or Master’s Degree in CS or related field

Tools

Docker
Elastic Search
MongoDB
Grafana
Kibana
Zabbix
Nagios

Job description

NVIDIA is looking for an outstanding engineering lead to join its Software Infrastructure and Operations team. The position will be part of a fast‑paced crew that develops and maintains sophisticated Kubernetes‑based development, compute and test environments for a multitude of platforms including Windows and Linux using OSS CICD tools GitHub, GitLab, Jenkins. You will be working with a team of passionate and skilled engineers that are continuously working to provide better tools to build and manage this infrastructure. With your help we would forge the next generation of compute infrastructure multiplying the power of the CPU, GPU and DPU for the age of AI. We need a motivated, hardworking and focused individual who has a real passion for operational excellence, infrastructure services, and automation.

What you’ll be doing:
  • Architect the scaling operation in our data centers. Deploy and support end‑to‑end container management solutions with Kubernetes, Docker, containerd. Design solutions with service discovery, networking, monitoring, logging, scheduling in Kubernetes.
  • Manage end‑to‑end OSS CICD tools GitLab/GitHub/Jenkins in on‑prem Kubernetes environment. Design and develop tools needed for automating CICD and developers’ workflow.
  • Design and build sophisticated automations and AI‑powered applications.
  • Use your depth in algorithms and system software background.
  • Work in teams to deploy new data center infrastructure.
  • Plan and implement critical metrics tracking using various data analytics mining methods and dashboards.
  • Reuse AI techniques to extract useful signals about machines and jobs from the data generated.
  • Take part in prototyping, crafting and developing cloud infrastructure for Nvidia.
What we need to see:
  • Strong Kubernetes understanding and background especially on‑premises setup and extensive experience with Kubernetes components & subsystems.
  • Experience of maintaining large‑scale on‑prem infrastructure applications & OSS CICD tools using Kubernetes.
  • Proven programming background in Python/Golang/Java and/or relevant scripting languages.
  • Excellent debugging and analytical skills and experience in databases both SQL (MySQL) and NoSQL (Elastic Search /MongoDB).
  • Proficient with configuration management tools like Ansible, Chef, Puppet and strong experience with Jenkins and/or other CI systems.
  • Hands‑on experience with VMs, Dockers, Kubernetes cluster.
  • Experience with analytics/visualization tools like Kibana, Grafana, Splunk etc. and experience with monitoring systems such as Zabbix and/or Nagios is nice to have.
  • 5+ years of proven experience.
  • Bachelor’s or Master’s Degree or equivalent experience in CS, Software Engineering, or related field.
Ways to stand out from the crowd:
  • Previous experience with DevOps/SRE teams.
  • Thrives in a multi‑tasking environment with constantly evolving priorities and documents work well.
  • Outstanding collaboration skills across organizational boundaries, experience with using and improving data centers and with computer algorithms and ability to choose the best possible algorithms to meet the scaling challenge.
  • Ability to divide complex problems into simple sub‑problems and then reuse available solutions to implement most of those.
  • Experience with designing simple systems that can work reliably without needing much support.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior CUDA Driver and DevOps Engineer
Senior CUDA Driver and DevOps Engineer

NVIDIA • India

On-site
INR 1,500,000 - 2,000,000
Competitive salaries
Comprehensive benefits package
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 300,000 - 550,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA • Pune District

On-site
INR 4,200,000 - 7,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior Cloud Software Engineer
Senior Cloud Software Engineer

NVIDIA • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NVIDIA • New Delhi

On-site
INR 4,000,000 - 8,000,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 900,000 - 1,500,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000