Software Engineer, Production Engineering

NVIDIA

Bengaluru

On-site

INR 3,000,000 - 4,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a senior engineer to design and develop data center automation tools core to large-scale supercomputer and cloud deployments.

You will enable reliable, scalable automation for infrastructure-as-code workflows, including CI/CD pipelines, compute resource management, and devices like DPUs.

Qualifications

  • 5+ years of automation platform experience in data centers for cloud services.
  • BS/BTech Degree in CS or equivalent.
  • Experience building software tools for DC operations: provisioning, image deployment, patching, packaging.

Responsibilities

  • Craft DC automation frameworks, workflows, and data models for automating processes.
  • Be a member of the 24/7 Production engineering team to support Production Services and reduce manual tasks.
  • Craft workflows to automate DC operations: provisioning, patching, image build, telemetry, and issue resolution.
  • Integrate DC systems such as DCIM, alert management, telemetry, and control plane of infrastructure (storage, compute, networking).
  • Build data center automation for AI workloads and applications.

Skills

Python
Golang
Linux
Kubernetes
DCIM
CI/CD
Jenkins
ArgoCD
AWX
DPUs/SmartNICs
Ansible
Terraform

Education

BS/BTech in CS

Tools

Ansible
Terraform
Jenkins
ArgoCD
AWX

Job description

NVIDIA is looking for a senior engineer to design and develop data center infrastructure automation tools core to large-scale supercomputer and cloud deployments. To enable reliable, scalable and efficient automation to support core infrastructure-as-code workflows and tools, including CI/CD pipelines, compute resource management flow for environments that incorporate GPUs, (DPUs) Data Processing Units, as well as other devices and resources.

We are looking for an engineer who has a deep understanding in building data center automation tools in a distributed systems environment, outstanding design skills and a track record in building and delivering large-scale software infrastructure

What you will be doing:
  • Crafting and developing DC automation frameworks, workflows, and data models for automating processes.

  • Being a member of the 24/7 Production engineering team to support Production Services, with a focus on reducing manual tasks.

  • Crafting workflows to automate DC operation tasks such as provisioning, patching, image build, telemetry, and resolving production issues for Production Engineering.

  • Performing integrations between various DC systems such as DCIM, alert management and telemetry, and the control plane of infrastructure (storage, compute, and networking).

  • Building data center automation specifically adapted for AI workloads and applications.

What we need to see:
  • Minimum 5+ years of demonstrated experience in developing automation platforms for data centers to support highly available, large-scale, cloud service environments.

  • BS/BTech Degree in CS or an equivalent combination of education, technical training, and work experience.

  • Expertise in building software tools for automating data center operations - OS provisioning, Image deployment, patching, packaging, and change implementation activities.

  • Proficiency in at least one of the programming languages: Python or Golang.

  • Deep knowledge of DCIM, asset database build/implementation, and DC automation tools such as Ansible or terraform.

  • Experience working with CI/CD tools like Jenkins, ArgoCD, or AWX.

  • Experience deploying and maintaining embedded devices such as SmartNICs, DPUs, and other devices that run Linux on less-than-traditional servers.

  • In depth knowledge in Linux and strong experience on K8s Containers Technologies.

  • Strong communication and soft skills, able to present to cross-team members in a persuasive manner.

  • Enthusiasm for influencing and establishing collaborations with software and functional groups such as development, server, storage, and security teams in a matrix organization.

Ways to stand out from the crowd:
  • Experience architecting, building, and deploying automation frameworks for large-scale (1000s of machines) environments, used by thousands of people.

  • Passion for innovation and investing in groundbreaking technologies.

  • Ability to learn new technologies quickly.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers the next generation inventions from artificial intelligence to autonomous cars. NVIDIA is looking for great people like you to help us build and Operate Datacenter infrastructure at scale, which would further accelerate the next wave of artificial intelligence services.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Cloud Software Engineer
Senior Cloud Software Engineer

NVIDIA • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 1,500,000 - 2,100,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA • Pune District

On-site
INR 1,500,000 - 2,100,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 4,000,000 - 6,500,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA Corporation • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000
Senior CAD Engineer
Senior CAD Engineer

NVIDIA • Hyderabad

Hybrid
INR 1,800,000 - 3,200,000