Software Engineer, Production Engineering

NVIDIA Gruppe

Bengaluru

On-site

INR 3,000,000 - 5,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe in Bengaluru is seeking an experienced Senior Kubernetes Engineer to join an on-site Production Engineering team responsible for large-scale Kubernetes services and automation. You will help reduce manual tasks and ensure reliable SLAs across on-prem and cloud-like environments.

The role emphasizes deep systems knowledge, incident management leadership, and collaboration with cross-functional teams.

Qualifications

  • 7+ years of experience administering large-scale production Kubernetes systems in HA Internet/Cloud/Data Center environments.
  • BS in Computer Science, Engineering, Mathematics, or equivalent experience.
  • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
  • Familiarity with GPU / DPU hardware and HPC cluster environments.
  • Strong Linux system administration, DNS, DHCP and core networking (IP Tables, routing, firewalls).
  • Experience with CI/CD tools like Jenkins, ArgoCD.
  • Scripting in Python, Go, or Rust preferred.
  • Strong communication and ability to present to cross-functional groups.

Responsibilities

  • Be a member of the 24/7 Production engineering team to support Production Kubernetes Services, with automation focus and weekend split shifts.
  • Perform large-scale K8s and systems administration to maintain service SLAs, integrity and reliability.
  • Utilize alerts and observability tools to monitor, detect, prevent and respond to incidents.
  • Analyze logs, metrics and system behavior to troubleshoot, lead root cause analysis, and implement resolutions.
  • Initiate and lead incident management calls to ensure timely detection, escalation, and resolution with SMEs.

Skills

Kubernetes administration
Linux system administration
SLURM
CI/CD tooling (Jenkins, ArgoCD)
Scripting (Python, Go, or Rust)
Networking (DNS, DHCP, IP Tables)
Communication / stakeholder management

Education

BS in Computer Science or related field

Tools

Jenkins
ArgoCD

Job description

What you will be doing:
  • Being a member of the 24/7 Production engineering team to support Production Kubernetes Services, with a focus on automation and on reducing manual tasks. Flexibility to work on split-weekend shifts.
  • Perform large scale K8s administration, systems administration, and security monitoring tasks to maintain service SLAs,integrity and reliability.
  • Utilize alerts, alarms, and observability tools to proactively monitor, detect, prevent, and respond to incidents.
  • Apply deep systems knowledge to analyze logs, metrics, and system behavior to troubleshoot issues, lead root cause analysis, and implement effective resolutions.
  • Initiate and lead incident management calls, ensuring timely detection, escalation, and resolution of issues by engaging subject matter experts and service owners as needed to resolve complex incidents efficiently.
What we need to see:
  • 7+ years of demonstrated experience administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center environments. Strong preference for on-prem expertise.
  • BS in Computer Science, Engineering, Mathematics, or equivalent experience.
  • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
  • Familiarity with GPU / DPU hardware and high-performance computing Cluster environments.
  • Strong Linux system administration, DNS, DHCP and core linux networking (IP Tables, routing, firewalls) experience with skills to troubleshoot and maintain Services on large-scale bare-metal infrastructure.
  • Experience working with CI/CD tools like Jenkins, ArgoCD.
  • Experience in scripting, Programming in Python or Golang or Rust preferred, but not required.
  • Strong communication and soft skills, able to present to cross-functional group members in a persuasive manner.
Ways to stand out from the crowd:
  • Experience architecting, building, and deploying K8s for large-scale environments, used by thousands of people.
  • Passion for innovation and experience in groundbreaking high-performance Cluster technologies.
  • Ability to learn new technologies quickly.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Coredge.io • Bengaluru

On-site
INR 1,200,000 - 1,800,000
System Administrator with Kubernetes
System Administrator with Kubernetes

TekVizion • Hyderabad

On-site
INR 800,000 - 1,200,000
Competitive salary and benefits
Opportunity to work with cutting-edge technologies
Collaborative engineering culture
Senior DevOps Engineer
Senior DevOps Engineer

Coredge • Dadri

On-site
INR 1,000,000 - 1,500,000
Senior Software Engineer - Devops
Senior Software Engineer - Devops

DDN • Pune District

On-site
INR 3,000,000 - 4,500,000
Azure Kubernetes Engineer
Azure Kubernetes Engineer

Neurealm • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior DevOps Engineer
Senior DevOps Engineer

Benchmarkit • Pune District

On-site
INR 1,400,000 - 2,400,000
Senior Platform Engineer
Senior Platform Engineer

viavisolutions • Chennai District

On-site
INR 2,500,000 - 4,200,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Senior Platform Engineer
Senior Platform Engineer

InfoVision Inc. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Software Engineer - Site Reliability
Senior Software Engineer - Site Reliability

Freshworks • Chennai District

On-site
INR 800,000 - 1,200,000