Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA is seeking an experienced professional to design, deploy, and manage Kubernetes solutions for large-scale data platforms. The role requires solid analytical troubleshooting skills and over 12 years of related experience. You will work on telemetry and observability for production systems to ensure reliability and performance.
The successful candidate will have a strong background in infrastructure automation and software design fundamentals. This opportunity offers a competitive salary ranging from $208,000 to $414,000 based on experience.
Systems Engineering is an engineering discipline focused on building, automating, and operating the platforms and tooling that deliver large-scale production systems with high efficiency, reliability, and velocity. It combines software and systems engineering practices across infrastructure automation, containerized platforms, storage, telemetry, and observability. Systems engineers are highly specialized and possess expertise across domains such as Kubernetes and container orchestration, infrastructure-as-code, CI/CD, storage systems, monitoring, and analytical troubleshooting. Their responsibilities center on deploying and operating reliable, automated platforms and on building the tools and services that keep storage and data infrastructure healthy and performant.
Our team at NVIDIA ensures that our internal and external facing GPU cloud services are deployed reliably, observable end-to-end, and continuously improved through automation. We enable developers to ship changes safely through repeatable CI/CD pipelines and Kubernetes-based deployments while keeping an eye on capacity, latency, and performance. A core part of this work is an SRE mindset: eliminating manual toil through automation, building self-service tooling, and growing the efficiency of production systems. We use a breadth of tools and approaches to tackle a broad spectrum of problems, and practices such as blameless postmortems, proactive identification of failure modes, and iterative improvement are key to product quality and to an interesting, dynamic day-to-day. Our culture of diversity, intellectual curiosity, problem-solving, and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame-free environment. We promote self-direction to work on meaningful projects while striving to build an environment that provides the support and mentorship needed to learn and grow.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative and autonomous, we want to hear from you!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 208,000 USD - 333,500 USD for Level 5, and 256,000 USD - 414,000 USD for Level 6.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until June 12, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.