Senior AI Tools Engineer, SRE Operations - GeForce NOW

Nvidia Corporation

Santa Clara (CA)

On-site

USD 144,000 - 230,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. You will build and deploy AI-powered tools to optimize a global GeForce Now service, transforming large-scale data streams into actionable intelligence for root-cause analysis and future trends.

You will lead the development of LLM- and agent-based systems, maintain data workflows, and influence platform choices to ensure the product's long-term success.

Qualifications

  • 5+ years of experience in SRE, AI tooling, or equivalent.
  • Strong proficiency in Python; Go or systems languages a plus.
  • Experience building, optimizing, and deploying AI tools.
  • Solid knowledge of LLM platforms and AI model development.
  • Hands-on with Kubernetes and AWS in production environments.
  • Active engagement with AI developments to inform decisions.
  • Expertise in automation and large-scale data pipelines.
  • Experience with monitoring/visualization tools like Grafana.

Responsibilities

  • Build and deploy AI-powered tools to support production data pipelines and GeForce Now operations.
  • Lead development of LLM- and agent-based systems to boost efficiency.
  • Maintain data workflows for large-scale data sources used in modeling.
  • Enhance LLM-based pipelines and align with product development timelines.
  • Advise on AI frameworks and platform choices for long-term sustainability.

Skills

Python
ML/AI Tools
Data Pipelines
Automation
Monitoring

Education

B.S. in Computer Science, Statistics, or Engineering

Tools

Kubernetes
AWS
Go
Docker

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people.

Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent.

As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

We are seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. We encourage applicants with SRE or equivalent experience.

What you will be doing:

You will build and deploy sophisticated AI‑powered tools and products. These tools support the operation and optimization of a critical production global GeForce Now service. This role is critical for transforming extensive production data streams—such as signals, metrics, and logs—into actionable intelligence.

The intelligence automates root cause analysis for incidents and predicts future service trends and patterns.

  • Build and implement robust AI/ML tools capable of analyzing production data to identify root causes for complex incidents and identify future operational trends.
  • Lead the development of brand‑new LLM‑ and Agent‑based systems to improve operational efficiency.
  • Establish and maintain excellent data management practices, including constructing workflows to convert and handle large‑scale data sources vital for model development.
  • Take charge of and enhance LLM‑based pipelines while integrating a strong grasp of LLM progress into product development.
  • Act as a resident authority on AI Frameworks, recommending the best platforms, toolsets, and architectural approaches to ensure the long‑term technical sustainability of the product.
What we need to see:
  • B.S. in Computer Science, Statistics, or Engineering (or equivalent experience), and 5+ years of experience.
  • Strong proficiency in Python; familiarity with Go or other systems languages is a plus.
  • Practical experience building, optimizing, and deploying AI tools.
  • Strong knowledge of the AI space and current developments, including understanding how LLM‑based platforms are built, optimized, and which platforms work best.
  • Hands‑on experience with container orchestration (Kubernetes) and cloud environments (AWS cloud).
  • Active engagement with developments in the AI field and the ability to distinguish meaningful advances from noise when making technical decisions.
  • Expertise in automation and handling large‑scale data pipelines.
  • Experience applying monitoring and visualization tools, such as Grafana, to interact with data.
  • Excellent ability to handle data sources and pipelines to transform and manage data.
Ways to stand out in a crowd:
  • Current experience in LLM improvement pipelines as well as a strong grasp of recent developments in LLM training.
  • Understanding of SRE concepts and experience managing production environments, as well as experience with Kubernetes, AWS, and other cloud technologies.
  • Someone with excellent knowledge of LLMs and AI Models who can reason and recommend an approach that sustains the team and product long‑term. (This person helps prevent grave mistakes by avoiding the wrong platform choice).
  • Proficiency in automation.

With a competitive salary package and benefits, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people in the world working for us. Are you a creative and autonomous AI Tools Engineer who loves challenges? Do you have a genuine passion for advancing the state of Site Reliability Engineering across a variety of industries? If so, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 144,000 USD - 230,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 31, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Tools Engineer, SRE Operations - GeForce NOW
Senior AI Tools Engineer, SRE Operations - GeForce NOW

NVIDIA AI • Santa Clara (CA)

On-site
USD 144,000 - 230,000
Equity
Benefits
Senior AI Tools Engineer, SRE Operations - GeForce NOW
Senior AI Tools Engineer, SRE Operations - GeForce NOW

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 144,000 - 230,000
Equity
Benefits
Senior AI Tools Engineer, SRE Operations - GeForce NOW
Senior AI Tools Engineer, SRE Operations - GeForce NOW

NVIDIA • United States

On-site
USD 144,000 - 230,000
Equity
Benefits
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior Software Engineer, AI Developer Tools
Senior Software Engineer, AI Developer Tools

NVIDIA • Santa Clara (CA)

On-site
USD 210,000 - 320,000
Equity
Benefits
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Principal Site Reliability Engineer
Principal Site Reliability Engineer

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Staff Site Reliability Engineer - AI Platform Runtime
Staff Site Reliability Engineer - AI Platform Runtime

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior Machine Learning Graphics Engineer, AI for Experiences
Senior Machine Learning Graphics Engineer, AI for Experiences

NVIDIA • Durham (NC)

On-site
USD 224,000 - 431,000
Equity
Benefits package
Senior Software Engineer – Platform Engineering
Senior Software Engineer – Platform Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Healthcare