Senior Cloud Operations Engineer

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 288,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a Senior Operations Engineer to enhance the efficiency, reliability, and scalability of our global business operations. You will help automate and support critical workflows, define implementations for automation and support solutions, and ensure teams operate seamlessly at scale.

The role involves building CI/CD pipelines, monitoring, offboarding, security practices, and cross-team collaboration to drive operational maturity across NVIDIA's IT landscape.

Qualifications

  • 8+ years of hands-on experience building/supporting complex services and BS/MS in Computer Science (or equivalent experience).
  • Knowledge in Python for automation, data handling, and tool development.
  • Experience with monitoring tools such as Prometheus, Grafana, Datadog, CloudWatch, Splunk and reporting.
  • Familiarity with ITSM practices, including incident, problem, and modification processes.
  • Ability to perform secure and compliant offboarding and access-related tasks.
  • Strong understanding of IT operations and system workflows.
  • Knowledge in core Java - Collections API, Streams API, Concurrency, I/O.
  • Knowledge in RDBMS and NoSQL (Cassandra, DynamoDb, Redis) databases.
  • Excellent communication skills with the ability to collaborate across multiple teams.
  • Excellent documentation, problem‑solving, and communication skills for cross‑team alignment.

Responsibilities

  • Drive day-to-day interactions with NVIDIA-wide IT subsystems, ensuring smooth operational workflows across infrastructure and applications.
  • Craft and maintain GitLab CI/CD pipelines to automate build, test, and deployment workflows.
  • Monitor system health, build/maintain dashboards, create alerts, and produce operational reports.
  • Perform user offboarding, access reviews, and compliance-related tasks across multiple systems.
  • Drive interactions with various IT subsystems, ensuring API performance and integration stability meet defined SLAs and SLOs.
  • Coordinate changes and releases between engineering, operations, and security teams.
  • Enforce security guidelines, manage vulnerability remediation, and collaborate with security teams on audits and assessments.
  • Maintain documentation, SOPs, and process improvements to enhance operational maturity.

Skills

Python automation
IT operations
Communication
Documentation
Security/compliance
Java core

Education

BS/MS in Computer Science

Tools

Prometheus
Grafana
Datadog
CloudWatch
Splunk
GitLab CI/CD

Job description

At NVIDIA, we are seeking a highly skilled Senior Operations Engineer to join our world-class NGC Cloud team. In this role, you will help drive the efficiency, reliability, and scalability of the systems that power our global business operations. This is an exceptional opportunity to shape how we automate, streamline, and support critical operational workflows across the organization. You will define how we implement innovative automation and support solutions, enabling teams to operate seamlessly and deliver impact at global scale- all within an encouraging and inclusive environment.

What you'll be doing:
  • Driving day-to-day interactions with NVIDIA wide IT subsystems, ensuring smooth operational workflows across infrastructure and applications.

  • Crafting and maintaining GitLab CI/CD pipelines to automate build, test, and deployment workflows.

  • Monitoring system health, building/maintaining dashboards, creating alerts, and producing operational reports.

  • Performing user offboarding, access reviews, and compliance-related tasks across multiple systems.

  • Drive interactions with various IT subsystems, ensuring API performance and integration stability meet defined SLAs and SLOs.

  • Coordinating changes and releases between engineering, operations, and security teams.

  • Enforcing security guidelines, managing vulnerability remediation, and collaborating with security teams on audits and assessments.

  • Maintaining documentation, SOPs, and process improvements to enhance operational maturity.

What we need to see:
  • 8+ years of hands-on experience building/supporting complex services and BS/MS in Computer Science (or equivalent experience).

  • Knowledge in Python for automation, data handling, and tool development.

  • Experience with monitoring tools (such as Prometheus, Grafana, Datadog, CloudWatch, Splunk) and reporting.

  • Familiarity with ITSM practices, including incident, problem, and modification processes.

  • Ability to perform secure and compliant offboarding and access-related tasks.

  • Strong understanding of IT operations and system workflows.

  • Knowledge in core Java - Collections API, Streams API, Concurrency, I/O.

  • Knowledge in RDBMS and NoSQL (Cassandra, DynamoDb, Redis) databases.

  • Excellent communication skills with the ability to collaborate across multiple teams.

  • Excellent documentation, problem‑solving, and communication skills for cross‑team alignment.

Ways to stand out from the crowd:
  • Experience designing or implementing automation pipelines or internal operational tools.

  • Background in customer support, technical support, or customer‑facing engineering roles.

  • Prior work in a security-conscious or compliance‑heavy environment.

  • Ability to build end‑to‑end monitoring solutions, dashboards, and automated reporting.

  • Strong documentation habits and a continuous‑improvement approach.

Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 22, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Operations Engineer
Senior Cloud Operations Engineer

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Senior Cloud Operations Engineer
Senior Cloud Operations Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Comprehensive benefits package
Senior Manager, Business Operations
Senior Manager, Business Operations

NVIDIA • Santa Clara (CA)

On-site
USD 240,000 - 379,500
Equity
Benefits
Senior Solutions Architect, NVIDIA Cloud Partner Operations
Senior Solutions Architect, NVIDIA Cloud Partner Operations

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Comprehensive benefits
Principal Software Engineer - Cloud Services
Principal Software Engineer - Cloud Services

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Distinguished Engineer, GPU Fleet Operations Automation, Distinguished Engineer, GPU Fleet Oper[...]
Distinguished Engineer, GPU Fleet Operations Automation, Distinguished Engineer, GPU Fleet Oper[...]

NVIDIA • Town of Texas (WI)

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Senior Software Engineer, Core Infrastructure Services - DGX Cloud

Socket.dev • Santa Clara (CA)

On-site
USD 168,000 - 322,000
equity (outlined)
Senior Solutions Architect, NVIDIA Cloud Partner Operations
Senior Solutions Architect, NVIDIA Cloud Partner Operations

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior Software Engineer, DGX Cloud Orchestration
Senior Software Engineer, DGX Cloud Orchestration

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Systems Operations and Administrator
Systems Operations and Administrator

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 112,000 - 219,000