Senior AI Infra Engineer — Scale AI Workloads & Equity

Socket.dev

Santa Clara (UT)

On-site

USD 184,000 - 357,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Health benefits

Job summary

NVIDIA as a leader in AI infrastructure seeks an experienced senior AI Infrastructure Software Engineer to design, build, and operate platforms enabling large-scale AI training, inference, and GenAI workloads. You will improve resiliency, scalability, and observability across distributed systems.

The role emphasizes problem-solving, API co-design, and reliability metrics with mentorship in a collaborative, blameless culture. Base salary discusses by location, with equity and benefits.

Qualifications

  • Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.
  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).
  • Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.
  • Proven track record in building and scaling large-scale distributed systems.
  • Experience with AI training and inferencing and data infrastructure services.
  • Familiar in Kubernetes and operating large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki).
  • Proficiency in programming languages such as Golang, Python, C/C++, script languages
  • Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential.

Responsibilities

  • Develop platform and tools for large-scale AI, LLM, and GenAI infrastructure.
  • Develop and optimize tools to improve AI/ML workload efficiency and resiliency.
  • Root cause and analyze and triage failures from the application level to the hardware level
  • Enhance infrastructure and products underpinning NVIDIA's AI platforms.
  • Co-design and implement APIs for integration with NVIDIA's resiliency stacks on the platform.
  • Define meaningful and actionable reliability metrics to track and improve system and service reliability.
  • Skilled in problem-solving, root cause analysis, and optimization.

Skills

Golang
Python
C/C++
Scripting languages

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
ELK
Prometheus
Loki

Job description

NVIDIA as a leader in AI infrastructure seeks an experienced senior AI Infrastructure Software Engineer to design, build, and operate platforms enabling large-scale AI training, inference, and GenAI workloads. You will improve resiliency, scalability, and observability across distributed systems.

The role emphasizes problem-solving, API co-design, and reliability metrics with mentorship in a collaborative, blameless culture. Base salary discusses by location, with equity and benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer — Equity Eligible, Scale Telemetry
Senior AI Infra Engineer — Equity Eligible, Scale Telemetry

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Senior AI Infra Systems Engineer - Equity
Senior AI Infra Systems Engineer - Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Equity
Senior Architect, Scaled AI Inference & Systems — Equity
Senior Architect, Scaled AI Inference & Systems — Equity

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Senior AI Infra Engineer | Equity Eligible, Observability Focus
Senior AI Infra Engineer | Equity Eligible, Observability Focus

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Engineer — Build Telemetry & Ops, Equity
Senior AI Infra Engineer — Build Telemetry & Ops, Equity

NVIDIA • Austin (TX)

On-site
USD 184,000 - 357,000
Senior AI Infrastructure Engineer – Scale & Equity
Senior AI Infrastructure Engineer – Scale & Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior AI Infra Engineer — Scale Cloud AI Platforms
Senior AI Infra Engineer — Scale Cloud AI Platforms

NVIDIA • United States

On-site
USD 184,000 - 357,000
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Senior AI Infra Engineer — Telemetry & ML Ops Equity
Senior AI Infra Engineer — Telemetry & ML Ops Equity

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Senior GPU Performance Engineer, Scale & AI — Equity
Senior GPU Performance Engineer, Scale & AI — Equity

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity and benefits