Cosmos Infra & End-to-End Performance Engineer — Equity

Nvidia Corporation in

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation in California seeks a Senior Software Engineer for Cosmos Infrastructure and End to End Performance. You will help build world foundation models and optimize performance across data center and edge deployments, collaborating with customers to enable the ecosystem and automate data ingestion, curation, and training workflows.

Ideal candidates hold a Master’s degree in a STEM field with 5+ years of related experience in distributed accelerator systems, distributed PyTorch, and

Qualifications

  • A Masters in Computer Engineering, Computer Science, Electrical Engineering or related STEM degree or equivalent experience. 5 years of relevant work experience.
  • Expertise in large scale parallel and distributed accelerator-based system systems.
  • Expertise optimizing performance and AI workloads on large scale systems. Experience with performance modeling and benchmarking at scale.
  • Proficiency in Distributed PyTorch; Python, C/C++. A strong background in Computer Architecture, Networking, Storage systems, Accelerators.
  • Understanding of DNNs and their use in emerging AI/ML applications and services.
  • Expertise with at least one of public CSP infrastructure (GCP, AWS, Azure, OCI, ...).
  • A deep understanding of World Foundation Models and their application to Physical AI.
  • Experience developing infrastructure to automate multimodal data ingestion and curation. Experience driving tokenization and data set preparation is a plus.
  • Prior experience building transformer models (autoregressive and diffusion) or building the framework for and driving Post training - fine tuning and RL algorithms.
  • Understanding of how to optimize for inference - export, quantization and containerization.

Responsibilities

  • Building SoTA, World foundation models (Cosmos3 etc).
  • Engage with end-to-end performance analysis and drive HW-SW codesign for data center and Edge deployments.
  • Engage with customers to ensure the Cosmos models are easy to use and enabling the ecosystem.
  • Develop infrastructure to improve and automate the entire process of data ingestion, curation, pre-training, post-training, export/quantization and deployment on the edge.
  • Design for robustness and fault tolerance.

Skills

Distributed PyTorch
Python
C/C++
Large-scale systems
Performance optimization
DNNs / AI workloads
World Foundation Models
Cloud CSP (GCP/AWS/Azure/OCI)
Data ingestion automation
Transformer models / post-training
Inference optimization

Education

Master's degree in STEM

Tools

CUDA
TensorFlow
JAX
Megatron-LM

Job description

NVIDIA Corporation in California seeks a Senior Software Engineer for Cosmos Infrastructure and End to End Performance. You will help build world foundation models and optimize performance across data center and edge deployments, collaborating with customers to enable the ecosystem and automate data ingestion, curation, and training workflows.

Ideal candidates hold a Master’s degree in a STEM field with 5+ years of related experience in distributed accelerator systems, distributed PyTorch, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cosmos Infrastructure & E2E Performance Engineer
Senior Cosmos Infrastructure & E2E Performance Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Cosmos Infra & End-to-End Performance Engineer
Senior Cosmos Infra & End-to-End Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Cosmos Infra Engineer - End-to-End Performance
Senior Cosmos Infra Engineer - End-to-End Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Equity
Benefits
Senior Software Engineer, Cosmos Infrastructure and End to End Performance
Senior Software Engineer, Cosmos Infrastructure and End to End Performance

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer, Cosmos Infrastructure and End to End Performance
Senior Software Engineer, Cosmos Infrastructure and End to End Performance

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, Cosmos Infrastructure and End to End Performance
Senior Software Engineer, Cosmos Infrastructure and End to End Performance

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, Cosmos Infrastructure and End to End Performance
Senior Software Engineer, Cosmos Infrastructure and End to End Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Equity
Benefits
Principal Software Engineer, E2E Performance and Goodput - CSP Engagements
Principal Software Engineer, E2E Performance and Goodput - CSP Engagements

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Comprehensive benefits
Senior Deep Learning Engineer
Senior Deep Learning Engineer

NVIDIA • Indiana (PA)

On-site
USD 100,000 - 150,000
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Oregon (WI)

On-site
USD 224,000 - 432,000
Equity
Benefits