Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems

San Jose (CA)

On-site

USD 137,000 - 254,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cadence Design Systems in San Jose seeks a Senior AI Systems Engineer to own the AI infrastructure lifecycle, from architecting GPU clusters to deploying advanced models.

The role combines hands-on GPU cluster management, LLM deployment, and production-grade agentic workflows, with emphasis on Docker, Kubernetes, PyTorch, TensorFlow, and CI/CD.

You will collaborate across on-prem and cloud (Azure OpenAI, GCP) and mentor engineers while ensuring security and performance.

Qualifications

  • 10+ years in a senior technical role, with at least 5 years focused on building and operating high-performance computing or AI infrastructure.
  • Expert-level knowledge of NVIDIA GPU architecture and technologies like CUDA and cuDNN.
  • Proven experience with public cloud AI services, specifically Azure OpenAI and Google Cloud Platform (GCP).
  • Extensive hands-on experience with Docker: image management, container orchestration, and troubleshooting.

Responsibilities

  • AI Infrastructure Architecture & Strategy: Lead the design and implementation of next-gen AI infrastructure for Agentic AI initiatives.
  • Cloud AI Service Integration: Manage access, usage, and billing for Azure OpenAI and GCP services.
  • Hands-on GPU Cluster Management: Configure, install, and optimize GPU server clusters, troubleshoot hardware/software, tune performance.
  • Full-Stack AI Tech Stack Development & Operations: Deploy and maintain PyTorch, TensorFlow, Docker, Kubernetes, and CI/CD pipelines.
  • Advanced LLM Deployment & Optimization: Deploy and optimize LLMs using vLLM, TGI, TensorRT-LLM for high throughput and low latency.
  • Agentic AI Workflow & Service Engineering: Build production-grade AI workflows integrating LLMs with tools, APIs, and databases; mentor engineers.
  • Automation & Monitoring: Create automation scripts (Python/Bash/Perl); monitor system health, GPU utilization, container performance.
  • AI Systems Support & Mentorship: Provide escalation engineering support and technical leadership on AI systems engineering and performance tuning.
  • Security and Compliance: Implement security best practices for AI systems and data.

Skills

NVIDIA GPU architecture
CUDA
cuDNN
Azure OpenAI
GCP
Docker
Kubernetes
Python
Bash
Linux system administration

Tools

LSF
Docker
Kubernetes
CI/CD

Job description

Cadence Design Systems in San Jose seeks a Senior AI Systems Engineer to own the AI infrastructure lifecycle, from architecting GPU clusters to deploying advanced models.

The role combines hands-on GPU cluster management, LLM deployment, and production-grade agentic workflows, with emphasis on Docker, Kubernetes, PyTorch, TensorFlow, and CI/CD.

You will collaborate across on-prem and cloud (Azure OpenAI, GCP) and mentor engineers while ensuring security and performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Infrastructure Architect – GPU Clusters
Senior AI Infrastructure Architect – GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
AI Infrastructure Lead
AI Infrastructure Lead

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior AI Factory Deployment Architect (Multi-GPU)
Senior AI Factory Deployment Architect (Multi-GPU)

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 236,000
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits