Senior AI Infrastructure Engineer – GPU Systems

Clockwork Systems, Inc.

Palo Alto (CA)

On-site

USD 130,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Great benefits package
Catered lunch

Job summary

Clockwork Systems, Inc. is seeking an experienced Systems Engineer to design and implement high-performance distributed GPU training systems. This role involves working with GPU clusters, high-speed networking, and ensuring fault tolerance in complex systems. Ideal candidates will have 8+ years in systems software, a strong background in C/C++, and familiarity with distributed systems. Attractive benefits include competitive compensation and catered lunch, all in a diverse, inclusive environment.

Qualifications

  • 8+ years building systems software.
  • Experience in designing and building complex systems.
  • Familiarity with distributed storage and databases.

Responsibilities

  • Design and implement low-level systems software for GPU clusters.
  • Work with internals of frameworks like PyTorch and CUDA.
  • Debug complex distributed systems to ensure reliability.

Skills

Strong C/C++ in systems contexts
Deep understanding of concurrency
Experience reasoning about distributed system behavior
Comfortable reading and modifying large code bases

Job description

Clockwork Systems, Inc. is seeking an experienced Systems Engineer to design and implement high-performance distributed GPU training systems. This role involves working with GPU clusters, high-speed networking, and ensuring fault tolerance in complex systems. Ideal candidates will have 8+ years in systems software, a strong background in C/C++, and familiarity with distributed systems. Attractive benefits include competitive compensation and catered lunch, all in a diverse, inclusive environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior System Software Engineer - GPU AI Inference Equity
Senior System Software Engineer - GPU AI Inference Equity

NVIDIA • United States

On-site
USD 152,000 - 242,000
Senior Distributed Systems Engineer for AI GPU Clusters
Senior Distributed Systems Engineer for AI GPU Clusters

NVIDIA Corporation • United States

On-site
USD 120,000 - 160,000
Senior GPU Performance Engineer for AI Training
Senior GPU Performance Engineer for AI Training

CareerArc • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits
Senior AI GPU Infra Engineer — Performance & Scale
Senior AI GPU Infra Engineer — Performance & Scale

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Distributed Systems Engineer - AI Infra & GPU Clusters
Distributed Systems Engineer - AI Infra & GPU Clusters

krea.ai • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Infra Platform Engineer - GPU Scale (Equity)
Senior AI Infra Platform Engineer - GPU Scale (Equity)

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 300,000 - 350,000
Equity
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Senior Systems Performance Architect – GPU/CPU EDA
Senior Systems Performance Architect – GPU/CPU EDA

NVIDIA • Durham (NC)

On-site
USD 184,000 - 288,000
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus