Distributed Systems Engineer - AI Infra & GPU Clusters

krea.ai

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI technology firm in San Francisco is seeking systems-oriented candidates to enhance their infrastructure used for advanced AI research. Ideal candidates will have a strong background in Python and experience managing distributed systems, particularly within GPU environments. The position offers the chance to work on cutting-edge projects involving large-scale data processing and custom infrastructure development. A passion for building and optimizing systems is a must.

Qualifications

  • Strong aptitude for building and optimizing distributed systems.
  • Good mental model of system interactions under varying conditions.
  • Experience with large-scale data processing and ETL systems.

Responsibilities

  • Work with distributed training and inference systems on GPU clusters.
  • Optimize data pipelines for research and operational tasks.
  • Collaborate with researchers to enhance ML infrastructure.

Skills

Python
Kubernetes
PyTorch
SQL
DuckDB
NumPy
Pandas

Tools

K8s
InfiniBand

Job description

A leading AI technology firm in San Francisco is seeking systems-oriented candidates to enhance their infrastructure used for advanced AI research. Ideal candidates will have a strong background in Python and experience managing distributed systems, particularly within GPU environments. The position offers the chance to work on cutting-edge projects involving large-scale data processing and custom infrastructure development. A passion for building and optimizing systems is a must.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Staff Engineer, Mid-Training Infra for Large-Scale AI
Staff Engineer, Mid-Training Infra for Large-Scale AI

Reflection • San Francisco (CA)

On-site
Top-tier compensation
Comprehensive health, dental, and vision insurance
Fully paid parental leave
+2
Senior Cloud Infra Engineer — AI GPU, Kubernetes SF On-site
Senior Cloud Infra Engineer — AI GPU, Kubernetes SF On-site

The Recruiting Guy • New York (NY)

On-site
USD 175,000 - 250,000
Staff Distributed Systems Engineer — AI Infra, Open Source
Staff Distributed Systems Engineer — AI Infra, Open Source

TrustIn • San Francisco (CA)

On-site
USD 400,000 - 500,000
Software Engineer, AI Infrastructure & Distributed Systems
Software Engineer, AI Infrastructure & Distributed Systems

Nexthop Systems Inc • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior Distributed Systems Engineer for AI GPU Clusters
Senior Distributed Systems Engineer for AI GPU Clusters

NVIDIA Corporation • United States

On-site
USD 120,000 - 160,000
Senior AI Infra Engineer - Distributed GPU Systems (Equity)
Senior AI Infra Engineer - Distributed GPU Systems (Equity)

NVIDIA • California (MO)

On-site
USD 170,000 - 288,000
Equity
Benefits
Senior AI Infra Engineer — Scalable GPU Clusters
Senior AI Infra Engineer — Scalable GPU Clusters

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Lead AI Systems Engineer: GPU Clusters & AI Ops
Lead AI Systems Engineer: GPU Clusters & AI Ops

Semiconductor Engineering • San Jose (CA)

On-site
USD 140,000 - 220,000