Engineer, Supercomputing & Distributed Systems

Krea

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Krea is building next-generation AI creative tools and operating the infrastructure for research and inference. The team handles distributed training, 1000+ GPU clusters, petabyte-scale data pipelines, and custom datastores.

The role focuses on designing and maintaining scalable AI infra and data processing systems in a fast-paced environment. You will work with Python, Kubernetes, Torch, and data tools like DuckDB and Arrow to optimize performance, reliability, and scale for research workloads

Qualifications

  • Experience in designing and operating distributed systems.
  • Ability to work with Python, Kubernetes, and ML infra tooling.
  • Familiarity with data tooling like DuckDB, Arrow, SQL, Pandas, NumPy.

Responsibilities

  • Design and operate large-scale AI inference and training infrastructure.
  • Collaborate with research teams to optimize data pipelines and training throughput.
  • Troubleshoot performance and reliability issues across GPU clusters and datacenters.

Skills

Distributed systems intuition
Strong Python skills

Tools

Python
Kubernetes
PyTorch
DuckDB
Arrow
SQL
Pandas
NumPy

Job description

About Krea

At Krea, we are building next-generation AI creative tools.


We are dedicated to making AI intuitive and controllable for creatives. Our mission is to build tools that empower human creativity, not replace it.


We believe AI is a new medium that allows us to express ourselves through various formats—text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium.


Supercomputing / AI Infra at Krea

We build and operate the infrastructure for Krea's research and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We build a lot of this from scratch — custom distributed datastores, job orchestration systems, and streaming pipelines that replace tools like Kafka and Ray for modern AI workloads at scale.


Example projects

Distributed data systems


  • Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets

  • Run classification models on billions of images

  • Deploy and combine LLMs to caption massive multimedia data


GPU infrastructure


  • Manage distributed training and inference on 1000+ GPU Kubernetes clusters

  • Solve orchestration and scaling for large-scale GPU job processing

  • Scale workloads and research between clusters in multiple datacenters


Distributed training


  • Profile and optimize dataloaders streaming thousands of images per second

  • Profile and debug InfiniBand networking on huge training runs

  • Build fault tolerance systems for large-scale pretraining

  • Collaborate with researchers on evolving RL infrastructure


Applied ML pipelines


  • Find clean scenes in millions of videos using distributed shot-boundary detection

  • Customize and train models to filter billions of images for questions like "is this a screenshot?"

  • Build the systems that bridge raw cluster capacity and research output


Who we're looking for

Systems people. If you've read a blog post about InfiniBand debugging or building a custom distributed database and thought "I want to do that" — this is that team.


You’ll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc. It's OK if you don't have K8s or ML experience — the main thing we hire for is an intuition for distributed systems, and a great mental model of how systems interact and function under different conditions.


Strong candidates may have experience with


  • Python, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy…

  • Kubernetes

  • Designing and implementing large-scale ETL systems

  • Fundamental knowledge of containerization, operating systems, file-systems, and networking

  • Distributed systems design

  • Distributed training systems (NCCL, InfiniBand, RDMA)

  • Streaming and event processing systems (Kafka, Pulsar, or similar)

  • PyTorch internals, custom dataloaders, and training infrastructure

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Systems & AI Infrastructure Engineer
Distributed Systems & AI Infrastructure Engineer

Krea • San Francisco (CA)

On-site
USD 180,000 - 240,000
Engineer, Supercomputing & Distributed Systems
Engineer, Supercomputing & Distributed Systems

krea.ai • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000