Distributed Systems Engineer, GPU Infrastructure

Coral Bricks AI

San Francisco, Northern (CA, KY)

Hybrid

USD 120,000 - 200,000

Full time

5 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision coverage
Flexible time off
Founding-team role

Job summary

Coral Bricks AI is seeking an experienced engineer to own the operational systems behind the inference platform — clusters, GPU fleet, deployment machinery, and the control plane ensuring models stay available and traffic flows smoothly.

You’ll integrate with cloud infrastructure, distributed systems, networking, and storage, shipping scalable, observable solutions. This founding-team role offers broad ownership and a fast-paced startup environment.

Qualifications

  • Strong backend or distributed-systems fundamentals and experience owning production services end to end.
  • Experience with Linux, containers, networking, and at least one major cloud platform.
  • Good instincts around reliability: staged rollouts, observability, failure isolation, incident response, and simple systems that are easy to operate.
  • Comfort working from symptoms to root cause. A failed launch, an unhealthy node, or a latency spike is a systems problem to investigate, not a ticket to hand off.
  • A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
  • A bias toward shipping and automation. You fix the immediate problem, then build the mechanism that keeps it from becoming routine work.

Responsibilities

  • Own our GPU clusters and fleet across cloud providers: capacity, provisioning, machine images, drivers, networking, storage, health, and cost.
  • Build the control-plane systems that place workloads, manage capacity, drain and replace unhealthy nodes, and recover cleanly from failures.
  • Turn model launches into a reliable process: bring up new weights, validate serving configurations, roll out safely, watch production behavior, and roll back when needed. You'll hear how a launch landed from the developers in our Discord, not only from the dashboards.
  • Build deployment and release systems for inference servers and the services around them, with fast feedback and clear failure modes.
  • Create the observability we need to operate the fleet: metrics, logs, traces, dashboards, alerts, and tools that make incidents diagnosable instead of mysterious.
  • Improve reliability at every layer - autoscaling, load balancing, failover, backpressure, graceful degradation, and capacity planning.
  • Automate recurring operational work so the fleet can grow faster than the team operating it.

Skills

Distributed systems
Backend development
Linux
Containers
Cloud platforms
Observability
Reliability engineering

Tools

Kubernetes
Nomad
Slurm
ECS
NCCL
TensorRT-LLM

Job description

Engineering San Francisco or remote - Full-time

Own the clusters, GPU fleet, and production systems that turn inference research into a reliable service

About Coral Bricks

Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here - but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.

We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them - rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts - many times the tokens per second at a fraction of the cost.

The team is small, technical, and shipping. We also build in the open: a lot of the day-to-day happens in our Discord, where the developers building on Coral Bricks tell us what broke, compare numbers with us, and push on what we work on next.

The role

You'll own the operational systems behind our inference platform: the clusters, GPU fleet, deployment machinery, and control plane that keep models available and traffic moving. When research produces a faster serving technique or a new model drops, you'll turn it into a repeatable, observable, production launch.

This role is distinct from our inference research role. You won't be measured on inventing a new attention kernel. You'll be measured on whether we can provision capacity, place workloads, ship changes, recover from failures, and operate a growing fleet without heroics.

This is a founding-team role with broad ownership. You'll work across cloud infrastructure, distributed systems, networking, storage, deployment, and the serving layer where they meet.

What you’ll work on
  • Own our GPU clusters and fleet across cloud providers: capacity, provisioning, machine images, drivers, networking, storage, health, and cost.
  • Build the control-plane systems that place workloads, manage capacity, drain and replace unhealthy nodes, and recover cleanly from failures.
  • Turn model launches into a reliable process: bring up new weights, validate serving configurations, roll out safely, watch production behavior, and roll back when needed. You'll hear how a launch landed from the developers in our Discord, not only from the dashboards.
  • Build deployment and release systems for inference servers and the services around them, with fast feedback and clear failure modes.
  • Create the observability we need to operate the fleet: metrics, logs, traces, dashboards, alerts, and tools that make incidents diagnosable instead of mysterious.
  • Improve reliability at every layer - autoscaling, load balancing, failover, backpressure, graceful degradation, and capacity planning.
  • Automate recurring operational work so the fleet can grow faster than the team operating it.
You probably have
  • Strong backend or distributed-systems fundamentals and experience owning production services end to end.
  • Experience with Linux, containers, networking, and at least one major cloud platform. You can debug across application, host, and infrastructure boundaries.
  • Good instincts around reliability: staged rollouts, observability, failure isolation, incident response, and simple systems that are easy to operate.
  • Comfort working from symptoms to root cause. A failed launch, an unhealthy node, or a latency spike is a systems problem to investigate, not a ticket to hand off.
  • A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
  • A bias toward shipping and automation. You fix the immediate problem, then build the mechanism that keeps it from becoming routine work.
Bonus
  • Experience operating GPU or accelerator fleets, including NVIDIA or AMD drivers, topology, health checks, and failure modes.
  • Experience with Kubernetes, Nomad, Slurm, ECS, or another cluster scheduler - especially if you've had to work below its happy path.
  • Familiarity with vLLM, SGLang, TensorRT-LLM, PyTorch distributed, NCCL, or other model-serving and collective-communication systems.
  • Experience with multi-cloud capacity, bare-metal provisioning, or scarce-resource scheduling.
  • You've built an internal platform, scheduler, deployment system, or piece of infrastructure that other engineers trusted in production.
Compensation

$120,000-$200,000 base salary, plus 0.25%-$2.0% equity. Where you land depends on experience, and cash and equity move together - take less of one and we'll weight the other.

Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.

Founding engineers shape the platform, the technical direction, and the team we build around it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

Coral Bricks AI • San Francisco (CA), Northern (KY)

Hybrid
USD 100,000 - 150,000
Founding Developer Advocate
Founding Developer Advocate

Coral Bricks AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Health, dental, and vision coverage
Flexible time off
Equity 0.25%–1.5%
Member of Technical Staff, Sandbox Infrastructure
Member of Technical Staff, Sandbox Infrastructure

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Meaningful equity
Office in San Francisco
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Member of Technical Staff, Product Engineering
Member of Technical Staff, Product Engineering

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 225,000 - 275,000
Relocation assistance
Equity
Member of Technical Staff, Product Engineering
Member of Technical Staff, Product Engineering

SF Tensor • San Francisco (CA)

On-site
USD 225,000 - 275,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

descript • United States

On-site
USD 220,000 - 292,000
Equity
Competitive benefits
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2