Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Coral Bricks AI is seeking an experienced engineer to own the operational systems behind the inference platform — clusters, GPU fleet, deployment machinery, and the control plane ensuring models stay available and traffic flows smoothly.
You’ll integrate with cloud infrastructure, distributed systems, networking, and storage, shipping scalable, observable solutions. This founding-team role offers broad ownership and a fast-paced startup environment.
Engineering San Francisco or remote - Full-time
Own the clusters, GPU fleet, and production systems that turn inference research into a reliable service
Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here - but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.
We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them - rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts - many times the tokens per second at a fraction of the cost.
The team is small, technical, and shipping. We also build in the open: a lot of the day-to-day happens in our Discord, where the developers building on Coral Bricks tell us what broke, compare numbers with us, and push on what we work on next.
You'll own the operational systems behind our inference platform: the clusters, GPU fleet, deployment machinery, and control plane that keep models available and traffic moving. When research produces a faster serving technique or a new model drops, you'll turn it into a repeatable, observable, production launch.
This role is distinct from our inference research role. You won't be measured on inventing a new attention kernel. You'll be measured on whether we can provision capacity, place workloads, ship changes, recover from failures, and operate a growing fleet without heroics.
This is a founding-team role with broad ownership. You'll work across cloud infrastructure, distributed systems, networking, storage, deployment, and the serving layer where they meet.
$120,000-$200,000 base salary, plus 0.25%-$2.0% equity. Where you land depends on experience, and cash and equity move together - take less of one and we'll weight the other.
Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.
Founding engineers shape the platform, the technical direction, and the team we build around it.