Turn this role into an interview — a resume and cover letter built around what this employer wants.
Coral Bricks AI is building an inference platform to make frontier models faster and cheaper for researchers and developers. You’ll work at the intersection of research and systems engineering, forming hypotheses, designing experiments, and implementing changes that improve time‑to‑first‑token and latency under real workloads.
You’ll help portfolio of models across NVIDIA/AMD GPUs, prototype cross‑platform changes, and push for practical, production‑ready improvements that customers directly
Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.
We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.
The team is small, technical, and shipping.
Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.
We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.
The team is small, technical, and shipping.
You’ll research and build new ways to make LLM inference faster and cheaper, then prove them against real agent workloads. The work sits between research and systems engineering: form a hypothesis about where time or memory is going, design the experiment, implement the change, and measure whether it survives contact with production‑shaped traffic.
This role is distinct from our GPU infrastructure role. You won’t own the day‑to‑day operation of the fleet or deployment platform. You’ll own the performance ideas that change what the serving system can do: new scheduling policies, cache strategies, parallelism approaches, quantization methods, and model‑specific optimizations.
This is a founding‑team role open to all experience levels, including new grads. You’ll work close to production, see your changes show up directly in customer cost and latency, and learn the deep end of the stack on the job.
$120,000–$200,000 base salary, plus 0.25%–2.0% equity. Both bands are wide on purpose: this role is open from new grad through senior, and we’d rather post the real span than a number that only fits one end of it. Where you land depends on experience, and cash and equity move together — take less of one and we’ll weight the other.
Equity vests over four years with a one‑year cliff. Health, dental, and vision coverage, and flexible time off.
Founding engineers shape the platform, the technical direction, and the team we build around it.