Inference Research Engineer (MTS)

Coral Bricks AI

San Francisco (CA)

On-site

USD 120,000 - 200,000

Full time

17 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health, dental and vision insurance
Flexible time off
Equity

Job summary

Coral Bricks AI is building an inference platform to make frontier models faster and cheaper for researchers and developers. You’ll work at the intersection of research and systems engineering, forming hypotheses, designing experiments, and implementing changes that improve time‑to‑first‑token and latency under real workloads.

You’ll help portfolio of models across NVIDIA/AMD GPUs, prototype cross‑platform changes, and push for practical, production‑ready improvements that customers directly

Qualifications

  • Strong software engineering fundamentals across systems, performance, or distributed computing.
  • Interest in LLM inference and running models locally or reviewing serving papers.
  • Curiosity to profile bottlenecks and understand what happens under the hood.
  • Comfort working in early-stage startup environments with fast pace.

Responsibilities

  • Research and build faster, cheaper LLM inference techniques.
  • Work between research and systems engineering to prototype changes.
  • Measure performance under production‑shaped traffic and refine ideas.
  • Collaborate on open‑weight model integration and porting across platforms.
  • Contribute to profiling, load testing, and system instrumentation.

Skills

Systems engineering
Performance optimization
Distributed computing
LLM inference
Curiosity
Shipping bias

Tools

vLLM
SGLang
llama.cpp

Job description

Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.

We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.

The team is small, technical, and shipping.

About Coral Bricks

Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.

We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.

The team is small, technical, and shipping.

The role

You’ll research and build new ways to make LLM inference faster and cheaper, then prove them against real agent workloads. The work sits between research and systems engineering: form a hypothesis about where time or memory is going, design the experiment, implement the change, and measure whether it survives contact with production‑shaped traffic.

This role is distinct from our GPU infrastructure role. You won’t own the day‑to‑day operation of the fleet or deployment platform. You’ll own the performance ideas that change what the serving system can do: new scheduling policies, cache strategies, parallelism approaches, quantization methods, and model‑specific optimizations.

This is a founding‑team role open to all experience levels, including new grads. You’ll work close to production, see your changes show up directly in customer cost and latency, and learn the deep end of the stack on the job.

What you’ll work on
  • Research ways to push throughput and bring down time‑to‑first‑token and inter‑token latency on workloads that are heavy on prompts, long‑running, and high‑fanout — batching, scheduling, parallelism, caching, and quantization.
  • Cross‑platform serving: we’re building for both NVIDIA and AMD GPUs — bring models up on each, close the performance gap between them, and keep the stack portable rather than vendor‑locked.
  • Prototype changes across the GPU and CPU sides of the serving stack when the research requires it, and carry successful ideas far enough to demonstrate them under production‑shaped load.
  • Bring up and tune new open‑weight model families as they release.
  • Build the profiling and load‑testing harness that tells us — honestly — what the system is doing under real agent load, not synthetic benchmarks.
  • Decide what we fork, what we contribute back upstream, and what we build ourselves.
You probably have
  • Strong software engineering fundamentals — systems, performance, or distributed computing, whether from work, internships, research, or personal projects. New grads are welcome.
  • A genuine interest in LLM inference. Maybe you’ve run models locally, poked at vLLM, SGLang, or llama.cpp, or just read the serving papers because you couldn’t help it.
  • Curiosity about what’s actually happening under the hood — you’d rather profile and find the real bottleneck than guess.
  • A high work ethic and excitement about early‑stage startups. The pace is fast, the problems are open‑ended, and everyone does a bit of everything.
  • A bias toward shipping. Most wins here come from a careful change that lands this week, not a six‑month rewrite.
Bonus
  • Open‑source contributions to vLLM, SGLang, TensorRT‑LLM, llama.cpp, or similar.
  • Experience with GPU programming on either vendor — CUDA, Triton, or ROCm / HIP — profiling tools (Nsight, rocprof, PyTorch Profiler), quantization, or MoE serving.
  • Experience designing and evaluating systems research: careful baselines, useful instrumentation, reproducible experiments, and skepticism about surprising results.
Compensation

$120,000–$200,000 base salary, plus 0.25%–2.0% equity. Both bands are wide on purpose: this role is open from new grad through senior, and we’d rather post the real span than a number that only fits one end of it. Where you land depends on experience, and cash and equity move together — take less of one and we’ll weight the other.

Equity vests over four years with a one‑year cliff. Health, dental, and vision coverage, and flexible time off.

Founding engineers shape the platform, the technical direction, and the team we build around it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Systems Engineer, GPU Infrastructure
Distributed Systems Engineer, GPU Infrastructure

Coral Bricks AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 200,000
Health, dental, and vision coverage
Flexible time off
Founding-team role
AI Infrastructure Engineer
AI Infrastructure Engineer

Coral Bricks AI • San Francisco (CA), Northern (KY)

Hybrid
USD 100,000 - 150,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Realmlabs • Sunnyvale (CA)

On-site
USD 210,000 - 350,000
Market aligned compensation
Founding engineer equity
Medical, Dental, Vision, and Life insurance
+2
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Remote option
Visa sponsorship
Relocation support
+2
Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Researcher
Machine Learning Researcher

SOLANA FOUNDATION • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff - Inference Research
Member of Technical Staff - Inference Research

Mixpeek • New York (NY)

On-site
USD 180,000 - 260,000
Member of Technical Staff - Research, Inference
Member of Technical Staff - Research, Inference

modal • New York (NY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Research, Inference
Member of Technical Staff - Research, Inference

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect AI • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Equity incentives
Visa sponsorship
Relocation support
+2