AI Infrastructure Engineer

Coral Bricks AI

San Francisco, Northern (CA, KY)

Hybrid

USD 100,000 - 150,000

Full time

11 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Coral Bricks AI in San Francisco or remote is seeking an early-career engineer to build and operate the cloud infrastructure behind our AI platform. You'll automate deployments, develop dashboards, and own concrete projects while learning the deeper parts of the stack.

You'll work with experienced systems engineers and founders; opportunities to grow toward GPU fleet operations or inference performance work as the company scales. This is a hands-on role in a fast-paced startup.

Qualifications

  • Around one to two years of professional software engineering, infrastructure, DevOps, or ML systems experience.
  • Solid programming fundamentals in at least one modern language; write maintainable code.
  • Experience with Linux, Git, containers, and cloud platforms.
  • Familiarity with CI/CD, IaC, monitoring or production support.
  • Strong debugging and problem-solving, capable of ownership.
  • Interest in early-stage startups and automation.

Responsibilities

  • Automate recurring infrastructure and internal workflow tasks.
  • Improve build, test, deploy, and rollback processes.
  • Operate AWS infrastructure: compute, containers, networking, storage, IAM, secrets.
  • Build observability: metrics, logs, dashboards, alerts, runbooks.
  • Benchmark and load-test serving configurations.
  • Support model and service launches with validated configurations.
  • Investigate production issues across systems and implement fixes.
  • Use AI coding agents carefully while learning the stack.
  • Keep infrastructure clean with clear code and documentation.

Skills

Python
TypeScript
Go
Debugging
Automation
Startup mindset

Tools

AWS
Docker
Kubernetes
Terraform
Pulumi
CI/CD pipelines

Job description

Engineering San Francisco or remote · Full-time

Automate the cloud infrastructure and operational workflows behind a fast-moving AI platform, and grow into the part of the stack that suits you

About Coral Bricks

Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.

We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.

The team is small, technical, and shipping. We also build in the open: a lot of the day-to-day happens in our Discord, where the developers building on Coral Bricks tell us what broke, compare numbers with us, and push on what we work on next.

The role

You'll help build and operate the cloud infrastructure around our AI platform. One day that might mean automating a deployment that still has manual steps; the next, adding the dashboard that makes a production issue obvious or writing a tool that turns a recurring operational task into a button or a command.

This is an early-career role designed for someone with around one to two years of professional experience. We don't expect you to arrive as an expert in GPU clusters, distributed systems, or LLM serving internals. We do expect you to be a solid programmer, comfortable in a terminal, eager to understand how production systems behave, and ready to take ownership of concrete projects while learning the deeper parts of the stack.

You'll work closely with experienced systems engineers and the founders. As you grow, so will the scope you own — and which direction it grows in is open. Some of this work leads toward the GPU fleet and the systems that operate it; some of it leads toward the inference performance work that makes the fleet fast.

What you'll work on
  • Automate recurring infrastructure, operational, and internal workflow tasks with scripts, services, and internal tools.
  • Improve how we build, test, deploy, and roll back services across development and production environments.
  • Help operate our AWS infrastructure, including compute, containers, networking, storage, permissions, and secrets.
  • Build useful observability: metrics, logs, dashboards, alerts, and runbooks that help the team find problems quickly.
  • Benchmark and load-test serving configurations so we argue about real numbers instead of guesses.
  • Support model and service launches by validating configurations, monitoring rollouts, and improving the launch process after each release.
  • Investigate production issues across application and infrastructure boundaries — often starting from a developer's report in our Discord rather than an alert — then turn what you learn into a lasting fix or better automation.
  • Use AI coding agents to move quickly, while reviewing their work carefully and understanding the systems you change.
  • Keep infrastructure understandable: clear code, small changes, useful documentation, and fewer one-off manual procedures.
You probably have
  • Around one to two years of professional software engineering, infrastructure, DevOps, site reliability, or ML systems experience. Strong internships, open-source work, or substantial personal projects can count too.
  • Solid programming fundamentals in Python, TypeScript, Go, or a similar language. You can write maintainable code, not just one-off shell commands.
  • Hands-on experience with Linux, Git, containers, and at least one cloud platform, whether from work or projects.
  • Experience with some part of the software delivery loop: CI/CD, deployment automation, infrastructure as code, monitoring, or production support.
  • A methodical approach to debugging. You follow the evidence, ask good questions, and keep going when the first explanation is wrong.
  • A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
  • A bias toward automation and shipping. When you do something twice, you start thinking about how the system should do it for you.
Bonus
  • Experience with AWS services such as ECS, EC2, ECR, IAM, CloudWatch, S3, or Amplify.
  • Familiarity with Terraform, Pulumi, CloudFormation, or another infrastructure-as-code tool.
  • Experience operating a production service, participating in incident response, or building tools used by other engineers.
  • Curiosity about LLM serving, GPUs, vLLM, SGLang, Kubernetes, or distributed systems. Prior production experience with them is not required.
  • Any experience profiling or benchmarking something and making it measurably faster.
  • A project where you used an AI coding agent to build something ambitious, automate a workflow, or explore an unfamiliar system.
Compensation

$100,000–$150,000 base salary, plus 0.1%–0.75% equity. Where you land depends on experience, and cash and equity move together — take less of one and we'll weight the other.

Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.

Early engineers shape the platform, the technical direction, and the team we build around it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Systems Engineer, GPU Infrastructure
Distributed Systems Engineer, GPU Infrastructure

Coral Bricks AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 200,000
Health, dental, and vision coverage
Flexible time off
Founding-team role
Founding Developer Advocate
Founding Developer Advocate

Coral Bricks AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Health, dental, and vision coverage
Flexible time off
Equity 0.25%–1.5%
Founding Applied AI Engineer
Founding Applied AI Engineer

Weekday (YC W21) • San Francisco (CA)

Hybrid
USD 220,000 - 300,000
Medical insurance
Dental insurance
Vision insurance
+5
AI Engineering Tech Lead / Architect
AI Engineering Tech Lead / Architect

Faros AI, Inc. • San Mateo (CA)

On-site
USD 210,000 - 250,000
AIML Engineer
AIML Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Performance bonus (up to 20% of base)
Equity participation
Health, dental, and vision insurance
+3
Software Engineer AI/ML Systems - USA
Software Engineer AI/ML Systems - USA

Socket.dev • Santa Clara (CA)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Entrepreneurial culture
+10
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Equity
RSUs
Paid time off
+13
Software Engineer (Entry-Level)
Software Engineer (Entry-Level)

Patapsco.AI • Elkridge (MD)

On-site
USD 70,000 - 110,000
Equity compensation
Paid time off
Paid company holidays
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range of $150-300k
Flexible work arrangement (remote or San Francisco office)
Full visa sponsorship and relocation support
+3
Member of Technical Staff - Training Platform
Member of Technical Staff - Training Platform

Kubelt • San Francisco (CA)

On-site
USD 150,000 - 300,000
Cash compensation $150K–$300K
Flexible work arrangement
Full visa sponsorship
+1