Founding Inference Engineer

General Compute

San Francisco (CA)

On-site

USD 180,000 - 320,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

General Compute is building a pioneering neocloud for alternative chips to accelerate AI inference. You will own the inference serving stack end-to-end, from request routing to autoscaling, while ensuring reliability under real customer load and tight latency demands.

This founding role sits at the intersection of hardware and software, demanding architectural judgment, collaboration with compiler and bring-up teams, and a relentless focus on throughput and cost-per-token reduction as the team

Qualifications

  • 5+ years building and operating production systems at the infrastructure layer.
  • Direct experience with LLM inference serving in production environments.
  • Comfortable owning reliability and being on-call for critical systems.
  • Strong systems fundamentals, including concurrency, networking, and scheduling.
  • Self-directed, able to work in ambiguity as a founding engineer.

Responsibilities

  • Own the inference serving stack end-to-end, including routing, batching, scheduling, and autoscaling.
  • Continuously tune batching, KV-cache handling, and hardware utilization for throughput advantages.
  • Build for reliability with monitoring, alerting, and failover for real customer load.
  • Define interfaces between model compile and live deployment with the compiler and bring-up teams.
  • Shape the serving roadmap, including multi-tenant isolation and new scheduling strategies.
  • Set technical bar and influence how the serving team grows and maintains code quality.

Skills

LLM inference serving
High-throughput systems
On-call reliability
Performance optimization
Ambiguity tolerance

Tools

vLLM
TensorRT-LLM
TGI
SGLang

Job description

About us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.

About the Role

Getting a model correct and fast on our silicon is only half the problem — the other half is serving it. You'll build and own the inference layer that sits between a bought-up model and a live customer request: request scheduling, batching, KV-cache management, autoscaling across our ASIC fleet, and the failure modes that only show up at real traffic and real scale.

This is a founding role on a small team, which means the scope is wide and the ownership is real: there's no separate SRE org to hand reliability to and no platform team to hand infra to. You'll design the serving architecture, then be the person paged when it breaks. The bet is that a serving stack built specifically for our hardware — not adapted from a GPU-first framework — is a durable edge, and you're the person who proves that out in production.

What You’ll Do:
  • Own the inference serving stack end-to-end. Design and build the system that takes a bring-up-verified model and serves it in production: request routing, batching, scheduling, and autoscaling a single model's serving replicas.

  • Push cost-per-token down. Continuously tune batching strategy, KV-cache handling, and hardware utilization to widen the throughput advantage over GPU-based serving.

  • Build for reliability from day one. Put in place the monitoring, alerting, and failover that make a fast-moving inference stack trustworthy under real customer load — and be the one who responds when it isn't.

  • Work at the boundary with the compiler and bring-up team. Define the interface between "a model is correct and compiled" and "a model is live and fast," and push issues back to the right side of that line.

  • Shape the roadmap, not just the backlog. As a founding engineer, you'll help decide what we build next in serving — multi-tenant isolation, speculative decoding, new scheduling strategies — not just execute a spec someone else wrote.

  • Set the technical bar for the team you're helping build. Early architecture and code-quality decisions you make here will shape how the serving team operates as it grows.

What We Need From You:
  • 5+ years building and operating production systems at the infrastructure layer, ideally including a high-throughput or low-latency serving system.

  • Direct experience with LLM inference serving — request batching, KV-cache management, continuous batching, or similar — in a production environment, not just research code.

  • Comfortable owning reliability: you've been on call for a system that mattered, and you design for failure rather than reacting to it after the fact.

  • Strong systems fundamentals — concurrency, networking, scheduling — deep enough to reason about performance at the hardware level, not just the application level.

  • Self-directed and comfortable with ambiguity. This is a founding role: there's no existing playbook to follow, and you'll help write it.

Nice-to-Haves:
  • Experience serving models on non-NVIDIA accelerators (TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar).

  • Familiarity with serving frameworks such as vLLM, TGI, TensorRT-LLM, or SGLang, and an opinion on where they fall short.

  • Experience running infrastructure at a small company or in a founding/early-engineer capacity before.

  • Exposure to capacity planning or fleet management for specialized hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Platform Engineer
Founding Platform Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Head of Infrastructure
Head of Infrastructure

General Compute • San Francisco (CA)

On-site
USD 190,000 - 280,000
Inference Engineering and Product Lead
Inference Engineering and Product Lead

United States Digital Space LLC • San Francisco (CA)

On-site
USD 240,000 - 360,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
Head of Inference, Perimeter Compute
Head of Inference, Perimeter Compute

Montauk Capital • New York (NY)

On-site
USD 150,000 - 200,000
Competitive compensation + equity
Studio support from Montauk Capital’s network
AI Inference Engineer
AI Inference Engineer

Fuse Energy • United States

On-site
USD 180,000 - 300,000
Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
+2
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Model Bring-up Engineer / ML Compiler Engineer
Model Bring-up Engineer / ML Compiler Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits