Software Engineer, Inference Platform

General Compute Inc.

San Francisco (CA)

On-site

USD 200,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

General Compute Inc. in San Francisco is seeking a senior systems engineer to build the platform layer for our inference cloud, including the OpenAI-compatible API surface and streaming infrastructure.

You will own surfaces, drive latency and throughput, and work with our hardware partner as we move more of the runtime in-house over the next 6–8 months. The first year emphasizes platform and serving, not kernels, with on-call duties for shipping work.

Qualifications

  • 5+ years writing production systems code.
  • Experience building and operating high-throughput API services with emphasis on latency.
  • Strong fundamentals in concurrency, memory, and queueing.
  • Knowledge of modern LLM inference concepts (transformers, attention, KV cache, batching).
  • Ability to measure and profile systems before relying on intuition.
  • Self-direction; capable of owning end-to-end surfaces without tickets.

Responsibilities

  • Own the OpenAI-compatible API surface: chat completions, streaming, tool use, error semantics, and edge cases for customers.
  • Build the request path that fronts our ASIC fleet — routing, admission control, queueing, retries, and degradation.
  • Integrate with OpenRouter and other distribution partners, manage request shapes and billing hooks.
  • Develop benchmarking and regression harnesses to catch latency and correctness drift.
  • Ship optimizations to improve TTFT, p99 latency, and throughput per dollar.
  • Collaborate with hardware runtime teams as layers move in-house; expand to batching, scheduling, KV-cache work.
  • Be oncall for what you ship.

Skills

Production systems
High-throughput API
Concurrency
Memory management
Queueing
Latency analysis
LLM inference basics
Batching
Profiling
Self-directed

Tools

OpenRouter
vLLM
TensorRT-LLM
llama.cpp
CUDA
Triton

Job description

Build the platform layer of our inference cloud — the OpenAI-compatible API surface, the request path, the streaming infrastructure, and the benchmarking harnesses that keep us honest. The runtime on our ASIC sits below you and is owned by our hardware partner today; your job is everything between an OpenRouter request landing on our edge and a token streaming back to the user.

This is a senior IC role on a small team. You'll own surfaces, not tickets. As we take more of the runtime in-house over the next 6-8 months, the work moves closer to the metal but the first year is platform and serving, not kernels.

Responsibilities
  • Own the OpenAI-compatible API surface: chat completions, streaming, tool use, error semantics, and the long tail of compatibility edge cases that matter to real customers.
  • Build the request path that fronts our ASIC fleet — routing, admission control, queueing, retries, and graceful degradation.
  • Own integration with OpenRouter and other distribution partners. Their request shapes, their billing hooks, their failure modes are your problem.
  • Build the benchmarking and regression harness that catches latency and correctness drift before customers do — TTFT, inter-token latency distributions, tokens/sec under load.
  • Ship optimizations that move real metrics: TTFT, p99 latency, throughput per dollar.
  • Work directly with our hardware partner's runtime team when the bottleneck is below your layer. As we take more of the stack in-house, your scope grows to include batching, scheduling, and KV-cache work directly.
  • Be on the oncall rotation for what you ship.
What we're looking for
  • 5+ years writing production systems code.
  • Have built and operated a high-throughput API service before — ideally one where tail latency mattered as much as throughput.
  • Strong fundamentals in concurrency, memory, queueing, and where latency actually comes from in a distributed system.
  • Working knowledge of modern LLM inference: transformers, attention, KV cache, batching, speculative decoding, quantization. You don't need to have written a kernel, but you should know why batching changes everything about a serving stack.
  • Comfortable with a profiler. You reach for measurement before intuition.
  • Self-directed. We don't have the bandwidth to assign you tickets — you'll find the work.
Nice to have
  • Have built or operated an OpenAI-compatible API at production scale.
  • Familiar with vLLM, TGI, TensorRT-LLM, SGLang, or llama.cpp internals.
  • Kernel-level work in CUDA, Triton, or on non-NVIDIA accelerators — relevant as the role grows down-stack.
  • Have shipped streaming infrastructure (SSE, gRPC streaming, WebSockets) under real load.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Platform
Software Engineer, Inference Platform

General Compute Inc. • New York (NY)

On-site
USD 130,000 - 160,000
Head of Infrastructure
Head of Infrastructure

General Compute Inc. • New York (NY)

On-site
USD 120,000 - 150,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Software Engineer, Inference & Training Platform) Perplexity AI San [...]
Member of Technical Staff (Software Engineer, Inference & Training Platform) Perplexity AI San [...]

Neura Market • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Distributed Systems Engineer, Real-Time Inference at Scale
Distributed Systems Engineer, Real-Time Inference at Scale

adaption • San Francisco (CA)

On-site
Model Bringup Engineer / ML Compiler Engineer
Model Bringup Engineer / ML Compiler Engineer

General Compute Inc. • New York (NY)

On-site
USD 180,000 - 300,000