Senior Platform Engineer, OpenAI-compatible Inference API

General Compute Inc.

San Francisco (CA)

On-site

USD 200,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

General Compute Inc. in San Francisco is seeking a senior systems engineer to build the platform layer for our inference cloud, including the OpenAI-compatible API surface and streaming infrastructure.

You will own surfaces, drive latency and throughput, and work with our hardware partner as we move more of the runtime in-house over the next 6–8 months. The first year emphasizes platform and serving, not kernels, with on-call duties for shipping work.

Qualifications

  • 5+ years writing production systems code.
  • Experience building and operating high-throughput API services with emphasis on latency.
  • Strong fundamentals in concurrency, memory, and queueing.
  • Knowledge of modern LLM inference concepts (transformers, attention, KV cache, batching).
  • Ability to measure and profile systems before relying on intuition.
  • Self-direction; capable of owning end-to-end surfaces without tickets.

Responsibilities

  • Own the OpenAI-compatible API surface: chat completions, streaming, tool use, error semantics, and edge cases for customers.
  • Build the request path that fronts our ASIC fleet — routing, admission control, queueing, retries, and degradation.
  • Integrate with OpenRouter and other distribution partners, manage request shapes and billing hooks.
  • Develop benchmarking and regression harnesses to catch latency and correctness drift.
  • Ship optimizations to improve TTFT, p99 latency, and throughput per dollar.
  • Collaborate with hardware runtime teams as layers move in-house; expand to batching, scheduling, KV-cache work.
  • Be oncall for what you ship.

Skills

Production systems
High-throughput API
Concurrency
Memory management
Queueing
Latency analysis
LLM inference basics
Batching
Profiling
Self-directed

Tools

OpenRouter
vLLM
TensorRT-LLM
llama.cpp
CUDA
Triton

Job description

General Compute Inc. in San Francisco is seeking a senior systems engineer to build the platform layer for our inference cloud, including the OpenAI-compatible API surface and streaming infrastructure.

You will own surfaces, drive latency and throughput, and work with our hardware partner as we move more of the runtime in-house over the next 6–8 months. The first year emphasizes platform and serving, not kernels, with on-call duties for shipping work.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer – Inference API & Streaming
Senior Platform Engineer – Inference API & Streaming

General Compute Inc. • New York (NY)

On-site
USD 130,000 - 160,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, AI Inference Infrastructure Platform
Software Engineer, AI Inference Infrastructure Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Platform Engineer for AI Inference & Optimization
Platform Engineer for AI Inference & Optimization

OpenAI • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 445,000
AI Infra Systems Engineer — GPT Deployment & Optimization
AI Infra Systems Engineer — GPT Deployment & Optimization

OpenAI • United States

On-site
USD 180,000 - 260,000
Inference Platform Backend Engineer (Equity & Benefits)
Inference Platform Backend Engineer (Equity & Benefits)

Together • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Health insurance
Competitive compensation
AI Systems Architect — Senior Platform Leader
AI Systems Architect — Senior Platform Leader

Accellor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Systems Architect - Inference & Platform
Senior AI Systems Architect - Inference & Platform

Worky • Mountain View (CA)

On-site
USD 260,000 - 320,000