Founding AI Inference Architect

General Compute

San Francisco (CA)

On-site

USD 180,000 - 320,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

General Compute is building a pioneering neocloud for alternative chips to accelerate AI inference. You will own the inference serving stack end-to-end, from request routing to autoscaling, while ensuring reliability under real customer load and tight latency demands.

This founding role sits at the intersection of hardware and software, demanding architectural judgment, collaboration with compiler and bring-up teams, and a relentless focus on throughput and cost-per-token reduction as the team

Qualifications

  • 5+ years building and operating production systems at the infrastructure layer.
  • Direct experience with LLM inference serving in production environments.
  • Comfortable owning reliability and being on-call for critical systems.
  • Strong systems fundamentals, including concurrency, networking, and scheduling.
  • Self-directed, able to work in ambiguity as a founding engineer.

Responsibilities

  • Own the inference serving stack end-to-end, including routing, batching, scheduling, and autoscaling.
  • Continuously tune batching, KV-cache handling, and hardware utilization for throughput advantages.
  • Build for reliability with monitoring, alerting, and failover for real customer load.
  • Define interfaces between model compile and live deployment with the compiler and bring-up teams.
  • Shape the serving roadmap, including multi-tenant isolation and new scheduling strategies.
  • Set technical bar and influence how the serving team grows and maintains code quality.

Skills

LLM inference serving
High-throughput systems
On-call reliability
Performance optimization
Ambiguity tolerance

Tools

vLLM
TensorRT-LLM
TGI
SGLang

Job description

General Compute is building a pioneering neocloud for alternative chips to accelerate AI inference. You will own the inference serving stack end-to-end, from request routing to autoscaling, while ensuring reliability under real customer load and tight latency demands.

This founding role sits at the intersection of hardware and software, demanding architectural judgment, collaboration with compiler and bring-up teams, and a relentless focus on throughput and cost-per-token reduction as the team

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Platform Engineer — AI Inference Cloud
Founding Platform Engineer — AI Inference Cloud

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Head of Inference Infrastructure
Head of Inference Infrastructure

General Compute • San Francisco (CA)

On-site
USD 190,000 - 280,000
Founding AI Inference Engineer
Founding AI Inference Engineer

Fuse Energy • United States

On-site
USD 180,000 - 300,000
Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
+2
Founding Platform Engineer
Founding Platform Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Chief Architect, Cluster-Scale AI Inference
Chief Architect, Cluster-Scale AI Inference

The Consensus • San Jose (CA)

On-site
USD 260,000 - 520,000
Medical, dental, and vision packages
Housing subsidy near Santana Row
Relocation support to San Jose
+3
Lead Architect, AI Inference & Disaggregated Serving
Lead Architect, AI Inference & Disaggregated Serving

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Cloudflare • Austin (TX)

Hybrid
USD 180,000 - 240,000
AI Inference Compute Engineer — Hardware/Software Co-Design
AI Inference Compute Engineer — Hardware/Software Co-Design

Google DeepMind • San Francisco (CA)

On-site
USD 174,000 - 252,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance