Inference Engineer

techire.®

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

techire.® is building an early-stage AI infrastructure platform in San Francisco, focused on turning a heterogeneous fleet of AI compute into reliable, scalable infrastructure customers can consume. This role is hands-on from 0→1, with no mature platform yet, and involves making architectural decisions that shape cloud behavior from day one.

You’ll work on building the actual product, defining how scheduling, serving and reliability components interact, and exploring trade-offs in how the

Qualifications

  • Experience building distributed systems at scale and in production
  • Ability to design cloud infrastructure and control planes
  • Familiarity with scheduling, routing, or large-scale serving systems

Responsibilities

  • Designing the control plane across heterogeneous compute
  • Building scheduling, routing and workload/model placement systems
  • Creating customer-facing APIs and serving infrastructure
  • Designing reliability, observability and failure handling from the start
  • Scaling the platform as workloads and customer demand grow

Job description

Your focus will include:
  • Designing the control plane across heterogeneous compute
  • Building scheduling, routing and workload/model placement systems
  • Creating customer-facing APIs and serving infrastructure
  • Designing reliability, observability and failure handling from the start
  • Scaling the platform as workloads and customer demand grow

Most infrastructure engineers joining an AI cloud inherit a platform.

Here, you’ll build it.

You’ll join an early-stage AI infrastructure company building a new kind of inference cloud, turning a heterogeneous fleet of AI compute into reliable, scalable infrastructure that customers can actually consume.

There’s no mature platform, large infrastructure team or established playbook waiting for you.

You’ll be one of the first engineers making the architectural decisions that shape how the cloud works from day one.

This is a genuinely hands-on 0→1 role. You won’t be managing a team or maintaining infrastructure somebody else designed. You’ll be building the actual product.

The interesting part is how much is still unsolved.

You’ll have the freedom and responsibility to decide how these systems should work, rather than fitting into an existing architecture. Decisions you make now around scheduling, serving and reliability could remain part of the platform for years.

You’ll suit this if you’ve built serious distributed systems, cloud infrastructure, control planes, schedulers or large-scale serving systems and understand what reliable production infrastructure actually requires.

More importantly, you’ll have experience creating systems from scratch rather than simply operating mature ones.

LLM serving experience with tools such as vLLM, TGI or Ray Serve would be useful. So would experience with GPU infrastructure or alternative AI accelerators, but none of those are essential.

The sweet spot is someone who sees an unsolved infrastructure problem and thinks, “I’ll build it.”

If you’ve built distributed infrastructure at scale but want considerably more technical ownership than you’d get inside an established AI or cloud company, this is worth a conversation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Cloud Architect: Build Scalable AI Infrastructure
Inference Cloud Architect: Build Scalable AI Infrastructure

techire.® • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding Platform Engineer — AI Inference Cloud
Founding Platform Engineer — AI Inference Cloud

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Founding Platform Engineer
Founding Platform Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
AI Infrastructure Engineer - Inference Platform
AI Infrastructure Engineer - Inference Platform

Hoonify Technologies Inc. • Albuquerque (NM)

On-site
USD 120,000 - 190,000
Infrastructure Engineer - Series A Startup
Infrastructure Engineer - Series A Startup

Epic Placements • San Francisco (CA)

On-site
USD 150,000 - 230,000
Infrastructure System Engineer
Infrastructure System Engineer

Berkley Hunt • New York (NY)

On-site
USD 150,000 - 210,000
Senior/Staff AI Engineer
Senior/Staff AI Engineer

Data Direct Networks • California (MO)

On-site
USD 150,000 - 230,000
Vacation plans
Paid holidays
Bonus programs
+5