Founding AI Inference Engineer — Scale AI Serving

Fuse Energy

United States

Remote

USD 200,000 - 350,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity sign-on bonus
Biannual bonus
Fully expensed tech
Meals allowance

Job summary

Fuse Energy is seeking a Founding AI Inference Engineer to define and build how we serve AI inference workloads at scale, reporting to the CTO. You will own the inference serving layer, ensuring throughput, latency, and uptime commitments across a large cluster, in collaboration with CUDA/GPU engineers.

The role combines architecture ownership with hands-on implementation, shaping a new function in a fast-growing energy-tech startup that aims to combine energy delivery with AI-scale compute.

Qualifications

  • Deep, hands-on experience with inference serving frameworks and techniques to optimise them (batching, KV-cache management, quantisation, speculative decoding).
  • Ability to reason about full path from incoming request to served response across a large cluster.
  • Comfort working with GPU/CUDA engineers to integrate low-level performance work.

Responsibilities

  • Define inference serving strategy and architecture from first principles.
  • Design and build the serving stack: routing, batching, scheduling, autoscaling for high-throughput, latency-sensitive workloads.
  • Own model-level optimisation strategy for serving to improve throughput and cost per token.

Skills

Inference serving
System design
Performance tuning
CUDA/GPU integration
Quantisation & distillation
ML inference frameworks

Tools

vLLM
TensorRT
Triton Inference Server
CUDA

Job description

Fuse Energy is seeking a Founding AI Inference Engineer to define and build how we serve AI inference workloads at scale, reporting to the CTO. You will own the inference serving layer, ensuring throughput, latency, and uptime commitments across a large cluster, in collaboration with CUDA/GPU engineers.

The role combines architecture ownership with hands-on implementation, shaping a new function in a fast-growing energy-tech startup that aims to combine energy delivery with AI-scale compute.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Fuse Energy • United States

Remote
USD 200,000 - 350,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Founding AI Inference Architect
Founding AI Inference Architect

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Founding Infrastructure Engineer, AI Inference Platform
Founding Infrastructure Engineer, AI Inference Platform

Collective Intuition, Inc. • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 200,000
Founding Platform Engineer — AI Inference Cloud
Founding Platform Engineer — AI Inference Cloud

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff AI Platform Engineer — Inference & Scale Leader
Staff AI Platform Engineer — Inference & Scale Leader

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 220,000 - 330,000
Founding AI Infra Engineer — Multi-Node Inference & Scale
Founding AI Infra Engineer — Multi-Node Inference & Scale

Piris Labs • San Francisco (CA)

On-site
USD 100,000 - 200,000
Equity
401(k)
Health insurance
+1
Founding Infrastructure Engineer – AI Inference Platform
Founding Infrastructure Engineer – AI Inference Platform

Observable Intuition, Inc. • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 200,000
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Senior AI Inference Architect — Scale & Disaggregated Serving
Senior AI Inference Architect — Scale & Disaggregated Serving

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Inference Cloud Architect: Build Scalable AI Infrastructure
Inference Cloud Architect: Build Scalable AI Infrastructure

techire.® • San Francisco (CA)

On-site
USD 180,000 - 240,000