Founding AI Inference Engineer

Fuse Energy

United States

On-site

USD 180,000 - 300,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
Private health insurance
Breakfast and dinner allowance

Job summary

Fuse Energy is hiring a founding engineer to own the inference serving layer, delivering high-throughput, low-latency model serving from kernel to endpoint. You will define the serving stack, partner with CUDA/GPU engineers, and shape the future of our energy AI platform.

In this founding role you will drive architecture decisions, experiment with quantisation and distillation techniques, and establish standards as the function scales, with equity available. This is a globally remote opportunity.

Qualifications

  • 4+ years building or operating large-scale inference serving systems.
  • Hands-on experience with inference serving frameworks and optimization techniques.
  • Strong systems thinking across a large cluster.
  • Collaborate with GPU/CUDA engineers to integrate performance work.
  • Proven track record of making architecture calls and owning outcomes.
  • Founding role; ability to shape a new function.

Responsibilities

  • Define inference serving strategy and architecture from first principles.
  • Design and build the serving stack for high-throughput, low-latency workloads.
  • Own model-level optimisation including quantisation and distillation.
  • Make architecture calls on serving frameworks and orchestration.
  • Translate throughput, latency and uptime into concrete specs and capacity plans.
  • Act as direct technical owner of inference performance and reliability.
  • Collaborate with CUDA/GPU teams to integrate custom kernels into serving.
  • Set standards, tooling and benchmarks as the function grows.

Skills

Inference serving
GPU/CUDA
Distributed systems
System design
Kubernetes/Slurm
LLMs

Tools

TensorRT
vLLM
Triton
SGLang

Job description

Fuse Energy is hiring a founding engineer to own the inference serving layer, delivering high-throughput, low-latency model serving from kernel to endpoint. You will define the serving stack, partner with CUDA/GPU engineers, and shape the future of our energy AI platform.

In this founding role you will drive architecture decisions, experiment with quantisation and distillation techniques, and establish standards as the function scales, with equity available. This is a globally remote opportunity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Fuse Energy • United States

On-site
USD 180,000 - 300,000
Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
+2
Founding AI Inference Architect
Founding AI Inference Architect

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Founding Platform Engineer — AI Inference Cloud
Founding Platform Engineer — AI Inference Cloud

General Compute • San Francisco (CA)

On-site
USD 180,000 - 260,000
CUDA Engineer: High-Perf GPU Kernels for Inference
CUDA Engineer: High-Perf GPU Kernels for Inference

Fuse Energy • United States

On-site
USD 140,000 - 210,000
Biannual bonus
Equity eligibility
Fully expensed tech
+2
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
LLM Inference Engineer — High-Performance AI Serving
LLM Inference Engineer — High-Performance AI Serving

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
Inference Engineer
Inference Engineer

Hyperbolic Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Lead Architect, AI Inference & Disaggregated Serving
Lead Architect, AI Inference & Disaggregated Serving

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
AI Inference Engineer
AI Inference Engineer

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000