AI Inference Engineer

Fuse Energy

United States

On-site

USD 180,000 - 300,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
Private health insurance
Breakfast and dinner allowance

Job summary

Fuse Energy is hiring a founding engineer to own the inference serving layer, delivering high-throughput, low-latency model serving from kernel to endpoint. You will define the serving stack, partner with CUDA/GPU engineers, and shape the future of our energy AI platform.

In this founding role you will drive architecture decisions, experiment with quantisation and distillation techniques, and establish standards as the function scales, with equity available. This is a globally remote opportunity.

Qualifications

  • 4+ years building or operating large-scale inference serving systems.
  • Hands-on experience with inference serving frameworks and optimization techniques.
  • Strong systems thinking across a large cluster.
  • Collaborate with GPU/CUDA engineers to integrate performance work.
  • Proven track record of making architecture calls and owning outcomes.
  • Founding role; ability to shape a new function.

Responsibilities

  • Define inference serving strategy and architecture from first principles.
  • Design and build the serving stack for high-throughput, low-latency workloads.
  • Own model-level optimisation including quantisation and distillation.
  • Make architecture calls on serving frameworks and orchestration.
  • Translate throughput, latency and uptime into concrete specs and capacity plans.
  • Act as direct technical owner of inference performance and reliability.
  • Collaborate with CUDA/GPU teams to integrate custom kernels into serving.
  • Set standards, tooling and benchmarks as the function grows.

Skills

Inference serving
GPU/CUDA
Distributed systems
System design
Kubernetes/Slurm
LLMs

Tools

TensorRT
vLLM
Triton
SGLang

Job description

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We’ve raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We’re building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We’re building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we’re looking for the founding engineer to own the latter. Reporting directly to the CTO, you’ll own the layer above kernels and hardware: how models actually get served, scaled and delivered against committed performance targets. Few companies can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of our offering.

Responsibilities
  • Define Fuse’s inference serving strategy and architecture from first principles
  • Design and build the serving stack: request routing, batching, scheduling and autoscaling for high-throughput, latency-sensitive inference workloads
  • Own model-level optimisation strategy for serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)
  • Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plans
  • Act as direct technical owner of inference performance and reliability
  • Work closely with the CUDA and GPU engineering teams to integrate custom kernels and hardware performance work cleanly into the serving layer
  • Set the standards, tooling and benchmarks this function will run on as it grows
Requirements
  • 4+ years building or operating large-scale inference serving systems, or equivalent strong project/industry experience
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
  • Strong systems thinking, able to reason about the full path from incoming request to served response across a large cluster
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system
  • A track record of making high-stakes architecture calls and owning the outcome
  • Comfort operating without a playbook: this is a founding role shaping a new function, not joining an established one
  • Bonus: Triton or custom ML inference/training frameworks; autoscaling or capacity planning for large-scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute
Benefits
  • Competitive salary and eligibility for equity
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Private health insurance
  • Breakfast and dinner allowance for office-based employees

As we hire globally, benefits vary by location.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding AI Inference Engineer
Founding AI Inference Engineer

Fuse Energy • United States

On-site
USD 180,000 - 300,000
Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
+2
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
CUDA Engineer
CUDA Engineer

Fuse Energy • United States

On-site
USD 140,000 - 210,000
Biannual bonus
Equity eligibility
Fully expensed tech
+2
VP of Engineering
VP of Engineering

Fuse Energy • San Francisco (CA)

On-site
USD 320,000 - 480,000
Equity ownership
Biannual bonus
Tech budget & equipment
+2
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
AI Inference Engineer
AI Inference Engineer

F5 • San Jose (CA)

On-site
USD 140,000 - 210,000
Inference Engineer
Inference Engineer

Hyperbolic Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1