Remote Reliability Engineer - Generative ML APIs

DataJobs

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

fal builds generative media model APIs, and this role helps keep them dependable, secure, and safe. You will own the reliability and operations side of the platform, ensuring model availability, performance, secure serving, and consistently safe generations for production traffic at scale.

You will collaborate with a team that iterates quickly on AI breakthroughs, with an emphasis on making speed reliable and safe for end users.

Qualifications

  • 5+ years of professional experience, including 2 years operating production ML or high-scale API systems.
  • Experience supporting diffusion models in production.
  • Strong systems fundamentals: distributed systems, networking, observability, and incident management.
  • Working knowledge of modern generative models (diffusion, transformers) and their production failure modes.
  • Familiarity with security and safety practices for ML systems, including abuse prevention or trust and safety engineering experience.
  • A bias toward automation, measurement, and blameless postmortems.

Responsibilities

  • Own availability, latency, and throughput SLOs across a large fleet of generative model APIs.
  • Build monitoring, alerting, and observability to detect ML-specific failures and regressions.
  • Harden deployment workflows with canary releases, shadow testing, and automated rollbacks.
  • Drive security posture for the model fleet including rate limiting and abuse prevention.
  • Operationalize safety systems for generative media at inference time without sacrificing performance.
  • Lead incident response for outages and postmortems to prevent recurrence.
  • Improve capacity planning, autoscaling, and GPU fleet efficiency.
  • Collaborate with model and infra teams to embed reliability, security, and safety into onboarding.

Skills

5+ years experience
Production ML
Diffusion models
Systems fundamentals
Security & safety
Automation & metrics

Tools

Python
Torch
Diffusers
Kubernetes
Fal Python SDK

Job description

fal builds generative media model APIs, and this role helps keep them dependable, secure, and safe. You will own the reliability and operations side of the platform, ensuring model availability, performance, secure serving, and consistently safe generations for production traffic at scale.

You will collaborate with a team that iterates quickly on AI breakthroughs, with an emphasis on making speed reliable and safe for end users.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Reliability
Machine Learning Engineer, Reliability

DataJobs • United States

Remote
USD 140,000 - 190,000
Applied ML Engineer - Generative Media Platform
Applied ML Engineer - Generative Media Platform

fal • San Francisco (CA)

On-site
USD 140,000 - 200,000
Health, dental, and vision insurance (
Regular team events and offsites
Applied ML Engineer - Generative Media & Models Production
Applied ML Engineer - Generative Media & Models Production

Speedrun Talent Network • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/Dental/Vision
Team offsites
Growth opportunities
Backend Reliability Engineer, AI-Powered Platform (Remote)
Backend Reliability Engineer, AI-Powered Platform (Remote)

Affirm • Madison (WI)

On-site
USD 173,000 - 255,000
Health care coverage
ESPP
Time off
+1
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)

Affirm • Riverside (OH)

Remote
USD 173,000 - 233,000
Health coverage
FSA Wallets
Time off
+1
Advanced AI Platform Engineer - ML Pipelines & Production
Advanced AI Platform Engineer - ML Pipelines & Production

LE0001 Relativity ODA LLC • Illinois

Hybrid
USD 103,000 - 155,000
Equity program
Parental leave
DTO (PTO)
Remote Backend Reliability Engineer - AI-Driven Platform
Remote Backend Reliability Engineer - AI-Driven Platform

Affirm • Boulder (CO)

On-site
USD 173,000 - 255,000
Health care coverage
Flexible Spending Wallets
Time off
+1
SRE Lead: GenAI Platform Reliability (Remote)
SRE Lead: GenAI Platform Reliability (Remote)

BMC Software, Inc. • United States

Remote
USD 153,000 - 255,000
Senior Generative ML Engineer - Build & Deploy Models
Senior Generative ML Engineer - Build & Deploy Models

CyberCoders • Pittsburgh

On-site
USD 140,000 - 190,000
Vacation/PTO
Medical
Dental
+3
Remote Senior Backend Engineer — AI-Driven Reliability
Remote Senior Backend Engineer — AI-Driven Reliability

Affirm • Richmond (VA)

On-site
USD 173,000 - 233,000
Health care coverage
Flexible Spending Wallets
Time off
+1