Machine Learning Engineer, Reliability

DataJobs

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

fal builds generative media model APIs, and this role helps keep them dependable, secure, and safe. You will own the reliability and operations side of the platform, ensuring model availability, performance, secure serving, and consistently safe generations for production traffic at scale.

You will collaborate with a team that iterates quickly on AI breakthroughs, with an emphasis on making speed reliable and safe for end users.

Qualifications

  • 5+ years of professional experience, including 2 years operating production ML or high-scale API systems.
  • Experience supporting diffusion models in production.
  • Strong systems fundamentals: distributed systems, networking, observability, and incident management.
  • Working knowledge of modern generative models (diffusion, transformers) and their production failure modes.
  • Familiarity with security and safety practices for ML systems, including abuse prevention or trust and safety engineering experience.
  • A bias toward automation, measurement, and blameless postmortems.

Responsibilities

  • Own availability, latency, and throughput SLOs across a large fleet of generative model APIs.
  • Build monitoring, alerting, and observability to detect ML-specific failures and regressions.
  • Harden deployment workflows with canary releases, shadow testing, and automated rollbacks.
  • Drive security posture for the model fleet including rate limiting and abuse prevention.
  • Operationalize safety systems for generative media at inference time without sacrificing performance.
  • Lead incident response for outages and postmortems to prevent recurrence.
  • Improve capacity planning, autoscaling, and GPU fleet efficiency.
  • Collaborate with model and infra teams to embed reliability, security, and safety into onboarding.

Skills

5+ years experience
Production ML
Diffusion models
Systems fundamentals
Security & safety
Automation & metrics

Tools

Python
Torch
Diffusers
Kubernetes
Fal Python SDK

Job description

fal builds generative media model APIs, and this role helps keep them dependable, secure, and safe. You will own the reliability and operations side of the platform, ensuring model availability, performance, secure serving, and consistently safe generations for production traffic at scale. You will also work with a team that iterates quickly on new AI breakthroughs, with an emphasis on making speed reliable.

What you’ll be responsible for
  • Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale.
  • Build monitoring, alerting, and observability to detect ML-specific failures, output quality degradation, pipeline breakage, and model regressions before customers experience issues.
  • Harden model deployment workflows using canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely.
  • Drive the security posture for the model fleet with secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns.
  • Operationalize safety systems for generative media, including content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without sacrificing performance.
  • Lead incident response for model API outages and degradations, run postmortems, and drive engineering changes to prevent recurrence.
  • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic.
  • Partner with model and infrastructure teams so reliability, security, and safety requirements are built into how new models are onboarded to the platform.
What you’ll bring
  • 5+ years of professional experience, including 2 years operating production ML or high-scale API systems, ideally with on-call ownership.
  • Experience supporting diffusion models in production.
  • Strong systems fundamentals: distributed systems, networking, observability, and incident management.
  • Working knowledge of modern generative models (diffusion, transformers) and their production failure modes.
  • Familiarity with security and safety practices for ML systems, including abuse prevention or trust and safety engineering experience (strong plus).
  • A bias toward automation, measurement, and blameless postmortems.
Tools you’ll use
  • Python, torch, diffusers, Kubernetes, fal Python SDK
Role details
  • Department: EngineeringML
  • Employment type: Full time
  • Location: Remote in APAC (India, Australia, or New Zealand)

You’ll have access to fal’s massive GPU cluster for inference and evaluation, and you’ll work with a team focused on quickly iterating and deploying AI improvements while maintaining reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Reliability Engineer - Generative ML APIs
Remote Reliability Engineer - Generative ML APIs

DataJobs • United States

Remote
USD 140,000 - 190,000
Machine Learning Engineer, Safety
Machine Learning Engineer, Safety

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Competitive salary and equity
Relocation assistance to San Francisco
Health, dental, and vision insurance (
+1
Senior Software Engineer, Machine Learning Infrastructure & Automation
Senior Software Engineer, Machine Learning Infrastructure & Automation

fal • United States

Remote
USD 140,000 - 190,000
Health, dental, and vision insurance (
Software Engineer, Applied Machine Learning
Software Engineer, Applied Machine Learning

fal • San Francisco (CA)

On-site
USD 140,000 - 200,000
Health, dental, and vision insurance (
Regular team events and offsites
Software Engineer, Applied Machine Learning
Software Engineer, Applied Machine Learning

Speedrun Talent Network • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/Dental/Vision
Team offsites
Growth opportunities
Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal - Features & Labels • United States

Remote
USD 120,000 - 180,000
Software Engineer, Distributed Systems
Software Engineer, Distributed Systems

fal - Features & Labels • United States

Remote
USD 140,000 - 190,000
Interesting work
Learning opportunities
Team offsites
Machine Learning Engineer
Machine Learning Engineer

Prodigy Resources • Denver (CO)

On-site
USD 160,000 - 210,000
Staff Machine Learning Systems & Reliability Engineer (Moveworks)
Staff Machine Learning Systems & Reliability Engineer (Moveworks)

ServiceNow • Mountain View (CA)

On-site
USD 250,000 - 320,000
Generous family leave
Annual learning stipend
Flexible PTO
+2
Machine Learning Engineer
Machine Learning Engineer

Evlo AI • Phoenix (AZ)

On-site
USD 120,000 - 165,000