Senior Software Engineer, Machine Learning Infrastructure & Automation

fal

United States

Remote

USD 140,000 - 190,000

Full time

43 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision insurance (

Job summary

fal is empowering the next generation of AI products by building scalable infra, tools, and model access for developers and enterprises. This role focuses on accelerating ML development by owning and improving CI/CD systems and automation for a growing suite of models and inference pipelines.

You will collaborate with Applied ML and ML Performance teams to ship new models and optimizations with confidence, designing automated model validation, performance benchmarks, and AI-driven workflows that

Qualifications

  • 5+ years of software engineering background with production infrastructure experience.
  • Strong Python proficiency and ability to build tooling for developers.
  • Experience designing and operating CI/CD systems, especially with GitHub Actions.
  • Deep understanding of automated testing, build systems, caching, and parallel execution.
  • Experience with containerized workloads (Docker) and cloud infrastructure.
  • Ability to design reliable distributed systems and diagnose complex failures.
  • Strong observability skills: logs, metrics, tracing, and alerting.
  • Proven track record of improving developer productivity through automation.

Responsibilities

  • Own ML CI/CD infrastructure for ML models and inference services.
  • Design, build, and maintain automated testing, validation, and deployment pipelines.
  • Reduce CI execution times via parallelization, caching, and efficient resource use.
  • Develop automated model validation to detect quality regressions across models and GPUs.
  • Build continuous performance testing for latency, throughput, GPU utilization, and cost.
  • Create automated pricing and deployment checks before production rollout.
  • Advance agentic engineering workflows to diagnose CI failures and propose fixes.
  • Improve deployment reliability with safeguards, verification, and rollback strategies.
  • Eliminate repetitive toil by building end-to-end automation for ML teams.

Skills

Python
CI/CD
Automation
Observability
Distributed systems

Tools

Docker
GitHub Actions
Cloud platforms

Job description

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

About this role:

Help fal's ML team move faster by building the automation, infrastructure, and developer tooling that makes developing, testing, and deploying generative AI models seamless.

You'll own and improve the CI/CD systems supporting our rapidly growing collection of ML models and inference pipelines. Your focus will be on eliminating manual work, accelerating development cycles, and building reliable systems that allow ML engineers to ship new models and optimizations with confidence.

This is a high-impact engineering role where you'll work closely with our Applied ML and ML Performance teams. You'll build everything from automated model validation and performance benchmarking to AI-powered development workflows that help engineers iterate faster.

The ideal candidate thinks beyond traditional CI/CD and sees automation as a force multiplier for the entire engineering organization.

What you’ll do:

  • Own ML CI/CD infrastructure Design, build, and maintain automated testing, validation, and deployment pipelines for our ML models and inference services.

  • Accelerate development cycles.Dramatically reduce CI execution times through intelligent parallelization, caching, test selection, and efficient use of compute resources.

  • Build automated model validation.Develop systems that test model outputs, detect quality regressions, and validate changes across different models, GPU architectures, and configurations.

  • Automate performance benchmarking. Build continuous performance testing that detects regressions in inference latency, throughput, GPU utilization, and cost.

  • Build automated pricing and deployment checks. Ensure model pricing, billing configurations, API schemas, and deployments are validated automatically before reaching production.

  • Extend our agentic engineering workflows. Develop AI-powered automation and agentic coding systems that automatically diagnose CI failures, identify regressions, propose fixes, and streamline engineering workflows.

  • Improve deployment reliability. Build automated safeguards, deployment verification, rollback mechanisms, and monitoring to ensure new model releases are reliable.

  • Eliminate engineering toil. Identify repetitive tasks across the ML team and build tools and systems that automate them, allowing engineers to focus on developing new models and improving performance.

Qualifications/Nice-to-haves:
  • 5+ years of software engineering background with proficiency in Python and experience building production infrastructure and developer tooling.

  • Experience designing and operating CI/CD systems using GitHub Actions or comparable technologies.

  • Deep understanding of automated testing, build systems, dependency management, caching, and parallel execution.

  • Experience working with containerized workloads, Docker, and cloud infrastructure.

  • Ability to design reliable distributed systems and debug complex infrastructure failures.

  • Strong understanding of observability, including logs, metrics, tracing, and automated alerting.

  • A passion for developer productivity and a demonstrated ability to eliminate manual processes through automation.

  • Comfortable working independently, identifying high-impact problems, and building end-to-end solutions.

  • Experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems.

  • Familiarity with NVIDIA GPU architectures and multi-GPU environments.

  • Experience building automated inference benchmarks or ML quality evaluation frameworks.

  • Experience with agentic coding tools such as Codex or Claude Code, or building custom AI engineering agents.

  • Experience optimizing CI/CD pipelines at scale, including distributed test execution and ephemeral compute environments.

  • Experience developing internal developer platforms or infrastructure-as-code tooling.

Tech Stack:

  • Python, PyTorch, Docker, GitHub Actions

  • NVIDIA GPUs and distributed GPU infrastructure

  • Model serving and inference pipelines

  • Observability and performance monitoring tools

  • AI coding agents and automated engineering workflows

  • Access to fal's massive GPU cluster for testing and development

What we offer at fal:

  • Interesting and challenging work

  • A lot of learning and growth opportunities

  • Health, dental, and vision insurance (US)

  • Regular team events and offsites

U.S. EQUAL EMPLOYMENT OPPORTUNITY INFORMATION:

fal provides equal employment opportunities to applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, or any other classification protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Applied Machine Learning
Software Engineer, Applied Machine Learning

Speedrun Talent Network • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/Dental/Vision
Team offsites
Growth opportunities
Software Engineer, Applied Machine Learning
Software Engineer, Applied Machine Learning

fal • San Francisco (CA)

On-site
USD 140,000 - 200,000
Health, dental, and vision insurance (
Regular team events and offsites
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Fal • United States

Remote
USD 180,000 - 250,000
Visa sponsorship
Relocation to San Francisco
Health insurance
Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal - Features & Labels • United States

Remote
USD 120,000 - 180,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation to San Francisco
Health, dental, and vision insurance (
Regular team events
+1
Product Marketer, Developer Platform
Product Marketer, Developer Platform

fal • San Francisco (CA)

On-site
USD 160,000 - 200,000
Visa sponsorship
Relocation to San Francisco
Health insurance
+3
Software Engineer, Infrastructure
Software Engineer, Infrastructure

The Consensus • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Software Engineer, Distributed Systems
Software Engineer, Distributed Systems

fal - Features & Labels • United States

Remote
USD 140,000 - 190,000
Interesting work
Learning opportunities
Team offsites
Technical Writer, Infrastructure
Technical Writer, Infrastructure

fal • San Francisco (CA)

On-site
USD 150,000 - 180,000
Visa sponsorship
Relocation to San Francisco
Health, dental, and vision insurance (
+1
Software Engineer, Machine Learning-Backend
Software Engineer, Machine Learning-Backend

Speedrun Talent Network • San Francisco (CA)

On-site
USD 120,000 - 190,000
Health, dental, and vision insurance (