Get more replies from employers
Send a job-specific resume in minutes.
fal is building a large-scale generative media platform and seeks an ML Engineering / Site Reliability Engineer to own reliability, security, and safety of model APIs. You will ensure uptime, latency, and safety across thousands of productions endpoints, while tracing failures and improving deployment pipelines.
The role combines production ML engineering with site reliability practices, requiring deep knowledge of modern ML systems, security, and scalable infrastructure.
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
This is a hybrid ML Engineering / Site Reliability Engineering role. You will own the reliability, security, and safety of fal's fleet of generative media model APIs, the production endpoints that thousands of developers and enterprises depend on every day. Your mission is simple to state and hard to do: keep a large, fast-moving fleet of image, video, and audio model APIs available, performant, secure, and safe at all times.
You understand both how generative models work and how production systems fail. You're as comfortable debugging a misbehaving diffusion pipeline as you are tracing a latency regression through an inference stack, and you treat model-specific failure modes; degraded output quality, drift, unsafe generations, abuse patterns; as first-class reliability concerns alongside uptime and latency.
This role will need to be based in India, Australia, or New Zealand
Remote (India, Australia, New Zealand)