MLOps Engineer

Fathom.io

Saudi Arabia

On-site

SAR 240,000 - 480,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Fathom.io is seeking a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. You will design infrastructure for model deployment, training, notebooks, functions, and RAG pipelines, enabling self-service workflows.

You will work at the intersection of platform engineering, ML infrastructure, distributed systems, and developer experience, delivering scalable, secure solutions for GPU-heavy workloads and edge deployments.

Qualifications

  • Must have strong experience in MLOps, AI platform engineering, ML infrastructure, or distributed systems.
  • Hands-on Kubernetes experience with stateful, training, notebook, serverless, or GPU workloads.
  • Experience building ML platforms supporting training, experiments, notebooks, feature/data workflows, model registries, and production serving.
  • Familiarity with model serving frameworks such as KServe, vLLM, Triton, or Ray Serve.
  • Experience with ML lifecycle tooling, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
  • Experience optimizing training and inference workloads for latency, throughput, and cost.
  • Understanding of GPU scheduling, quantization, batching, autoscaling, and multi-model serving.

Responsibilities

  • Design and build the infrastructure powering our Intelligence layer.
  • Enable automated workflows for training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create AI agents and RAG pipelines through the platform UI.
  • Develop capabilities for notebooks, functions, experiments, training jobs, and model registries.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance and cost.
  • Explore deployment of models across centralized GPUs and edge devices.
  • Automate workflows to reduce manual MLOps effort with safe, self-service features.
  • Collaborate with backend, product, and AI teams to turn infrastructure into platform features.
  • Help define security, multi-tenancy, resource isolation, and model governance standards.

Skills

MLOps
AI platform engineering
Distributed systems
Kubernetes
Platform design
Product-minded
Self-service UX

Tools

Kubernetes
Knative
KServe
vLLM
MLflow
LangFuse

Job description

About The Role

We're looking for a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. This is not a traditional MLOps role focused only on maintaining pipelines and deployments. You will help create the underlying infrastructure that makes sophisticated AI workflows accessible through a simple, self-service product experience: model deployment, training, notebooks, functions, UI-driven agent creation, and RAG pipelines.

You'll work at the intersection of platform engineering, machine learning infrastructure, distributed systems, and developer experience.

What You'll Do
  • Design and build the infrastructure powering our Intelligence layer.
  • Enable reliable, automated workflows for model training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create, configure, and operate AI agents and RAG pipelines through the platform UI.
  • Develop the platform capabilities behind managed notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
  • Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
  • Automate workflows that otherwise require manual MLOps effort, with a focus on safe, self-service capabilities for platform users.
  • Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
  • Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.
Our current stack
You'll Work With Technologies Including
  • Kubernetes and Knative
  • KServe and vLLM
  • MLflow
  • LangFuse
  • GPU infrastructure and model-serving workloads
  • RAG architectures, vector retrieval, agent workflows, and LLM applications

Our intelligence backend is built in Rust, so Rust experience is a strong advantage.

What We're Looking For
  • Strong experience in MLOps, AI platform engineering, machine learning infrastructure, or distributed systems.
  • Hands-on Kubernetes experience, including deploying and operating stateful, training, notebook, serverless, or GPU-intensive workloads.
  • Experience building ML platforms that support some combination of training, experimentation, notebooks, feature/data workflows, model registries, and production serving.
  • Experience with model serving frameworks such as KServe, vLLM, Triton, Ray Serve, or similar.
  • Familiarity with ML lifecycle tooling such as MLflow, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
  • Practical experience optimizing training and inference workloads for latency, throughput, availability, and cost.
  • Understanding of GPU scheduling, resource allocation, model quantization, batching, autoscaling, and multi-model serving.
  • Experience with RAG systems, LLM applications, AI agents, or their supporting infrastructure.
  • Strong software engineering fundamentals and a product-minded approach to platform design.
  • Ability to make complex operational workflows reliable and approachable for users.
Nice to have
  • Rust experience or interest in working with Rust-based backend systems.
  • Experience running notebook environments such as JupyterHub, Kubeflow Notebooks, VS Code Server, or equivalent.
  • Experience creating secure, scalable execution environments for user-defined functions or jobs.
  • Experience deploying AI workloads on edge devices.
  • Familiarity with model optimization techniques, including quantization, compilation, pruning, and hardware-aware serving.
  • Experience with multi-tenant AI platforms, security controls, and model governance.
  • Familiarity with observability and evaluation tooling for ML and LLM systems.
What Success Looks Like

You will help us evolve from manually operated AI infrastructure to a platform where users can train, evaluate, deploy, and operate models; work in managed notebooks; run functions; create agents; and configure RAG workflows - with the platform handling as much of the operational complexity as possible.

By submitting this application, I agree that my personal data will be collected, processed, and retained by the company solely for the purposes of managing and assessing my candidacy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Fathom.io • Dhahran Compound

On-site
SAR 300,000 - 600,000
MLOps + DevOps Engineer - Agentic AI & Platform
MLOps + DevOps Engineer - Agentic AI & Platform

SYNC • Saudi Arabia

On-site
SAR 280,000 - 560,000
MLOps/LLMOps Engineer
MLOps/LLMOps Engineer

datascience • Riyadh

On-site
SAR 260,000 - 380,000
MLOps Platform Engineer - Build AI Workflows
MLOps Platform Engineer - Build AI Workflows

Fathom.io • Saudi Arabia

On-site
SAR 240,000 - 480,000
AI Engineer
AI Engineer

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
AI Engineer
AI Engineer

azmtalent • Saudi Arabia

On-site
SAR 240,000 - 360,000
Machine Learning Engineer
Machine Learning Engineer

Infinniti • Riyadh

On-site
SAR 360,000 - 540,000
ML, Engineer
ML, Engineer

Master Works • Riyadh

On-site
SAR 180,000 - 300,000
AI Production Engineer — MLOps & DevOps on AWS
AI Production Engineer — MLOps & DevOps on AWS

SYNC • Saudi Arabia

On-site
SAR 280,000 - 560,000
AI/ML Automation Analyst
AI/ML Automation Analyst

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 180,000 - 300,000