MLOps Engineer

Fathom.io

Dhahran Compound

On-site

SAR 300,000 - 600,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Fathom.io in Dhahran, Saudi Arabia, seeks a mid-to-senior MLOps engineer to build the intelligence layer of our AI platform, enabling self-service deployment, training, notebooks, functions, and RAG workflows.

You will work at the crossroads of platform engineering, ML infrastructure, and distributed systems to deliver scalable, observable infrastructure for AI agents and LLM applications.

Qualifications

  • Strong experience in MLOps, AI platform engineering, or distributed systems.
  • Hands-on Kubernetes experience with stateful, training, notebooks, serverless, or GPU workloads.
  • Experience building ML platforms supporting training, notebooks, pipelines, registries, and production serving.
  • Familiarity with ML lifecycle tooling like MLflow, experiment tracking, registries, and CI/CD for ML systems.
  • Experience optimizing training and inference for latency, throughput, and cost.
  • Understanding of GPU scheduling, quantization, batching, autoscaling, and multi-model serving.
  • Experience with RAG systems, AI agents, or supporting infrastructure.
  • Strong software engineering fundamentals and a product-minded platform design.
  • Ability to make complex workflows reliable and approachable for users.

Responsibilities

  • Design and build the infrastructure powering our Intelligence layer.
  • Enable automated workflows for training, deployment, lifecycle, and inference.
  • Build scalable foundations for users to create AI agents and RAG pipelines via the UI.
  • Develop platform capabilities for notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve serving, observability, versioning, evaluation, promotion, and rollback.

Skills

MLOps
Kubernetes
Distributed systems
Rust
GPU inference
MLflow
KServe
Notebooks
CI/CD for ML
LLM apps

Tools

LangFuse
Knative
KServe
MLflow
Triton
Ray Serve

Job description

About The Role

We’re looking for a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. This is not a traditional MLOps role focused only on maintaining pipelines and deployments. You will help create the underlying infrastructure that makes sophisticated AI workflows accessible through a simple, self-service product experience: model deployment, training, notebooks, functions, UI-driven agent creation, and RAG pipelines.

You’ll work at the intersection of platform engineering, machine learning infrastructure, distributed systems, and developer experience.

What You’ll Do
  • Design and build the infrastructure powering our Intelligence layer.
  • Enable reliable, automated workflows for model training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create, configure, and operate AI agents and RAG pipelines through the platform UI.
  • Develop the platform capabilities behind managed notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
  • Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
  • Automate workflows that otherwise require manual MLOps effort, with a focus on safe, self-service capabilities for platform users.
  • Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
  • Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.
Our current stack
You’ll Work With Technologies Including
  • Kubernetes and Knative
  • KServe and vLLM
  • MLflow
  • LangFuse
  • GPU infrastructure and model-serving workloads
  • RAG architectures, vector retrieval, agent workflows, and LLM applications

Our intelligence backend is built in Rust, so Rust experience is a strong advantage.

What We’re Looking For
  • Strong experience in MLOps, AI platform engineering, machine learning infrastructure, or distributed systems.
  • Hands-on Kubernetes experience, including deploying and operating stateful, training, notebook, serverless, or GPU-intensive workloads.
  • Experience building ML platforms that support some combination of training, experimentation, notebooks, feature/data workflows, model registries, and production serving.
  • Experience with model serving frameworks such as KServe, vLLM, Triton, Ray Serve, or similar.
  • Familiarity with ML lifecycle tooling such as MLflow, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
  • Practical experience optimizing training and inference workloads for latency, throughput, availability, and cost.
  • Understanding of GPU scheduling, resource allocation, model quantization, batching, autoscaling, and multi-model serving.
  • Experience with RAG systems, LLM applications, AI agents, or their supporting infrastructure.
  • Strong software engineering fundamentals and a product-minded approach to platform design.
  • Ability to make complex operational workflows reliable and approachable for users.
Nice to have
  • Rust experience or interest in working with Rust-based backend systems.
  • Experience running notebook environments such as JupyterHub, Kubeflow Notebooks, VS Code Server, or equivalent.
  • Experience creating secure, scalable execution environments for user-defined functions or jobs.
  • Experience deploying AI workloads on edge devices.
  • Familiarity with model optimization techniques, including quantization, compilation, pruning, and hardware-aware serving.
  • Experience with multi-tenant AI platforms, security controls, and model governance.
  • Familiarity with observability and evaluation tooling for ML and LLM systems.
What Success Looks Like

You will help us evolve from manually operated AI infrastructure to a platform where users can train, evaluate, deploy, and operate models; work in managed notebooks; run functions; create agents; and configure RAG workflows - with the platform handling as much of the operational complexity as possible.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Fathom.io • Saudi Arabia

On-site
SAR 240,000 - 480,000
MLOps + DevOps Engineer - Agentic AI & Platform
MLOps + DevOps Engineer - Agentic AI & Platform

SYNC • Saudi Arabia

On-site
SAR 280,000 - 560,000
MLOps/LLMOps Engineer
MLOps/LLMOps Engineer

datascience • Riyadh

On-site
SAR 260,000 - 380,000
AI Engineer
AI Engineer

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
MLOps Platform Engineer - Build AI Workflows
MLOps Platform Engineer - Build AI Workflows

Fathom.io • Saudi Arabia

On-site
SAR 240,000 - 480,000
AI/ML Automation Analyst
AI/ML Automation Analyst

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 180,000 - 300,000
AI Production Engineer — MLOps & DevOps on AWS
AI Production Engineer — MLOps & DevOps on AWS

SYNC • Saudi Arabia

On-site
SAR 280,000 - 560,000
Senior ML Platform Engineer, Inference & MLOps
Senior ML Platform Engineer, Inference & MLOps

Mollkom • Riyadh

Hybrid
SAR 150,000 - 210,000
AI Engineer
AI Engineer

azmtalent • Saudi Arabia

On-site
SAR 240,000 - 360,000
Machine Learning Engineer
Machine Learning Engineer

Infinniti • Riyadh

On-site
SAR 360,000 - 540,000