MLOps Engineer Featured Office: Saudi Arabia

Fathom

As Sudiyah

On-site

OMR 12,000 - 24,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Fathom is building an AI platform that makes complex workflows accessible via a self-service UI. We’re seeking a mid-to-senior MLOps engineer to design infrastructure for the Intelligence layer, enabling model training, deployment, notebooks, and RAG pipelines.

You'll work at the intersection of platform engineering and ML infrastructure, deploying scalable GPU workloads, building observability, and aligning security and multi-tenancy with product goals.

Qualifications

  • Mid-to-senior level MLOps or AI platform engineering experience.
  • Hands-on Kubernetes experience with stateful workloads.
  • Experience building ML platforms that support training, notebooks, feature/data workflows, and model registries.
  • Familiarity with ML lifecycle tooling such as MLflow and CI/CD for ML systems.
  • Strong software engineering fundamentals and a product-minded approach to platform design.
  • Ability to simplify complex operational workflows for users.

Responsibilities

  • Design and build the infrastructure powering our Intelligence layer.
  • Enable automated workflows for model training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create and operate AI agents and RAG pipelines through the platform UI.
  • Develop the platform capabilities behind notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
  • Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
  • Automate workflows with a focus on safe, self-service capabilities for platform users.
  • Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
  • Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.

Skills

MLOps
Kubernetes
Rust
Distributed systems
Platform engineering

Tools

KServe
vLLM
MLflow
LangFuse

Job description

We’re looking for a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. This is not a traditional MLOps role focused only on maintaining pipelines and deployments. You will help create the underlying infrastructure that makes sophisticated AI workflows accessible through a simple, self-service product experience: model deployment, training, notebooks, functions, UI-driven agent creation, and RAG pipelines.

You’ll work at the intersection of platform engineering, machine learning infrastructure, distributed systems, and developer experience.

What you’ll do
  • Design and build the infrastructure powering our Intelligence layer.
  • Enable reliable, automated workflows for model training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create, configure, and operate AI agents and RAG pipelines through the platform UI.
  • Develop the platform capabilities behind managed notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
  • Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
  • Automate workflows that otherwise require manual MLOps effort, with a focus on safe, self-service capabilities for platform users.
  • Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
  • Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.
Our current stack

You’ll work with technologies including:

  • Kubernetes and Knative
  • KServe and vLLM
  • MLflow
  • LangFuse
  • GPU infrastructure and model-serving workloads
  • RAG architectures, vector retrieval, agent workflows, and LLM applications

Our intelligence backend is built in Rust, so Rust experience is a strong advantage.

What we’re looking for
  • Strong experience in MLOps, AI platform engineering, machine learning infrastructure, or distributed systems.
  • Hands-on Kubernetes experience, including deploying and operating stateful, training, notebook, serverless, or GPU-intensive workloads.
  • Experience building ML platforms that support some combination of training, experimentation, notebooks, feature/data workflows, model registries, and production serving.
  • Experience with model serving frameworks such as KServe, vLLM, Triton, Ray Serve, or similar.
  • Familiarity with ML lifecycle tooling such as MLflow, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
  • Practical experience optimizing training and inference workloads for latency, throughput, availability, and cost.
  • Understanding of GPU scheduling, resource allocation, model quantization, batching, autoscaling, and multi-model serving.
  • Experience with RAG systems, LLM applications, AI agents, or their supporting infrastructure.
  • Strong software engineering fundamentals and a product-minded approach to platform design.
  • Ability to make complex operational workflows reliable and approachable for users.
Nice to have
  • Rust experience or interest in working with Rust-based backend systems.
  • Experience running notebook environments such as JupyterHub, Kubeflow Notebooks, VS Code Server, or equivalent.
  • Experience creating secure, scalable execution environments for user-defined functions or jobs.
  • Experience deploying AI workloads on edge devices.
  • Familiarity with model optimization techniques, including quantization, compilation, pruning, and hardware-aware serving.
  • Experience with multi-tenant AI platforms, security controls, and model governance.
  • Familiarity with observability and evaluation tooling for ML and LLM systems.
What success looks like

You will help us evolve from manually operated AI infrastructure to a platform where users can train, evaluate, deploy, and operate models; work in managed notebooks; run functions; create agents; and configure RAG workflows - with the platform handling as much of the operational complexity as possible.

By submitting this application, I agree that my personal data will be collected, processed, and retained by the company solely for the purposes of managing and assessing my candidacy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Engineer
Senior ML Engineer

NTG • Muscat

On-site
OMR 30,000 - 50,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Employment • Muscat

On-site
OMR 58,000 - 73,000
Senior Machine Learning Engineer / Technical Lead
Senior Machine Learning Engineer / Technical Lead

PhazeRo • Oman

On-site
OMR 40,000 - 55,000
Senior ML Platform Engineer — National-Scale MLOps
Senior ML Platform Engineer — National-Scale MLOps

NTG • Muscat

On-site
OMR 30,000 - 50,000
AI Platform Engineer — Self-Service ML & RAG
AI Platform Engineer — Self-Service ML & RAG

Fathom • As Sudiyah

On-site
OMR 12,000 - 24,000
Senior AI Engineer
Senior AI Engineer

NTG • Muscat

On-site
OMR 40,000 - 60,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Phaze • Muscat

On-site
OMR 23,070 - 30,760
Senior/Lead Software Engineer - Backend
Senior/Lead Software Engineer - Backend

Intelmatix • As Sudiyah

Hybrid
OMR 26,915 - 34,605
Senior AI Engineer
Senior AI Engineer

Esbaar • Muscat

On-site
OMR 17,000 - 23,000
GENERATIVE-AI ENGINEER (OMANI NATIONAL)
GENERATIVE-AI ENGINEER (OMANI NATIONAL)

Nagarro • Muscat

On-site
OMR 16,000 - 23,000