MLOps Engineer

GCS Recruitment Specialists

United Arab Emirates

On-site

AED 250,000 - 420,000

Full time

37 hours ago
Be an early applicant
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

GCS Recruitment Specialists in the United Arab Emirates seeks a Senior MLOps Engineer to lead the development and management of infrastructure for training, deploying, and maintaining ML models in production.

You will design scalable pipelines for inference serving, model evaluation, and continuous delivery, collaborate with researchers and engineers, and own reliability, cost governance, and observability across hybrid cloud environments.

Qualifications

  • 5+ years in MLOps, ML infrastructure, or ML engineering with production lifecycles.
  • Hands-on experience in model inference serving across a range of sizes and architectures.
  • Strong Python skills; C/C++ experience is a plus.

Responsibilities

  • Design and manage infrastructure for training, deployment, and monitoring of ML models.
  • Lead inference serving strategies across model ranges and SLA requirements.
  • Build and maintain automated pipelines for training, evaluation, versioning, and CD/CI (MLflow, Kubeflow, SageMaker Pipelines).
  • Ensure reliability, observability, alerting, and incident response for ML services.
  • Implement cost governance and multi-provider routing to optimize spend and performance.

Skills

MLOps
ML Infra
Python
Cloud

Education

Bachelor's or Master's in CS/ML

Tools

MLflow
Kubeflow
SageMaker Pipelines
DeepSpeed
Accelerate

Job description

Senior MLOps Engineer
Role Summary

We are seeking a Senior MLOps Engineer to lead the development and management of infrastructure designed for training, deploying, and maintaining ML models. This role plays a critical function in operationalizing state-of-the-art systems to ensure high-performance delivery across research and production environments.

The successful candidate will be responsible for designing and implementing infrastructure to support efficient model deployment, inference, monitoring, and retraining. This includes close collaboration with cross-functional teams to integrate machine learning models into scalable and secure production pipelines, enabling the delivery of real-time, data-driven solutions across various domains.

Key Responsibilities
  • Inference Serving Across a Wide Model Range: Deploy and scale self-hosted open-weight models from ~7B up to ~376B parameters using engines such as vLLM, Triton, or TGI, choosing serving strategies (continuous batching, tensor/pipeline parallelism, quantization) appropriate to each model's size and SLA.
  • Multi-Provider Gateway & Cost Governance: Operate and improve the routing layer that spans self-hosted models and external APIs (OpenAI, Anthropic, OpenRouter). Build token accounting, budget controls, and cost/latency/quality-aware routing for model consumption predictably and within budget.
  • Training & Fine-Tuning Infrastructure: Build and maintain automated pipelines for fine-tuning, evaluation, versioning, and continuous delivery (MLflow, SageMaker Pipelines, or Kubeflow), including distributed training with DeepSpeed, FSDP, or Accelerate.
  • Reliability & Observability: Own production reliability for ML services, including monitoring, logging, alerting, incident response, and safe rollback, to meet latency, throughput, and availability targets. Participate in on-call for the serving platform.
  • Evaluation & Regression Safety: Stand up evaluation and verification harnesses that catch quality and performance regressions before they reach users, ensuring model or infrastructure changes ship with evidence rather than hope.
  • Platform as a Product: Treat internal engineers and external end-users as customers by reducing friction through sensible defaults, self-service tooling, and clear contracts. Deliver Infrastructure-as-Code (IaC) and CI/CD for repeatable, secure deployments.
  • Model & Cost Efficiency: Apply quantization, pruning, and multi-GPU/distributed inference to reduce latency and cost, especially at the high end of the model range.
Qualifications
  • Professional Experience: 5+ years in MLOps, ML Infrastructure, or ML Engineering, owning end-to-end model lifecycles in production.
  • Production Ownership: Experience supporting ML services in production and handling real incidents, including degraded inference, GPU OOMs, and cost escalations.
  • Inference Serving Depth: Hands-on experience serving models across a range of sizes, with real decisions made around quantization and parallelism trade-offs under latency and cost constraints.
  • Cost & Capacity Discipline: Demonstrable track record of measuring and optimizing GPU and cloud spend for ML workloads.
  • Programming Skills: Strong Python skills; additional C/C++ experience for performance-sensitive workloads is advantageous.
  • Cloud & Orchestration: Strong experience with cloud services (e.g., AWS SageMaker, EC2, EKS, Lambda), Docker, and Kubernetes. Experience across major hyperscalers is beneficial.
  • Tooling & Distributed Training: Proficiency with MLOps frameworks (MLflow, Kubeflow, or SageMaker Pipelines) and distributed training frameworks (DeepSpeed, FSDP, Accelerate).
  • Bonus: Experience building evaluation and verification harnesses, or multi-provider LLM gateways with token and cost management.
  • Educational Background: Bachelor's or Master's degree in Computer Science, Machine Learning, Data Engineering, or a related field, or equivalent practical experience.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Inception42 • Abu Dhabi Emirate

On-site
AED 350,000 - 550,000
MLOps Engineer
MLOps Engineer

Salt Digital Recruitment • Abu Dhabi

On-site
AED 180,000 - 240,000
MLOps Engineer
MLOps Engineer

SUNDUS MANAGEMENT CONSULTANCY & STUDIES BUREAUL.L.C • Abu Dhabi

On-site
AED 280,000 - 420,000
MLOps Engineer
MLOps Engineer

ai71 • Abu Dhabi

On-site
AED 420,000 - 720,000
Competitive compensation
Flexible working environment
Health insurance
Machine Learning Engineer
Machine Learning Engineer

Client of Discovered MENA • Abu Dhabi

On-site
AED 240,000 - 480,000
Senior MLOps Architect: Scale ML Infra & On-Prem/Cloud
Senior MLOps Architect: Scale ML Infra & On-Prem/Cloud

Greenhouse Software, Inc. • Abu Dhabi

On-site
AED 300,000 - 420,000
Flexible MLOps Architect for Scalable AI Infra
Flexible MLOps Architect for Scalable AI Infra

ai71 • Abu Dhabi

On-site
AED 420,000 - 720,000
Competitive compensation
Flexible working environment
Health insurance
MLOps Engineer New Abu Dhabi, UAE
MLOps Engineer New Abu Dhabi, UAE

Greenhouse Software, Inc. • Abu Dhabi

On-site
AED 300,000 - 420,000
Artificial Intelligence (AI) & Machine Learning Engineer
Artificial Intelligence (AI) & Machine Learning Engineer

Jaheziya • Abu Dhabi

On-site
AED 300,000 - 520,000
Machine Learning Engineer
Machine Learning Engineer

Dicetek LLC • Abu Dhabi

On-site
AED 300,000 - 550,000