MLOps Engineer New Abu Dhabi, UAE

Greenhouse Software, Inc.

Abu Dhabi

On-site

AED 300,000 - 420,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

AI71 is seeking an experienced MLOps Engineer to lead ML infrastructure strategy across SaaS and on-prem deployments. You will own deployment architectures, model serving at scale, and reliability targets while mentoring engineers across teams.

You will drive cross-team initiatives for performance, cost efficiency, and scalable inference, partnering with researchers and leadership to shape multi-quarter ML infrastructure strategy in a distributed cloud/on-prem environment.

Qualifications

  • 10+ years in MLOps, ML infrastructure, or ML engineering with architectural ownership.
  • Proven track record deploying large-scale models, including LLMs, and building ML infra at scale.
  • Deep cloud expertise across AWS/Azure/GCP and strong Python proficiency.
  • Mentorship history; ability to lead engineers to operate independently.
  • Experience with SaaS and on-prem deployments, including air-gapped environments.
  • Kubernetes at architectural depth; GPU scheduling and distributed workloads.

Responsibilities

  • Define ML infrastructure architecture: model deployment, pipelines, cloud-native infra.
  • Set reliability targets: monitoring, latency, throughput, incident response.
  • Mentor senior MLOps engineers and raise the operational bar.
  • Drive cross-team initiatives for inference performance and cost-efficiency.
  • Partner with researchers, product, and leadership on multi-quarter strategy.
  • Ensure scales across managed SaaS and on-prem deployments.

Tools

MLflow
Kubeflow
vLLM
Triton
TGI
DeepSpeed
FSDP
Accelerate
Slurm
CUDA
NCCL
Megatron-LM
NVLink
InfiniBand

Job description

AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.

The Role:

As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.

What You'll Do:
  • Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)
  • Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.
  • Mentor senior MLOps engineers; raise the operational bar across multiple teams.
  • Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
  • Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
  • Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.
What You'll Bring:
  • 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.
  • Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
  • Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency
  • Mentorship record — engineers you have grown now operate independently at higher levels.
  • Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.
  • Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.
  • Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams.
Strong Preference:
  • Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
  • Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
  • Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
  • Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.
  • Er…
  • Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.
  • On-prem / air-gap ML delivery architecture experience at scale.
  • Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.
  • Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.
Nice to Have:
  • Conference speaking, technical writing, or industry thought leadership.
  • Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.
  • C/C++ or CUDA kernel experience for performance-critical paths.
Why AI71:
  • Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.
  • Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.
  • Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.
  • World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

ai71 • Abu Dhabi

On-site
AED 420,000 - 720,000
Competitive compensation
Flexible working environment
Health insurance
Senior Engineering Manager
Senior Engineering Manager

AI71 • Abu Dhabi

On-site
AED 360,000 - 540,000
Flexible work environment
Competitive compensation
Career growth
+1
AI/ML DevOps Specialist
AI/ML DevOps Specialist

Netision Technology LLP • United Arab Emirates

On-site
AED 320,000 - 520,000
Competitive salary
Cutting-edge AI/ML tech
Career growth opportunities
Associate MLOps Engineer
Associate MLOps Engineer

AppliedAI • Abu Dhabi

On-site
AED 180,000 - 300,000
Health insurance
Visa sponsorship
21 days paid annual leave
+2
Staff Machine Learning Engineer
Staff Machine Learning Engineer

GCS • Abu Dhabi

Hybrid
AED 250,000 - 420,000
Flexible MLOps Architect for Scalable AI Infra
Flexible MLOps Architect for Scalable AI Infra

ai71 • Abu Dhabi

On-site
AED 420,000 - 720,000
Competitive compensation
Flexible working environment
Health insurance
Associate ML Ops Engineer
Associate ML Ops Engineer

Opus • Abu Dhabi

On-site
AED 120,000 - 180,000
21 days paid annual leave
Company health insurance
Visa sponsorship for international
Associate ML Ops Engineer
Associate ML Ops Engineer

AppliedAI • Abu Dhabi

On-site
AED 156,000 - 234,000
Health insurance
Visa sponsorship
On-site Abu Dhabi HQ
MLOps Engineer
MLOps Engineer

Inception42 • Abu Dhabi Emirate

On-site
AED 350,000 - 550,000
Associate ML Ops Engineer
Associate ML Ops Engineer

Tanqeeb • Abu Dhabi

On-site
AED 120,000 - 240,000
Health insurance
Visa sponsorship
21 days annual leave