MLOps Engineer

ai71

Abu Dhabi

On-site

AED 420,000 - 720,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Flexible working environment
Health insurance

Job summary

AI71 is seeking an experienced MLOps Engineer to define and own the ML infrastructure strategy across platform deployments, including LLMs and distributed models. You will shape deployment, reliability, and scale across SaaS and air-gapped on-prem environments.

Collaborate with researchers, product and engineering leaders to plan multi-quarter infrastructure roadmaps, mentor engineers, and drive cost-efficient performance improvements in a dynamic, security-conscious setting.

Qualifications

  • 10+ years in MLOps, ML infra, or ML engineering with architecture ownership.
  • Experience deploying large-scale models (LLMs) and ML infra at scale.
  • Deep cloud expertise across AWS, Azure, or GCP; strong Python.
  • Mentor engineers to operate independently at higher levels.

Responsibilities

  • Define ML infra architecture: model deployment, pipelines, and cloud infra.
  • Set reliability targets: monitoring, latency, availability, incident response.
  • Mentor senior MLOps engineers across teams.
  • Drive cross-team initiatives to improve inference performance and cost-efficiency.
  • Partner with ML researchers, product, and engineering on infra strategy.
  • Ensure infra scales for SaaS and on-prem air-gapped deployments.

Skills

MLOps leadership
ML infrastructure
Python proficiency
Kubernetes expertise
Cloud platforms (AWS,Apex,GCP)
Communication & stakeholder management

Tools

Kubeflow
MLflow
DeepSpeed
FSDP
Megatron-LM
CUDA

Job description

AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.

About AI71

AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.

The Role

As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.

What You'll Do
  • Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)
  • Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.
  • Mentor senior MLOps engineers; raise the operational bar across multiple teams.
  • Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
  • Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
  • Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.
What You'll Bring
  • 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.
  • Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
  • Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency
  • Mentorship record — engineers you have grown now operate independently at higher levels.
  • Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.
  • Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.
  • Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams.
Strong Preference
  • Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
  • Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
  • Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
  • Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.
  • Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.
  • On-prem / air-gap ML delivery architecture experience at scale.
  • Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.
  • Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.
Nice to Have
  • Conference speaking, technical writing, or industry thought leadership.
  • Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.
  • C/C++ or CUDA kernel experience for performance-critical paths.
  • Arabic language skills.
Why AI71
  • Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.
  • Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.
  • Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.
  • World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer New Abu Dhabi, UAE
MLOps Engineer New Abu Dhabi, UAE

Greenhouse Software, Inc. • Abu Dhabi

On-site
AED 300,000 - 420,000
Senior Engineering Manager
Senior Engineering Manager

AI71 • Abu Dhabi

On-site
AED 360,000 - 540,000
Flexible work environment
Competitive compensation
Career growth
+1
Flexible MLOps Architect for Scalable AI Infra
Flexible MLOps Architect for Scalable AI Infra

ai71 • Abu Dhabi

On-site
AED 420,000 - 720,000
Competitive compensation
Flexible working environment
Health insurance
MLOps Engineer
MLOps Engineer

Inception42 • Abu Dhabi Emirate

On-site
AED 350,000 - 550,000
MLOps Engineer
MLOps Engineer

SUNDUS MANAGEMENT CONSULTANCY & STUDIES BUREAUL.L.C • Abu Dhabi

On-site
AED 280,000 - 420,000
Staff Software Engineer - Full Stack
Staff Software Engineer - Full Stack

ai71 • Abu Dhabi

On-site
AED 360,000 - 600,000
Flexible working environment
Career growth opportunities
Competitive compensation
AI/ML DevOps Specialist
AI/ML DevOps Specialist

Netision Technology LLP • United Arab Emirates

On-site
AED 320,000 - 520,000
Competitive salary
Cutting-edge AI/ML tech
Career growth opportunities
Lead AI Scientist / Head of AI Solutions
Lead AI Scientist / Head of AI Solutions

EstateSight AI • Abu Dhabi

On-site
AED 450,000 - 900,000
Opportunity to lead AI innovation with societal impact
Research-driven environment
Leadership exposure across teams
Associate MLOps Engineer
Associate MLOps Engineer

AppliedAI • Abu Dhabi

On-site
AED 180,000 - 300,000
Health insurance
Visa sponsorship
21 days paid annual leave
+2
Staff Machine Learning Engineer
Staff Machine Learning Engineer

GCS • Abu Dhabi

Hybrid
AED 250,000 - 420,000