MLOps + DevOps Engineer - Agentic AI & Platform

SYNC

Saudi Arabia

On-site

SAR 280,000 - 560,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SYNC is seeking an MLOps + DevOps Engineer to build and operate an AI-native platform at the intersection of ML systems, backend infrastructure, and distributed workflows. You will ensure models, agents, APIs, and end-to-end workflows run reliably, scale predictably, and stay observable in production.

You will design cloud architectures, manage ML platforms on AWS, and implement CI/CD pipelines beyond basic GitHub Actions, while ensuring security and governance across services.

Qualifications

  • Expertise in AWS cloud services, EC2, S3, Lambda, ECS/EKS, and managed AI platforms.
  • Hands-on Docker and Kubernetes at production grade.
  • Experience with Apache Kafka or similar event-driven platforms.
  • Networking protocols, VPC, load balancers, security concepts.
  • Distributed systems and reliability engineering experience.
  • CI/CD pipelines with Jenkins, GitLab CI, or GitHub Actions.
  • Observability stack: Grafana, Prometheus, ELK, Sentry, OpenTelemetry.
  • Strong ownership, communication and code-review skills.

Responsibilities

  • Design cloud architectures for AI/ML workloads and production-grade systems.
  • Build and manage ML platforms using AWS services (EC2, ECS/EKS, Lambda, S3, RDS, VPC, IAM).
  • Leverage Bedrock, SageMaker or similar for model training and deployment.
  • Use Terraform, CloudFormation or CDK to automate cloud provisioning.
  • Work with vector and graph databases to support RAG and knowledge solutions.
  • Deploy and manage Small/Medium Language Models and external LLMs.
  • Build pipelines for model versioning, evaluation and fine-tuning.
  • Support RAG, embeddings and vector DB infra.
  • Design and operate event-driven backend infra – Kafka.
  • Handle async workflows, retries, ordering, idempotency and API communication.
  • Own Kubernetes cluster design, scaling strategies, workload isolation.
  • End-to-end CI/CD and multi-service deployment, environment isolation, blue/green.

Skills

Distributed systems
Reliability engineering
Observability
Ownership & communication
CI/CD practices
Security concepts

Tools

Docker
Kubernetes
Apache Kafka
Jenkins
GitLab CI
GitHub Actions
Grafana
Prometheus
ELK
OpenTelemetry

Job description

MLOps + DevOps Engineer - Agentic AI & Platform

We are building an AI-native platform where agentic systems, backend services, and real workflows operate together. This role sits at the intersection of ML Systems, Backend infrastructure and Distributed system operations. You will be responsible for ensuring that models, agents, APIs and workflows run reliably, scale predictably and remain observable end to end. This is not a traditional ML/DevOps role, but it is about operating intelligent systems in production.

What You'll Work On:

1. Design and implement cloud architectures supporting AI/ML workloads and production-grade systems

2. Build and manage ML platforms using AWS services including EC2, ECS/EKS, Lambda, S3, RDS, VPC, and IAM

3. Leverage AWS Bedrock, SageMaker, or similar managed AI services for model training and deployment

4. Use Infrastructure-as-Code tools such as Terraform, CloudFormation, or CDK to automate cloud provisioning

5. Work with vector databases (Milvus, Pinecone, Weaviate) and graph databases (Neo4j) to support retrieval-based and knowledge-driven AI solutions

6.1.Deploy and manage Small and Medium Language Models (SLMs)

6.2.Manage external LLM integrations

6.3. Build pipeline for model versioning, evaluation and fine-tuning

6.4. Support RAG systems, embeddings and Vector database infra

7. Agentic System Runtime

7.1.Enable execution of multi-agent workflows

7.3.Ensure consistency, fault-tolerance and latency control

8. Design and operate event-driven backend infrastructure – Apache Kafka (or equivalent)

9. Handle async workflows, retries, ordering, idempotency and enable reliable communication between backend and AI

10.Own Kubernetes cluster design, scaling strategies and workload isolation

  • Support micro-services and model-serving workloads
  • Design and manage API Gateway between Frontend and Backend layers
  • Implement routing, auth, rate limiting and service protection at Edge layers
  • Build end to end CI/CD (beyond basic GitHub Actions) and support multi-service deployment, environment isolation, rollback strategies, secret/vault management, blue/green deployment, feature flag support
  • Designing, managing and handling data storage infra layer for RDBMS, Vector DB, Document DB, Caching, Object storage.
  • Instrumentation, Observability and Debugging
    • You will design and own end to end observability across AI + Backend + Frontend + Infra
    • Implement structures instrumentation across APIs, async workflows, AI agents and capture request lifecycle, agent decision paths and execution timelines
    • Design centralised logging – structured logs (JSON) and contextual logging (correlation IDs)
    • Distributed tracing across tiers and service layers – user journey, agent decisions and backend actions
    • Build strong monitoring with metrics for system health, API performance, LLM token burns, Queue lag and model latency
  • Owning and commanding triage & debugging with engineering teams for multiservice failures, AI <> Backend inconsistencies, root cause analysis and replaying failures
  • You will be managing security and secrets by securing APIs, model configs, infra credentials
  • You will be implementing RBAC, secret rotation and environment isolation
Must-have:
  • Expertise in AWS cloud services, EC2, S3, Lambda, ECS/EKS, managed AI platforms, model deployment and distributed systems
  • Strong hands-on experience with Docker, Kubernetes at Production grade
  • Experience with Apache Kafka or similar event-driven
  • Networking protocols, Security concepts, VPC, Load balancers
  • Strong experience in distributed systems and reliability engineering
  • Hands-on with CI/CD pipeline - Jenkins, GitLab CI, GitHub Actions, automated testing
  • Experience with Observability stack (logs, metrices, tracing) – Grafana, Prometheus, ELK, Sentry, Phoenix, Arize, Open Telemetry.
  • Strong ownership, communication, and code review skills.
  • You're based in Saudi Arabia and holding transferable Iqama.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer
AI Engineer

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
AI Engineer
AI Engineer

Latitude • Riyadh

On-site
SAR 420,000 - 660,000
Senior DevOps Engineer
Senior DevOps Engineer

Mirai Arabian International Company Limited • Riyadh

On-site
SAR 180,000 - 230,000
AI Production Engineer — MLOps & DevOps on AWS
AI Production Engineer — MLOps & DevOps on AWS

SYNC • Saudi Arabia

On-site
SAR 280,000 - 560,000
Ai/ml/llm Systems Engineer - Enterprise Ai Platform Engineer
Ai/ml/llm Systems Engineer - Enterprise Ai Platform Engineer

Saudi Aramco (ASC) • Saudi Arabia

On-site
SAR 299,000 - 450,000
AI Engineer - Agentic Systems & Applied AI
AI Engineer - Agentic Systems & Applied AI

SYNC • Jeddah

On-site
SAR 240,000 - 320,000
Senior AI Platform Engineer — MLOps, LLMs & RAG Systems
Senior AI Platform Engineer — MLOps, LLMs & RAG Systems

Latitude • Riyadh

On-site
SAR 420,000 - 660,000
AI/ML Associate Manager
AI/ML Associate Manager

Accenture Middle East • Riyadh

On-site
SAR 360,000 - 600,000
AI Engineer - Agentic
AI Engineer - Agentic

Master Works • Riyadh

On-site
SAR 240,000 - 360,000
AI & Machine Learning Consultant
AI & Machine Learning Consultant

Accenture Middle East • Riyadh

On-site
SAR 250,000 - 500,000