Position Summary
Systemsltd is seeking an experienced, highly technical, and accomplished DevOps / Infrastructure / SRE / Platform Engineer with professional cloud infrastructure, GPU cluster administration, and artificial intelligence platform engineering experience to join our client engagement team on a full-time, onsite basis in Dubai. In this specialized AI infrastructure and site reliability role, you will spearhead designing, deploying, and optimizing robust Kubernetes clusters, NVIDIA GPU infrastructure, LLM/ML operational pipelines, and automated CI/CD deployment workflows. You will work closely with platform engineering squads, infrastructure architects, and technical leaders to ensure high availability, low latency, and zero-downtime production reliability for enterprise artificial intelligence systems. Ideal candidates bring a robust academic degree in computer science or engineering, deep practical mastery of cloud DevOps, Prometheus, Grafana, PostgreSQL, and immediate availability for onsite deployment in Dubai.
Detailed Job Description
As a DevOps / Infrastructure / SRE / Platform Engineer at Systemsltd in Dubai, you will take full ownership of AI infrastructure orchestration, platform engineering scalability, production troubleshooting, root cause analysis (RCA), and observability frameworks. Your day-to-day responsibilities encompass managing Kubernetes clusters, provisioning NVIDIA GPU compute resources, configuring Prometheus and Grafana monitoring stacks, and administering Azure Cloud environments. You will deploy and scale local LLM runtimes such as Ollama and vLLM, automate workflow orchestration via n8n, manage PostgreSQL databases, and maintain enterprise CI/CD deployment pipelines. Working in a fast-paced technology consulting environment in Dubai, you will drive operational excellence, system resilience, and high-performance AI platform reliability.
Key Responsibilities
- Design, architect, deploy, and manage enterprise-grade AI infrastructure and AI platform engineering solutions on Kubernetes.
- Provision, configure, optimize, and maintain NVIDIA GPU compute clusters and specialized hardware accelerators for heavy ML workloads.
- Build, scale, and troubleshoot LLM and machine learning infrastructure pipelines, ensuring low-latency model serving.
- Configure, monitor, and maintain comprehensive observability platforms utilizing Prometheus and Grafana for proactive alerting and metrics tracking.
- Administer enterprise cloud infrastructure on Azure Cloud, managing virtual networks, IAM policies, and storage tiers.
- Deploy, manage, and scale local LLM runtimes and inference servers including Ollama, vLLM, and related serving engines.
- Automate workflow integrations and operational tasks utilizing n8n and advanced infrastructure-as-code (IaC) tooling.
- Manage, tune, and maintain PostgreSQL databases supporting high-throughput platform applications.
- Spearhead continuous integration and continuous delivery (CI/CD) deployment pipelines and release automation.
- Conduct rigorous production troubleshooting, incident response, and root cause analysis (RCA) to maintain high system availability.
Required Qualifications & Skills
- Bachelor’s or Master’s degree in Computer Science, Software Engineering, Information Technology, or a related technical discipline.
- Minimum 5+ years of professional DevOps, site reliability engineering (SRE), infrastructure engineering, or platform engineering experience.
- Strong technical proficiency and hands-on operational mastery of Kubernetes container orchestration and cloud DevOps methodologies.
- Extensive practical experience managing NVIDIA GPU hardware infrastructure and machine learning/LLM operational pipelines.
- Proven expertise in monitoring and observability stacks utilizing Prometheus and Grafana.
- Strong operational mastery of Azure Cloud environments, PostgreSQL database administration, and Linux system engineering.
- Practical experience deploying and managing local LLM runtimes (Ollama, vLLM) and workflow automation tools (n8n).
- Exceptional production troubleshooting capabilities, incident management, and root cause analysis (RCA) expertise.
- Professional availability for full-time onsite employment in Dubai, UAE (Application email: heena.shaikh@systemsltd.com).
Nice-to-Have Skills
- Professional certifications such as Certified Kubernetes Administrator (CKA), Certified Kubernetes Security Specialist (CKS), or Microsoft Certified: Azure DevOps Engineer Expert.
- Advanced experience with infrastructure-as-code tools (Terraform, Ansible) and container registry hardening.
- Prior working experience in tier-1 artificial intelligence enterprises, cloud service providers, or global IT consulting firms in the UAE/GCC region.
- Advanced scripting proficiency in Python, Bash, or Go for infrastructure automation.
- Bilingual proficiency in English and Arabic.
Application Information
- Recruiter: Systemsltd Talent Acquisition Practice
- Contact Name: Heena Shaikh - CHRAM - Global Talent Acquisition Lead
- Email: heena.shaikh@systemsltd.com
- Phone: Unspecified
- Application URL: Unspecified
- Salary/Rate: Competitive and commensurate with experience
- Deadline: Open until filled
- Notice Period: Immediate to short notice preferred
- Contract Duration: Permanent / Full-Time Onsite (Dubai, UAE)
Recruitment Pro Tip
When applying for platform engineering and AI infrastructure roles in Dubai, ensure your resume explicitly highlights your Kubernetes cluster scale, NVIDIA GPU provisioning depth, and production RCA track record.