Job Title: Senior AI Solutions Engineer (AI Infrastructure, MLOps & Generative AI Platforms) | Quantum Talent Group | Abu Dhabi, UAE
Recruiting Company: Quantum Talent Group
Job Location: Abu Dhabi, United Arab Emirates
Job Type: Full-Time
Application Method
- Location: Abu Dhabi, UAE
- Customer-Facing Solution Design & Delivery Role
- Experience Required: 8+ Years Infrastructure / Platform / DevOps Engineering
- AI/ML Experience Required: 2+ Years
- Enterprise AI Platform Engineering Opportunity
- Immediate Hiring Requirement
Position Summary
Quantum Talent Group is seeking a highly experienced Senior AI Solutions Engineer to design, architect, and deliver enterprise-scale AI platforms supporting advanced AI, Machine Learning, and Generative AI workloads. This customer-facing role requires a strong blend of AI infrastructure expertise, cloud-native platform engineering, Kubernetes administration, GPU computing, and solution architecture capabilities.
Detailed Job Description
As a Senior AI Solutions Engineer, you will lead the design and implementation of modern AI infrastructure platforms that enable enterprise AI, MLOps, Generative AI, and Agentic AI initiatives. You will work closely with customers, architects, data scientists, and engineering teams to design scalable solutions supporting Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and GPU-accelerated workloads. The role demands deep expertise in Kubernetes, OpenShift, NVIDIA GPU infrastructure, automation frameworks, and AI platform operations. You will facilitate workshops, gather requirements, translate business challenges into technical architectures, and oversee end-to-end implementation delivery. This is an excellent opportunity to work at the forefront of enterprise AI transformation and next-generation intelligent computing platforms.
Key Responsibilities
- Design and deliver enterprise AI infrastructure platforms supporting machine learning and Generative AI workloads.
- Architect scalable Kubernetes and OpenShift environments for AI, MLOps, and GPU-intensive applications.
- Design and implement NVIDIA GPU-enabled platforms for LLM training, inference, and AI workloads.
- Lead customer workshops, solution discovery sessions, and technical architecture discussions.
- Build and optimize infrastructure supporting Retrieval-Augmented Generation (RAG) and Agentic AI solutions.
- Implement infrastructure automation using Terraform, Ansible, GitOps, and Infrastructure as Code practices.
- Develop Python-based automation scripts and platform management solutions.
- Design highly available, scalable, and secure AI platform architectures.
- Support GPU resource allocation, scheduling, sharing, and cluster optimization activities.
- Collaborate with AI engineers, data scientists, cloud architects, and DevOps teams.
- Provide technical leadership throughout solution delivery from design to production deployment.
- Ensure platform observability, operational readiness, security, governance, and performance optimization.
Required Qualifications & Skills
- Bachelor’s Degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
- Minimum 8 years of experience in Infrastructure Engineering, Platform Engineering, DevOps Engineering, or related disciplines.
- Minimum 2 years of hands-on experience supporting AI/ML workloads or GPU-accelerated computing environments.
- Strong production experience with Kubernetes and/or OpenShift deployments.
- Expertise managing Kubernetes/OpenShift on bare-metal and virtualized infrastructure.
- Extensive experience with NVIDIA GPU infrastructure, CUDA, GPU scheduling, and GPU sharing technologies.
- Strong Infrastructure as Code (IaC) expertise using Terraform and Ansible.
- Experience implementing GitOps methodologies and automation frameworks.
- Advanced Python scripting and automation experience.
- Hands-on experience supporting Large Language Model (LLM) serving environments.
- Experience designing and supporting Retrieval-Augmented Generation (RAG) architectures.
- Strong solution architecture, stakeholder engagement, and customer workshop facilitation skills.
- Excellent troubleshooting, communication, and technical leadership capabilities.
Nice-to-Have Skills
- Experience with MLOps platforms and AI model lifecycle management.
- Experience deploying large-scale model serving and inference platforms.
- Knowledge of high-performance storage architectures for AI workloads.
- Experience with InfiniBand, RoCE, and high-speed networking technologies.
- Familiarity with AI observability, monitoring, and performance optimization frameworks.
- Experience supporting Agentic AI, Generative AI, and enterprise AI transformation programs.
Recruitment Pro Tip
For AI infrastructure and platform engineering roles, ensure your CV prominently highlights Kubernetes, OpenShift, NVIDIA GPUs, CUDA, AI Infrastructure, MLOps, Terraform, Ansible, GitOps, Python, LLM Serving, RAG Architectures, Agentic AI, Model Deployment, GPU Clusters, High-Performance Computing (HPC), and Enterprise AI Platforms. Quantifiable achievements such as reduced inference latency, improved GPU utilization, accelerated AI deployments, or successful enterprise AI platform implementations will significantly strengthen your application.