DevOps / IT Infra Engineer

Evollabs Tech

Dubai

On-site

AED 480,000 - 720,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Evollabs Tech seeks a Platform Engineer to own end-to-end on-prem AI infra, architecting and operating Kubernetes GPU clusters, and enabling reliable CI/CD. You will manage servers, networks, and core services, and drive observability, security, and disaster recovery.

You will build reusable IaC modules, implement GitOps, and collaborate with developers to accelerate AI workloads while maintaining strict data residency and air-gapped controls.

Qualifications

  • 5+ years in DevOps, SRE, or Platform Engineering with hands-on ownership of on-prem environments.
  • Advanced troubleshooting in Linux networking, storage I/O, and kernel parameters.
  • Hands-on experience with physical server racks, network switches, cabling, and hardware lifecycle management.
  • Experience with GPU server management, drivers, firmware, and thermal/power considerations.
  • Experience designing and operating air-gapped/isolated networks with strict access controls (FreeIPA/LDAP, Kerberos, SSSD).
  • Production observability setups with Prometheus and Grafana, dashboards, and alerting.
  • Experience with Proxmox VE clusters, hypervisors, VM storage, and virtual networking.
  • Proficiency with Docker and Kubernetes for containerized workloads and cluster ops.
  • Python and Bash scripting for automation and API integrations.
  • CI/CD pipelines across GitLab CI, GitHub Actions, Jenkins, etc.
  • Experience integrating AI tooling and LLM endpoints (Cursor/Claude/Copilot) and cost tracking for AI infra.
  • Knowledge of auditing, reproducible environments, and runbooks for engineers.

Responsibilities

  • Design and operate on-prem infrastructure as code: Terraform/Ansible/Helm modules; GitOps workflows for auditable changes.
  • Build and run Kubernetes for AI: multi-tenant GPU clusters, quotas, and workload isolation.
  • Administer servers, networks, and core services: Linux, identity/SSO (Keycloak/LDAP), secrets (Vault), DNS/DHCP/NTP, artifact registries, mirrors.
  • Enable CI/CD: collaborate with developers to create fast, reproducible pipelines and to deploy to GPU/CPU nodes.
  • Observability & reliability: metrics/logs/traces with Prometheus/Grafana, define SLOs, incident response, post-mortems, and runbooks.
  • Backup & disaster recovery: define RPO/RTO, backups, restores, and site failover tests.

Skills

DevOps
SRE
Platform Engineering
Linux
On-prem Hardware
GPU Hardware
Proxmox
Kubernetes
CI/CD
Terraform
Ansible
Helm
GitLab CI
NVIDIA CUDA
Vault
Keycloak/LDAP
Prometheus/Grafana

Tools

Proxmox VE
Terraform
Ansible
Helm
Kubernetes
Docker
GitLab CI
GitHub Actions
Jenkins
NVIDIA CUDA

Job description

About Us

We are a young high-tech company incorporated in the heart of one of the world's fastest growing tech hubs - Dubai, UAE. As the exclusive software partner to one of the world's largest ODMs in the networking equipment space, we develop the Network Operating Systems that power critical data centre and telecom routing & switching infrastructure. Building on this foundation, we've recently launched an AI division focused on designing our own chips to accelerate inference and training workloads.

About Us

We are a young high-tech company incorporated in the heart of one of the world's fastest growing tech hubs - Dubai, UAE. As the exclusive software partner to one of the world's largest ODMs in the networking equipment space, we develop the Network Operating Systems that power critical data centre and telecom routing & switching infrastructure. Building on this foundation, we've recently launched an AI division focused on designing our own chips to accelerate inference and training workloads.

Your Mission

Own the end-to-end design and operation of our on-premise infrastructure for AI and enterprise workloads - built as code, automated, observable, and secure. You will architect and run Kubernetes clusters for training/inference, manage servers, networks, and core services, and enable developers with reliable CI/CD and platform tooling. This is where minutes, time-to-recovery and cost-per-job directly impact AI velocity at scale.

Responsibilities
  • Design and operate on-prem infrastructure as code: author reusable Terraform/Ansible/Helm modules; build GitOps workflows for repeatable, audited changes across environments.
  • Build and run Kubernetes for AI: configure multi-tenant GPU clusters (MIG/GPUDirect RDMA, NVIDIA device plugins/DCGM), scheduling/quotas, and workload isolation.
  • Administer servers, networks, and core services: OS lifecycle (Linux), identity/SSO (Keycloak/LDAP), secrets (Vault), DNS/DHCP/NTP, artifact registries, and internal package mirrors.
  • Enable CI/CD: partner with developers to design fast, reproducible pipelines (GitLab CI), caching, and artifact provenance, and to deploy to GPU/CPU nodes.
  • Observability & reliability: build metrics/logs/traces (Prometheus/Grafana), define SLOs/error budgets, incident response/on-call, post-mortems, and runbooks.
  • Backup & disaster recovery: formulate RPO/RTO targets, implement cluster/app backups, test restores and site failover procedures regularly.
  • Document clearly: architecture diagrams, playbooks, and self-service guides for engineers and stakeholders.
Must-Have
  • Overall Experience: 5+ years in DevOps, SRE, or Platform Engineering with direct hands-on ownership and operation of physical, on-premises environments.
  • Linux Systems Engineering: Advanced troubleshooting across kernel parameters, storage I/O, networking stacks (TCP/IP, VLANs, DNS), and performance tuning.
  • On-prem Hardware & Datacenter Operations: Hands-on experience managing physical server racks, network switches, patch cabling (Fiber, SFP+/QSFP, Cat6), and hardware lifecycle management.
  • GPU Hardware & Firmware: Practical experience with server GPU installation, driver management (e.g., NVIDIA CUDA), firmware updates, thermal/power considerations, and hardware-level troubleshooting.
  • Air-gapped Environments & Access Management: Experience designing, operating, and securing air-gapped/isolated networks, enforcing strict access controls alongside FreeIPA (LDAP, Kerberos, SSSD, DNS, and RBAC). Securing inference engines and enforcing strict data residency controls to ensure sensitive corporate data, code, and prompts never leave on-prem or air-gapped boundaries.
  • Observability & Monitoring: Production setup and maintenance of Prometheus and Grafana for metrics collection, system monitoring, dynamic dashboards, and alerting.
  • Proxmox Virtualization: Hands-on experience configuring and operating Proxmox VE clusters, hypervisors, VM storage backends, and virtual networking.
  • Containers & Orchestration: Production experience with Docker and Kubernetes (K8s/K3s) for containerizing workloads, managing pods, and maintaining cluster operations.
  • Scripting & Core Automation: High proficiency in Python and Bash for writing custom administrative scripts, automation tools, and API integrations.
  • CI/CD Pipelines: Experience building and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, etc.) to automate code deployments and operational tasks.
  • Developer AI & vLLM Integration: Experience configuring, supporting, and promoting Cursor/Claude/Copilot and AI-assisted coding tools for development teams.
  • Token Observability & Cost Tracking: Experience setting up proxies, rate-limiting, telemetry, and cost/usage monitoring for LLM endpoints (e.g., via LiteLLM, OpenRouter, or custom API gateways).
  • Agentic Workflows for DevOps: Experience designing, deploying, or utilizing autonomous agent systems (e.g., LangGraph, CrewAI, or custom Python agents) to automate operational workflows like log analysis, issue triage, or system alerts.
Nice-to-Have
  • Infrastructure as Code (IaC): Knowledge of Ansible, Terraform, or OpenTofu to help transition the team toward automated infrastructure provisioning in the future.
  • Local LLM Infrastructure: Experience running self-hosted inference engines (vLLM, Ollama) on high-end HPC machines and setting up GPU passthrough (PCIe/vGPU) within Proxmox VMs or containers.
  • Storage Systems: Enterprise experience configuring and managing Ceph, ZFS, NFS, or SAN/NAS storage topologies.
  • Secrets Management: Experience with secrets management solutions (e.g., HashiCorp Vault) for safe handling of API keys, tokens, and system credentials in isolated setups.
  • Internal AI Systems: Basic understanding of vector databases (Qdrant, Pgvector, Milvus) for building retrieval-augmented generation tools over internal infrastructure docs and build AI tooling securely for day to day task automation.
What Success Really Means
  • Reproducible environments by default: any engineer can spin up an identical dev/test stack (K8s namespace, storage, secrets, runners) from Git in ≤30 minutes, with audit trails for every change.
  • Solid CI/CD for AI workflows: model/build/test pipelines are deterministic and cache-efficient; median pipeline time down 30-50%, with artifact provenance (SBOM, signatures) and traceable datasets/checkpoints.
  • Lab-to-cluster continuity: hardware bring-up images, drivers, and firmware are versioned and promoted through the same pipelines; new boards/nodes join clusters with push-button automation.
  • Actionable observability: dashboards and alerts reflect SLOs meaningful to researchers (throughput, time-to-first-token, I/O wait, GPU mem pressure);
  • Cost & toil reduction: infra tasks automated to eliminate recurring manual work; fewer \"custom one-offs,\" more reusable modules; quarterly infra spend per GPU hour trends down.
  • Clear docs & self-service: engineers rely on concise runbooks and service catalogs; >80% of routine requests resolved via self-service workflows rather than ad-hoc ops support.

Join us in our mission to democratize AI compute - where your platform engineering turns ideas into production at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

SUNDUS MANAGEMENT CONSULTANCY & STUDIES BUREAUL.L.C • Abu Dhabi

On-site
AED 280,000 - 420,000
AI Engineer
AI Engineer

techcarrot • Dubai

On-site
AED 350,000 - 650,000
AI/ML DevOps Specialist
AI/ML DevOps Specialist

Netision Technology LLP • United Arab Emirates

On-site
AED 320,000 - 520,000
Competitive salary
Cutting-edge AI/ML tech
Career growth opportunities
MLOps Engineer
MLOps Engineer

Inception42 • Abu Dhabi Emirate

On-site
AED 350,000 - 550,000
Senior MLOps & DevOps Engineer
Senior MLOps & DevOps Engineer

Telcovas Solutions & Services • Dubai

On-site
AED 350,000 - 480,000
Artificial Intelligence & Machine Learning Architect SME
Artificial Intelligence & Machine Learning Architect SME

Remotedxb • Dubai

On-site
AED 420,000 - 900,000
Health plan
401K
Paid time off
+3
Artificial Intelligence (AI) & Machine Learning Engineer
Artificial Intelligence (AI) & Machine Learning Engineer

Jaheziya • Abu Dhabi

On-site
AED 300,000 - 520,000
Lead AI Scientist / Head of AI Solutions
Lead AI Scientist / Head of AI Solutions

Recenso • Abu Dhabi

On-site
Opportunity to lead AI innovation
Research-driven environment
Leadership exposure across teams
Forward Deployed Engineer, UAE
Forward Deployed Engineer, UAE

Telnyx • Dubai

On-site
AED 350,000 - 600,000
Forward Deployment Engineer (m/f/d)
Forward Deployment Engineer (m/f/d)

Halian • Abu Dhabi

On-site
AED 300,000 - 600,000