Manager DevOps

Fatima Group

Lahore

On-site

PKR 24,965,000 - 41,609,000

Full time

3 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Fatima Group seeks a hands-on Manager DevOps to lead our DevOps, Cloud, Platform Engineering, and MLOps operations. You will ensure AI workloads move from development to secure, observable, and production-ready environments while owning deployment automation and platform reliability.

You will collaborate with Software Engineering, AI/ML, Data, Security, and Architecture teams to implement secure, scalable, and automated systems across on-prem and cloud environments.

Qualifications

  • Hands-on leadership of DevOps, cloud, platform engineering, and ML ops in production environments.
  • Experience building and running CI/CD pipelines across multi-cloud and on-prem platforms.
  • StrongLinux administration, networking, and security controls implementation.
  • Hands-on experience with container platforms (Docker, Kubernetes) and IaC tooling.

Responsibilities

  • Lead day-to-day DevOps, Cloud, Platform Engineering, and production operations.
  • Design and improve CI/CD pipelines across multiple tools and cloud platforms.
  • Own containerized platforms, deployment scaling, and reliability.
  • Manage cloud infrastructure mainly on AWS with Azure/GCP familiarity.
  • Implement IaC using Terraform, Ansible, Helm, CloudFormation.
  • Maintain observability with CloudWatch, Prometheus, Grafana, OpenTelemetry.

Skills

CI/CD pipelines
Cloud engineering
DevOps leadership
Linux administration
Automation scripting
Monitoring & observability
Security practices
MLOps familiarity
Networking fundamentals
Platform reliability

Education

Bachelor's or Master's degree in CS/SE/IT
AWS certifications (CloudOps/SysOps/DevOps/Security)

Tools

Docker
Kubernetes
GitLab CI/CD
Jenkins
GitHub Actions
Azure DevOps
Terraform
Ansible
Helm
CloudFormation
OpenTelemetry

Job description

We are looking for a hands-on Manager DevOps to lead and strengthen our DevOps, Cloud, Platform Engineering, and MLOps operations. The role is responsible for ensuring that software, data, AI/ML, and Generative AI workloads can move reliably from development to secure, observable, and production-ready environments. The ideal candidate is a strong DevOps and Cloud engineering leader who understands how modern AI/ML systems are productionized and operated. This person will own deployment automation, cloud and on-premises platform operations, containerized workloads, infrastructure reliability, observability, security controls, backup and recovery, and the operational foundations required for MLOps and AI platform delivery. The role remains technically hands-on and works closely with Software Engineering, AI/ML, Data, Security, Infrastructure, Product, and Architecture teams

Key Job Responsibilities
  • Lead day-to-day DevOps, Cloud, Platform Engineering, and production operations across development, staging, and production environments.
  • Design, implement, maintain, and continuously improve CI/CD pipelines using platforms such as GitLab CI/CD, Jenkins, GitHub Actions, Azure DevOps, or equivalent tools.
  • Own containerized application platforms using Docker and Kubernetes, including workload deployment, configuration, scaling, availability, resource management, and troubleshooting.
  • Manage and optimize cloud infrastructure, primarily on AWS, with practical familiarity across Azure and GCP services where required.
  • Establish Infrastructure as Code and configuration automation using Terraform, Ansible, CloudFormation, Helm, or equivalent technologies to make environments repeatable and auditable.
  • Operate Linux-based infrastructure and production services, including networking, reverse proxies, web servers, systemd services, DNS, SSL/TLS, VPN connectivity, firewalls, load balancing, and secure remote access.
  • Build and maintain reliable observability across applications, infrastructure, and AI workloads using tools such as CloudWatch, Prometheus, Grafana, Loki, ELK/OpenSearch, OpenTelemetry, or equivalent platforms.
  • Establish alerting, incident response, root cause analysis, post-incident reviews, operational runbooks, and preventive actions to improve platform reliability and reduce recurring failures.
  • Own backup, restore, disaster recovery, business continuity, capacity planning, patching, certificate management, and operational readiness for critical systems.
  • Apply DevSecOps practices across CI/CD and infrastructure, including IAM, secrets management, least-privilege access, vulnerability remediation, secure configuration, key management, auditability, and coordination with security teams.
  • Support the productionization and operation of AI/ML workloads, including model-serving environments, batch and real-time inference services, training jobs, model artifacts, data pipelines, background workers, and GPU or CPU intensive workloads.
  • Build and support MLOps and LLMOps workflows using platforms such as AWS SageMaker, MLflow, Kubeflow, Azure Machine Learning, Vertex AI, or equivalent tools, based on project needs.
  • Enable reliable deployment and monitoring of AI-enabled applications, including Generative AI, RAG, vector database, embedding, agentic, and API-based model services without taking ownership of model research or algorithm development.
  • Work closely with AI/ML Engineers and Data teams to automate model packaging, deployment, versioning, environment promotion, rollback, observability, and retraining or refresh workflows where applicable.
  • Operate and troubleshoot supporting platform components such as PostgreSQL, TimescaleDB, Redis, RabbitMQ, Celery, message queues, caches, and related data services from an infrastructure and reliability perspective.
  • Manage source-control and deployment governance, including repository standards, branch protections, runners, environment controls, release approvals, rollback procedures, and secure CI/CD variables.
  • Drive cloud and infrastructure cost optimization through right-sizing, utilization monitoring, life cycle management, storage optimization, automation, and appropriate use of managed services.
  • Maintain accurate infrastructure diagrams, dependency maps, access and ownership records, runbooks, recovery procedures, environment documentation, and operational knowledge required for continuity.
  • Coordinate with internal Infrastructure, Network, Security, application teams, vendors, and client stakeholders to resolve dependencies and ensure stable service delivery.
  • Mentor DevOps and platform engineers, establish engineering standards, review technical implementations, and promote a culture of automation, ownership, reliability, security, and continuous learning.
Detailed Description
  • Strong hands-on experience in DevOps, Cloud Engineering, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering in production environments.
  • Strong Linux administration and troubleshooting skills, including networking fundamentals, DNS, HTTP/HTTPS, TLS, SSH, reverse proxies, firewalls, VPNs, load balancers, and system services.
  • Hands-on experience with AWS infrastructure and services such as EC2, VPC, IAM, S3, RDS, CloudWatch, ECR, ECS/EKS, Lambda, Route 53, or equivalent production services.
  • Strong experience with Docker and practical production experience with Kubernetes, including EKS, AKS, GKE, or on-premises Kubernetes environments.
  • Strong CI/CD experience with GitLab CI/CD, Jenkins, GitHub Actions, Azure DevOps, or similar platforms.
  • Hands-on experience with Terraform or equivalent Infrastructure as Code tooling. Experience with Ansible, Helm, CloudFormation, or similar configuration and deployment automation is strongly valued.
  • Strong understanding of monitoring, logging, alerting, observability, availability, performance, incident response, backup, restore, and disaster recovery practices.
  • Practical experience with cloud security, IAM, secrets management, network security, vulnerability management, access control, and DevSecOps practices.
  • Working knowledge of relational databases, time-series databases, caches, message brokers, and background job platforms such as PostgreSQL, TimescaleDB, Redis, RabbitMQ, or Celery.
  • Proficiency in automation and scripting using Bash, Python, PowerShell, or a comparable language.
  • Practical understanding of MLOps and the production life cycle of machine learning systems, including model deployment, model serving, training or inference infrastructure, model artifacts, monitoring, and release automation.
  • Experience supporting AI/ML workloads on cloud or container platforms. Exposure to AWS SageMaker, MLflow, Kubeflow, Azure Machine Learning, Vertex AI, or equivalent tooling is expected.
  • Ability to troubleshoot complex production issues across infrastructure, applications, networking, CI/CD, databases, containers, and AI/ML services.
  • Strong technical leadership, documentation, communication, prioritization, and stakeholder-management skills, with the ability to remain hands-on while leading operational delivery.
  • Active AWS certifications aligned to Cloud Operations and DevOps are mandatory for this role. AWS Certified CloudOps Engineer - Associate, or an active legacy AWS Certified SysOps Administrator - Associate credential, is required.
  • AWS Certified DevOps Engineer - Professional is mandatory.
  • AWS Certified Security - Specialty is mandatory as the AWS Specialty certification for this role.
Requirements
  • Experience operating AI/ML, Generative AI, RAG, LLM, vector database, or agentic application workloads in production.
  • Experience with GPU-enabled workloads, NVIDIA infrastructure, inference servers such as NVIDIA Triton or vLLM, or GPU scheduling on Kubernetes or cloud platforms.
  • Experience with GitOps and deployment platforms such as Argo CD, Flux, Tekton, or equivalent tools.
  • Exposure to data and workflow orchestration platforms such as Apache Airflow, Spark, Kafka, or similar technologies.
  • Experience operating hybrid environments that combine on-premises infrastructure with public cloud services.
  • Experience supporting industrial, manufacturing, IoT, time-series, or high-availability application environments is an advantage.
  • Additional certifications such as CKA/CKAD, HashiCorp Terraform Associate, Azure DevOps Engineer, Google Professional Cloud DevOps Engineer, or equivalent platform certifications are preferred.
  • Experience improving an existing DevOps environment through automation, standardization, stronger observability, Infrastructure as Code, security controls, and operational governance.
Education & Experience
  • Bachelor's or Master's degree in Computer Science, Software Engineering, Information Technology, Computer Engineering, or a related technical discipline.
  • Mandatory AWS certification baseline: AWS Certified CloudOps Engineer - Associate, or active legacy AWS Certified SysOps Administrator - Associate, AWS Certified DevOps Engineer - Professional, and AWS Certified Security - Specialty.
  • 7 to 10+ years of relevant experience across DevOps, Cloud, Platform Engineering, SRE, Infrastructure Engineering, or closely related roles.
  • At least 2 to 4+ years of experience in a technical lead, team lead, or management capacity with responsibility for production systems and engineering delivery.
  • Demonstrated experience operating business-critical production environments and leading incident resolution, platform improvement, and cross-functional technical coordination.
  • Hands-on exposure to MLOps, AI platform operations, or production AI/ML workloads is required. Deep model-development or data-science experience is not required.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Lead
AI/ML Lead

Digifloat • Lahore

On-site
PKR 3,600,000 - 7,200,000
Infrastructure & DevSecOps Lead
Infrastructure & DevSecOps Lead

KnowledgeCity • Pakistan

On-site
PKR 1,800,000 - 3,200,000
Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Islamabad

On-site
PKR 4,000,000 - 7,000,000
Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Karachi Division

On-site
PKR 2,500,000 - 4,200,000
Head of AI/ML Engineering
Head of AI/ML Engineering

Fatima Group • Lahore

On-site
PKR 8,000,000 - 16,000,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Abacus Global • Lahore

On-site
PKR 1,500,000 - 2,500,000
Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Islamabad

On-site
PKR 2,400,000 - 4,200,000
Infrastructure & DevSecOps Lead Pakistan
Infrastructure & DevSecOps Lead Pakistan

KnowledgeCity • Pakistan

On-site
PKR 3,500,000 - 5,000,000
Senior Software Engineer - Data Engineering & AI
Senior Software Engineer - Data Engineering & AI

Devsinc, LLC • Islamabad

On-site
PKR 1,674,000 - 3,348,000
AI Platform Operations Engineer
AI Platform Operations Engineer

Datamatics Technologies • Karachi Division

On-site
PKR 2,000,000 - 4,000,000