We are looking for a hands-on Manager DevOps to lead and strengthen our DevOps, Cloud, Platform Engineering, and MLOps operations. The role is responsible for ensuring that software, data, AI/ML, and Generative AI workloads can move reliably from development to secure, observable, and production-ready environments. The ideal candidate is a strong DevOps and Cloud engineering leader who understands how modern AI/ML systems are productionized and operated. This person will own deployment automation, cloud and on-premises platform operations, containerized workloads, infrastructure reliability, observability, security controls, backup and recovery, and the operational foundations required for MLOps and AI platform delivery. The role remains technically hands-on and works closely with Software Engineering, AI/ML, Data, Security, Infrastructure, Product, and Architecture teams
Key Job Responsibilities
- Lead day-to-day DevOps, Cloud, Platform Engineering, and production operations across development, staging, and production environments.
- Design, implement, maintain, and continuously improve CI/CD pipelines using platforms such as GitLab CI/CD, Jenkins, GitHub Actions, Azure DevOps, or equivalent tools.
- Own containerized application platforms using Docker and Kubernetes, including workload deployment, configuration, scaling, availability, resource management, and troubleshooting.
- Manage and optimize cloud infrastructure, primarily on AWS, with practical familiarity across Azure and GCP services where required.
- Establish Infrastructure as Code and configuration automation using Terraform, Ansible, CloudFormation, Helm, or equivalent technologies to make environments repeatable and auditable.
- Operate Linux-based infrastructure and production services, including networking, reverse proxies, web servers, systemd services, DNS, SSL/TLS, VPN connectivity, firewalls, load balancing, and secure remote access.
- Build and maintain reliable observability across applications, infrastructure, and AI workloads using tools such as CloudWatch, Prometheus, Grafana, Loki, ELK/OpenSearch, OpenTelemetry, or equivalent platforms.
- Establish alerting, incident response, root cause analysis, post-incident reviews, operational runbooks, and preventive actions to improve platform reliability and reduce recurring failures.
- Own backup, restore, disaster recovery, business continuity, capacity planning, patching, certificate management, and operational readiness for critical systems.
- Apply DevSecOps practices across CI/CD and infrastructure, including IAM, secrets management, least-privilege access, vulnerability remediation, secure configuration, key management, auditability, and coordination with security teams.
- Support the productionization and operation of AI/ML workloads, including model-serving environments, batch and real-time inference services, training jobs, model artifacts, data pipelines, background workers, and GPU or CPU intensive workloads.
- Build and support MLOps and LLMOps workflows using platforms such as AWS SageMaker, MLflow, Kubeflow, Azure Machine Learning, Vertex AI, or equivalent tools, based on project needs.
- Enable reliable deployment and monitoring of AI-enabled applications, including Generative AI, RAG, vector database, embedding, agentic, and API-based model services without taking ownership of model research or algorithm development.
- Work closely with AI/ML Engineers and Data teams to automate model packaging, deployment, versioning, environment promotion, rollback, observability, and retraining or refresh workflows where applicable.
- Operate and troubleshoot supporting platform components such as PostgreSQL, TimescaleDB, Redis, RabbitMQ, Celery, message queues, caches, and related data services from an infrastructure and reliability perspective.
- Manage source-control and deployment governance, including repository standards, branch protections, runners, environment controls, release approvals, rollback procedures, and secure CI/CD variables.
- Drive cloud and infrastructure cost optimization through right-sizing, utilization monitoring, life cycle management, storage optimization, automation, and appropriate use of managed services.
- Maintain accurate infrastructure diagrams, dependency maps, access and ownership records, runbooks, recovery procedures, environment documentation, and operational knowledge required for continuity.
- Coordinate with internal Infrastructure, Network, Security, application teams, vendors, and client stakeholders to resolve dependencies and ensure stable service delivery.
- Mentor DevOps and platform engineers, establish engineering standards, review technical implementations, and promote a culture of automation, ownership, reliability, security, and continuous learning.
Detailed Description
- Strong hands-on experience in DevOps, Cloud Engineering, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering in production environments.
- Strong Linux administration and troubleshooting skills, including networking fundamentals, DNS, HTTP/HTTPS, TLS, SSH, reverse proxies, firewalls, VPNs, load balancers, and system services.
- Hands-on experience with AWS infrastructure and services such as EC2, VPC, IAM, S3, RDS, CloudWatch, ECR, ECS/EKS, Lambda, Route 53, or equivalent production services.
- Strong experience with Docker and practical production experience with Kubernetes, including EKS, AKS, GKE, or on-premises Kubernetes environments.
- Strong CI/CD experience with GitLab CI/CD, Jenkins, GitHub Actions, Azure DevOps, or similar platforms.
- Hands-on experience with Terraform or equivalent Infrastructure as Code tooling. Experience with Ansible, Helm, CloudFormation, or similar configuration and deployment automation is strongly valued.
- Strong understanding of monitoring, logging, alerting, observability, availability, performance, incident response, backup, restore, and disaster recovery practices.
- Practical experience with cloud security, IAM, secrets management, network security, vulnerability management, access control, and DevSecOps practices.
- Working knowledge of relational databases, time-series databases, caches, message brokers, and background job platforms such as PostgreSQL, TimescaleDB, Redis, RabbitMQ, or Celery.
- Proficiency in automation and scripting using Bash, Python, PowerShell, or a comparable language.
- Practical understanding of MLOps and the production life cycle of machine learning systems, including model deployment, model serving, training or inference infrastructure, model artifacts, monitoring, and release automation.
- Experience supporting AI/ML workloads on cloud or container platforms. Exposure to AWS SageMaker, MLflow, Kubeflow, Azure Machine Learning, Vertex AI, or equivalent tooling is expected.
- Ability to troubleshoot complex production issues across infrastructure, applications, networking, CI/CD, databases, containers, and AI/ML services.
- Strong technical leadership, documentation, communication, prioritization, and stakeholder-management skills, with the ability to remain hands-on while leading operational delivery.
- Active AWS certifications aligned to Cloud Operations and DevOps are mandatory for this role. AWS Certified CloudOps Engineer - Associate, or an active legacy AWS Certified SysOps Administrator - Associate credential, is required.
- AWS Certified DevOps Engineer - Professional is mandatory.
- AWS Certified Security - Specialty is mandatory as the AWS Specialty certification for this role.
Requirements
- Experience operating AI/ML, Generative AI, RAG, LLM, vector database, or agentic application workloads in production.
- Experience with GPU-enabled workloads, NVIDIA infrastructure, inference servers such as NVIDIA Triton or vLLM, or GPU scheduling on Kubernetes or cloud platforms.
- Experience with GitOps and deployment platforms such as Argo CD, Flux, Tekton, or equivalent tools.
- Exposure to data and workflow orchestration platforms such as Apache Airflow, Spark, Kafka, or similar technologies.
- Experience operating hybrid environments that combine on-premises infrastructure with public cloud services.
- Experience supporting industrial, manufacturing, IoT, time-series, or high-availability application environments is an advantage.
- Additional certifications such as CKA/CKAD, HashiCorp Terraform Associate, Azure DevOps Engineer, Google Professional Cloud DevOps Engineer, or equivalent platform certifications are preferred.
- Experience improving an existing DevOps environment through automation, standardization, stronger observability, Infrastructure as Code, security controls, and operational governance.
Education & Experience
- Bachelor's or Master's degree in Computer Science, Software Engineering, Information Technology, Computer Engineering, or a related technical discipline.
- Mandatory AWS certification baseline: AWS Certified CloudOps Engineer - Associate, or active legacy AWS Certified SysOps Administrator - Associate, AWS Certified DevOps Engineer - Professional, and AWS Certified Security - Specialty.
- 7 to 10+ years of relevant experience across DevOps, Cloud, Platform Engineering, SRE, Infrastructure Engineering, or closely related roles.
- At least 2 to 4+ years of experience in a technical lead, team lead, or management capacity with responsibility for production systems and engineering delivery.
- Demonstrated experience operating business-critical production environments and leading incident resolution, platform improvement, and cross-functional technical coordination.
- Hands-on exposure to MLOps, AI platform operations, or production AI/ML workloads is required. Deep model-development or data-science experience is not required.