We are looking for an experienced DevOps Engineer with strong expertise in AWS/Azure, cloud infrastructure, CI/CD, containerization, Kubernetes, Infrastructure as Code, automation, monitoring and production operations, with exposure to GenAI/LLM platforms and AI workloads.
The role will focus on building secure and scalable cloud environments, automating application and AI deployments, implementing DevSecOps, GitOps and observability practices, and supporting reliable production operations for GenAI/LLM applications.
Key Responsibilities
- Design, deploy and manage applications and AI workloads across AWS and Azure.
- Manage cloud infrastructure, networking, compute, storage, databases and production environments.
- Work with AWS EC2, ECS/EKS, Lambda, S3, IAM, VPC, CloudWatch and equivalent Azure services such as AKS, App Services, Azure Functions, Storage, Entra ID and Azure Monitor.
- Implement scalable, highly available and secure cloud architectures.
- Design, build and maintain automated CI/CD pipelines for application and GenAI platform deployments.
- Work with Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps / Azure Pipelines and CircleCI.
- Manage source control using Git, GitHub, GitLab or Bitbucket.
- Deploy and manage workloads using Kubernetes, Amazon EKS and Azure AKS.
- Work with Helm, Kubernetes Operators and Kubernetes networking.
- Provision and manage infrastructure using Terraform.
- Work with AWS CloudFormation, Azure Bicep / ARM Templates, Pulumi or CDK.
- Implement reusable Terraform modules and environment-specific configurations.
- Automate server and configuration management using Ansible, Puppet or Chef.
- Implement infrastructure versioning, automated provisioning and environment consistency.
- Implement GitOps-based deployment models using Argo CD or Flux CD.
- Manage Kubernetes configurations using Git-based deployment workflows.
- Implement automated environment promotion and configuration synchronization.
- Work with Argo Rollouts / progressive delivery where applicable.
- Implement infrastructure and application monitoring using Prometheus, Grafana, Datadog, Dynatrace, New Relic or Splunk.
- Configure dashboards, alerts, health checks, SLIs/SLOs and production monitoring.
- Troubleshoot application, infrastructure, network and deployment-related incidents.
- Integrate security into CI/CD pipelines using SonarQube, Snyk, Trivy, Checkov and OWASP ZAP.
- Implement container image, dependency, source-code and Infrastructure-as-Code security scanning.
- Support vulnerability remediation, compliance and security best practices.
- Troubleshoot production incidents across cloud, Kubernetes, networking, application and CI/CD environments.
- Implement high availability, disaster recovery, backup, autoscaling and resilience mechanisms.
- Perform root-cause analysis and implement preventive actions.
- Support performance, reliability and cloud-cost optimization.
Good-to-Have Skills / Development Areas
- GitOps & Deployment: Argo CD, Flux CD, Helm, Kustomize
- Infrastructure Automation: Ansible, Pulumi, AWS CDK, Azure Bicep
- DevSecOps: SonarQube, Snyk, Trivy, Checkov, HashiCorp Vault
- Monitoring & Observability: Datadog, Dynatrace, New Relic, Splunk, OpenTelemetry
- AI / LLMOps: Azure OpenAI, Azure AI Foundry, AWS Bedrock, LangChain, LangGraph, LangSmith, Langfuse, MLflow
- AI Operations: LLM evaluation, AI observability, token/cost monitoring, model performance and accuracy monitoring
- Security & Governance: IAM, RBAC, API security, secrets management, OPA / Kyverno
- Automation & Distributed Systems: Python, Bash, PowerShell, Kafka, RabbitMQ, Redis
- Reliability: AIOps, incident management, autoscaling, high availability and disaster recovery