About The Role
The DevOps Engineer will build and operate the infrastructure, deployment systems, and observability platform that keep production services reliable at scale. The role spans AWS infrastructure, Kubernetes orchestration, infrastructure as code, CI/CD automation, and operational tooling for distributed applications.
The engineer will partner with software engineers and SREs to improve release velocity without compromising availability or security. The team needs someone who can automate repeatable work, troubleshoot complex production incidents, and turn operational lessons into more resilient platforms and processes.
Key Responsibilities
- Design and manage highly available AWS infrastructure using Terraform, including networking, IAM, compute, storage, and managed database services
- Build and maintain CI/CD pipelines with GitHub Actions, Jenkins, or GitLab CI to automate testing, security checks, deployments, and rollback workflows
- Operate Kubernetes environments using Helm and related tooling; improve cluster capacity, workload scheduling, service reliability, and deployment standards
- Implement observability with Prometheus, Grafana, and centralized logging platforms such as ELK or OpenSearch; define actionable metrics, dashboards, and alerts
- Automate infrastructure provisioning, configuration management, and routine operational tasks using Python, Bash, or Go
- Participate in on-call rotations, lead incident response, and document root-cause analyses with specific remediation plans
- Harden production environments through least-privilege IAM, secrets management, vulnerability remediation, backup validation, and disaster-recovery testing
What We Are Looking For
- 3–8 years of experience in DevOps, SRE, platform engineering, or a closely related infrastructure role supporting production systems
- Hands-on experience with AWS services such as EC2, EKS, VPC, IAM, S3, RDS, CloudWatch, and Route 53
- Strong proficiency with Terraform and Git-based infrastructure and application delivery workflows
- Production experience operating Kubernetes, including deployments, services, ingress, autoscaling, resource management, and troubleshooting
- Working knowledge of Linux systems, networking fundamentals, containers, and scripting with Python or Bash
- Experience building observability and incident-response practices around metrics, logs, traces, alerting, SLIs, and SLOs
- Bonus: experience with Go, Argo CD, Ansible, service meshes, FinOps, compliance automation, or formal education in computer science, engineering, or a related technical field