About the Company
Our client is a technical consulting and engineering services company supporting organizations in the high-tech industry. Their team helps platform and infrastructure groups manage complex cloud environments, execute large-scale migrations, and improve the reliability and efficiency of application deployments.
They work across modern infrastructure and cloud technologies, partnering closely with engineering and operations teams to solve complex technical challenges and deliver reliable, scalable solutions.
About the Role
Our client is looking for a hands‑on DevOps Engineer to help scale, secure, and maintain its Kubernetes‑based production platform. Most of their applications run on Kubernetes, primarily through managed on‑premises clusters and Amazon EKS, with this role focused on supporting system reliability, operational excellence, and continuous improvement across the cloud‑native environment.
The DevOps Engineer will work closely with our client's Engineering and Operations teams to operate production workloads, improve automation, strengthen security practices, and enhance CI/CD pipelines while taking end‑to‑end ownership of production infrastructure.
What You'll Be Doing
Kubernetes & Container Platforms
- Manage and support production workloads hosted on Kubernetes, including managed on‑premises clusters and Amazon EKS.
- Deploy, update, and administer applications through Helm or comparable package management solutions.
- Diagnose and resolve Kubernetes cluster and workload issues involving networking, storage, scaling, and resource limitations.
- Apply and maintain Kubernetes security practices, including RBAC and namespace isolation.
- Improve resource allocation and autoscaling configurations.
Cloud & Infrastructure
- Manage and support AWS infrastructure covering EKS, EC2, S3, IAM, VPC, networking, and security groups.
- Build and maintain Infrastructure as Code using Terraform and AWS CloudFormation.
- Contribute to infrastructure design reviews and capacity planning activities.
Reliability, Security & Production Operations
- Help maintain the availability and performance of production environments.
- Take part in a structured on‑call rotation supporting production systems.
- Investigate incidents, conduct root cause analysis (RCA), and implement improvements aimed at preventing recurrence.
- Participate in post‑incident reviews centered on lessons learned and improving long‑term system resilience.
- Continuously identify and remediate vulnerabilities affecting infrastructure, container images, and dependencies.
- Keep base images, Helm charts, system packages, and third‑party libraries updated in a timely manner.
- Work with development teams to address security findings, including CVEs, within code repositories.
- Improve infrastructure and container security across all environments.
CI/CD, Automation & Platform Improvement
- Maintain and improve CI/CD pipelines used for containerized application deployments.
- Automate build, testing, and Kubernetes deployment processes.
- Refine release processes to reduce downtime and lower deployment risks.
- Use GitHub Actions, Jenkins, and Ansible to automate operational activities and minimize manual work.
- Encourage DevOps and cloud‑native best practices across teams.
What We're Looking For
- At least 2 years of experience in DevOps, SRE, or Systems Engineering.
- Hands‑on experience operating production workloads on Kubernetes.
- Experience working with Amazon EKS.
- Experience deploying applications through Helm or comparable tooling.
- Experience using Docker for containerization.
- Experience with vulnerability scanning and remediation processes.
- Experience scripting with Bash and/or Python.
- Strong troubleshooting capabilities and an operational mindset.
- Willingness to join a structured on‑call rotation.
Desirable Experience
- Experience managing Kubernetes clusters across multiple environments, including development, staging, and production.
- Experience implementing autoscaling approaches such as HPA and cluster autoscaler.
- Experience with monitoring and observability technologies such as Prometheus, Grafana, ELK Stack, and Graylog.
- Familiarity with Kubernetes networking concepts, including Ingress, Services, and CNI.
- Experience supporting production systems using monitoring and alerting.
- Cloud certifications such as AWS Certified DevOps Engineer or an equivalent credential.
Work Setup
- Full-time, onsite
- Rotational shift schedule with monthly rotation across morning, afternoon, and night shifts
- Hybrid work flexibility available on select weekends and holidays, subject to company policies
- Additional remote work privileges may be granted based on performance and internal eligibility requirements
- Participation in an on‑call support rotation may be required depending on operational needs.