We are working with a client that develops and operates complex technology platforms where scalability, reliability, and availability are critical.
They are looking for a DevOps Engineer to help automate and improve the infrastructure, development, and deployment processes supporting a cloud-native, event-driven microservices environment.
You’ll work across cloud infrastructure, Kubernetes, CI/CD, Infrastructure as Code, monitoring, and automation, partnering closely with software development, QA, and IT teams. The role also participates in a 24/7 on-call rotation supporting business-critical production systems.
Responsibilities
- Design, deploy, and maintain scalable infrastructure across AWS, Azure, or Google Cloud
- Manage cloud environments with a focus on availability, reliability, security, performance, and cost
- Operate Kubernetes platforms, including stateful workloads, distributed SQL databases, and event-streaming technologies such as Kafka or NATS
- Build and maintain containerized CI/CD pipelines using tools such as Dagger, GitHub Actions, or GitLab CI
- Automate build, testing, and deployment processes across applications and services
- Implement Infrastructure as Code using Terraform or Ansible, alongside Helm and GitOps workflows
- Manage Kubernetes deployments using tools such as Argo CD
- Automate infrastructure provisioning, configuration, and scaling to create consistent and repeatable environments
- Manage containerized applications using Docker and Kubernetes
- Implement and maintain monitoring, logging, tracing, and alerting using technologies such as Prometheus, Grafana, OpenTelemetry, and ELK
- Monitor production systems, troubleshoot performance issues, and respond to incidents as part of the on-call rotation
- Work closely with developers, system administrators, and QA engineers to improve deployment and operational processes
- Integrate security practices into infrastructure and deployment workflows, including secrets management, access controls, and policy-as-code
- Support compliance with relevant standards including SOC 2, ISO 27001, and PCI DSS
Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Software Engineering, or a related field, or equivalent practical experience
- 3+ years of experience within DevOps, Site Reliability Engineering (SRE), or a similar role
- Strong hands-on experience with at least one major cloud platform: AWS, Azure, or Google Cloud
- Experience designing, building, and maintaining CI/CD pipelines
- Hands-on experience with CI/CD technologies such as Dagger, GitHub Actions, GitLab CI, or Jenkins
- Strong Infrastructure as Code experience using technologies such as Terraform, Ansible, and Helm
- Experience with GitOps deployment approaches and tools such as Argo CD
- Strong knowledge of Docker and Kubernetes
- Automation and scripting experience using Python, Bash, or PowerShell
- Experience with monitoring and observability technologies such as Prometheus, Grafana, OpenTelemetry, or ELK
- Strong knowledge of Git, GitHub, or GitLab
- Strong troubleshooting skills across infrastructure, deployments, and production systems
- Ability to collaborate effectively across development, QA, operations, and infrastructure teams
- Paid vacation, sick leave, and bereavement leave
- Health and dental plans
- Employee and Family Assistance Program (EFAP)