AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people‑first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a DevOps Engineer to maintain production and staging environments, execute application migrations, and improve the reliability and automation of cloud infrastructure.
WHAT YOU WILL DO
- Migrate applications between environments; these migrations typically take place on weekends, so weekend availability is required.
- Monitor and support production and staging environments in real time, ensuring high availability, performance, and stability.
- Respond to incidents, perform triage and root cause analysis, and contribute to post‑incident reviews and remediation efforts.
- Participate in an on‑call rotation with defined SLAs.
- Handle ad‑hoc and unplanned operational requests from Product, Support, and internal teams.
- Maintain and enhance monitoring, alerting, dashboards, logs, and metrics; improve signal‑to‑noise ratio and standardise observability practices.
- Support CI/CD pipelines, production releases, and GitOps workflows.
- Contribute to automation efforts to reduce operational toil.
- Maintain and improve Kubernetes‑based infrastructure and containerised workloads.
- Support Infrastructure as Code practices and ongoing environment improvements.
MUST HAVES
- 2+ years of experience in Site Reliability Engineering, DevOps, or Production Operations.
- Demonstrable AWS EKS experience supporting production environments.
- Experience supporting production SaaS applications.
- Strong understanding of CI/CD systems (GitHub Actions, Jenkins, CircleCI, or similar).
- GitOps experience and strong Git fundamentals.
- Experience using GitHub, Jira, and Confluence in collaborative engineering environments.
- Kubernetes experience (EKS, kOps, or similar).
- Observability stack experience (Grafana, Prometheus, Loki, PagerDuty, or similar).
- Scripting experience (Bash, Python, or Go).
- Infrastructure as Code experience (Terraform, Helm, or similar).
- Working knowledge of relational databases (e.g., PostgreSQL, MySQL) for troubleshooting and operational support.
- Comfortable working within structured operational processes and SLAs.
- Strong written and verbal English communication skills; able to clearly explain technical concepts.
- Self‑driven with a growth mindset.
- Weekend availability for application migrations between environments when needed.
NICE TO HAVES
- Experience in multi‑tenant SaaS environments.
- Experience working in globally distributed teams.
- Familiarity with ChatOps practices.
- Experience improving monitoring quality and reducing alert fatigue.
- Background in operational cost optimisation.
PERKS AND BENEFITS
- Professional growth: Accelerate your professional journey with mentorship, TechTalks, and personalised growth roadmaps.
- Competitive compensation: We match your ever‑growing skills, talent, and contributions with competitive USD‑based compensation and budgets for education, fitness, and team activities.
- A selection of exciting projects: Join projects with modern solutions development and top‑tier clients that include Fortune 500 enterprises and leading product brands.
- Flextime: Tailor your schedule for an optimal work‑life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.
THE ROLE: Infrastructure / DevOps Engineer
We are looking for an Infrastructure / DevOps Engineer to join our Cloud Platform Engineering Team.
In this role, you will own cloud infrastructure, deployment automation, and system reliability, with a strong focus on cloud‑agnostic portability, Kubernetes administration, Infrastructure as Code, and automated CI/CD workflows.
We are looking for someone with strong hands‑on experience in Kubernetes, Terraform, GitHub Actions, multi‑cloud environments, and IAM solutions such as Keycloak.
RESPONSIBILITIES
- Kubernetes Administration: Deploy, manage, scale, and troubleshoot workloads on Kubernetes, initially using Google Kubernetes Engine (GKE), with Helm and Kustomize.
- Infrastructure as Code: Build modular and cloud‑agnostic infrastructure using Terraform, enabling consistent deployments across GCP, other public clouds, and on‑premise environments.
- CI/CD Pipeline Ownership: Design, maintain, and optimise reusable GitHub Actions workflows for automated testing, containerisation, and deployment.
- Manage container images and deployment workflows using GitHub Container Registry (GHCR).
- IAM & Platform Deployment: Deploy, operate, and scale highly available Keycloak clusters.
- Configure GitHub Workload Identity Federation (OIDC) to provide GitHub Actions with secure, keyless access to GCP resources.
- Implement and maintain secure secrets management practices.
- Collaborate with technical leadership to evaluate and improve deployment strategies, including traditional CI/CD and GitOps approaches using tools such as ArgoCD or Flux.
- Contribute to the reliability, scalability, security, and portability of the overall cloud platform.
REQUIREMENTS
- Advanced hands‑on experience managing production Kubernetes environments, including container networking.
- Strong proficiency with Terraform, Helm, and Infrastructure as Code (IaC) methodologies.
- Deep experience designing and maintaining deployment pipelines using GitHub Actions.
- Experience working with GitHub Container Registry (GHCR).
- Experience deploying and maintaining production Keycloak/IAM environments.
- Strong understanding of OIDC, identity management, authentication, and authorization concepts.
- Experience with secrets management and cloud security best practices.
- Experience working with GCP/GKE and an understanding of multi‑cloud or cloud‑agnostic infrastructure approaches.
- Strong understanding of containerised environments, deployment automation, and system reliability.
CORE TECH STACK
- Kubernetes (GKE) | Terraform | Helm | Kustomize | GitHub Actions | GHCR | GCP | Multi‑Cloud | Keycloak | OIDC | IaC | ArgoCD / Flux
ABOUT THE ROLE
We are looking for a Senior Cloud/DevOps Engineer to operate and improve the cloud infrastructure and reliability layers behind an enterprise data platform in a regulated healthcare environment.
The mandatory requirements are 5+ years of experience in Cloud Engineering, DevOps, or Site Reliability Engineering, advanced hands‑on experience with AWS EKS and Kubernetes, experience administering and troubleshooting Argo Workflows, and strong English communication skills.
MUST HAVES
- 5+ years of professional experience in Cloud Engineering, DevOps or Site Reliability Engineering.
- Strong hands‑on experience operating AWS infrastructure in production environments.
- Advanced experience with Kubernetes and Amazon EKS, including workload operations, troubleshooting, access, observability, capacity, and reliability.
- Hands‑on experience administering and troubleshooting Argo Workflows or comparable workflow orchestration platforms.
- Strong Infrastructure as Code experience with Terraform and source‑controlled infrastructure practices.
- Experience building, hardening, and supporting CI/CD pipelines and production release processes.
- Strong experience with monitoring, logging, alerting, and incident‑routing tools such as Splunk, PagerDuty, Opsgenie, or comparable platforms.
- Demonstrated ability to lead complex incident resolution, perform root‑cause analysis, and translate findings into preventive improvements.
- Proficiency in automation and scripting using Python, Shell, Bash, or similar languages.
- Ability to make well‑reasoned technical decisions, identify tradeoffs, estimate work, and drive improvements across a complex platform.
- Experience mentoring engineers and collaborating effectively with Data Engineering, Security, Governance, Analytics, and business stakeholders.
- Strong written and verbal English communication skills, with the ability to work directly with client stakeholders.
- Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed on‑call rotation.
NICE TO HAVES
- Experience supporting data‑platform infrastructure involving Snowflake, dbt, Fivetran, HVR, Tableau Cloud, or custom ingestion pipelines.
- Familiarity with data‑specific observability platforms such as SYNQ.
- Experience modernising or migrating legacy orchestration and ingestion solutions such as Boomi or AWS Data Pipeline.
- Experience with service‑management and change‑control tools such as Freshservice and Jira.
- Experience operating in healthcare, life sciences, financial services, or another regulated environment.
- Familiarity with HIPAA, GDPR, FDA‑related controls, least‑privilege access, separation of duties, and audit‑ready operational practices.
WHAT YOU WILL DO
- Provide senior technical ownership for the Cloud / DevOps service tower during the LatAm coverage window, including day‑to‑day operations, complex troubleshooting, and L2/L3 escalation.
- Operate, maintain, and improve AWS infrastructure supporting the Data Platform, including Amazon EKS, S3, EventBridge, SQS, API Gateway, Lambda, and related services.
- Administer Kubernetes‑hosted workloads and Argo Workflows, including deployment, scheduling, monitoring, troubleshooting, capacity management, resiliency, and recovery.
- Define and improve standards for Infrastructure as Code, configuration management, CI/CD, release execution, rollback, and environment consistency, primarily using Terraform and Git‑based delivery practices.
- Lead the consolidation and improvement of observability across infrastructure and data workloads, linking alerts to operational evidence from Argo, dbt, Snowflake, and supporting runbooks.
- Improve alert routing and escalation workflows across tools such as Splunk, Opsgenie, PagerDuty, Microsoft Teams, and data‑specific observability platforms.
- Lead or support major incident response, root‑cause analysis, post‑incident reviews, and corrective actions, with clear communication to technical and service stakeholders.
- Design and implement reliability improvements such as selective auto‑remediation, dependency‑aware alert correlation, impact analysis, and automation of repetitive operational work.
- Track and contribute to service metrics including availability, SLA compliance, alert volumes, workflow reliability, deployment outcomes, and mean time to restore service.
- Apply disciplined change‑management, access‑control, secrets‑management, auditability, and documentation practices appropriate for a HIPAA‑, GDPR‑, and FDA‑regulated environment.
- Create and maintain runbooks, operating procedures, architecture context, recovery procedures, and knowledge‑transfer materials.
- Mentor Middle‑level engineers, review technical work, improve team practices, and promote consistent execution across the distributed team.
- Participate in the Cloud / DevOps on‑call rotation for critical incidents outside staffed service hours.
PERKS AND BENEFITS
- Professional growth: Mentorship, TechTalks, and personalised growth roadmaps.
- Competitive compensation: USD‑based pay with education, fitness, and team activity budgets.
- Exciting projects: Modern solutions with Fortune 500 and top product companies.
- Flextime: Flexible schedule with remote and office options.