Senior Cloud/Devops Engineer Id92208

Agileengine

San Carlos de Bariloche

Remote

ARS 182,485,000 - 228,106,000

Full time

42 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Professional growth : Mentorship, Tech
USD-based pay with education, fitness,
Exciting projects : Fortune 500
Flextime : Remote and office options

Job summary

AgileEngine is seeking a Senior Cloud/DevOps Engineer to operate and improve the cloud infrastructure behind an enterprise data platform in a regulated healthcare environment.

The role requires 5+ years in Cloud Engineering, DevOps, or SRE, with hands-on AWS/EKS, Argo Workflows, Terraform, and strong English communication. Remote LatAm coverage window and on-call rotation are included.

Qualifications

  • 5+ years of professional experience in Cloud Engineering, DevOps or Site Reliability Engineering.
  • Strong hands-on experience operating AWS infrastructure in production environments.
  • Advanced experience with Kubernetes and Amazon EKS, including workload operations, troubleshooting, access, observability, capacity, and reliability.
  • Hands-on experience administering and troubleshooting Argo Workflows or comparable workflow orchestration platforms.
  • Strong Infrastructure as Code experience with Terraform and source-controlled infrastructure practices.
  • Experience building, hardening, and supporting CI/CD pipelines and production release processes.
  • Strong experience with monitoring, logging, alerting, and incident-routing tools such as Splunk, PagerDuty, Opsgenie, or comparable platforms.
  • Demonstrated ability to lead complex incident resolution, perform root-cause analysis, and translate findings into preventive improvements.
  • Proficiency in automation and scripting using Python, Shell, Bash, or similar languages.
  • Ability to make well-reasoned technical decisions, identify tradeoffs, estimate work, and drive improvements across a complex platform.
  • Experience mentoring engineers and collaborating effectively with Data Engineering, Security, Governance, Analytics, and business stakeholders.
  • Strong written and verbal English communication skills, with the ability to work directly with client stakeholders.
  • Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed on-call rotation.

Responsibilities

  • Provide senior technical ownership for the Cloud / DevOps service tower during the LatAm coverage window, including day-to-day operations, complex troubleshooting, and L2/L3 escalation.
  • Operate, maintain, and improve AWS infrastructure supporting the Data Platform, including Amazon EKS, S3, EventBridge, SQS, API Gateway, Lambda, and related services.
  • Administer Kubernetes-hosted workloads and Argo Workflows, including deployment, scheduling, monitoring, troubleshooting, capacity management, resiliency, and recovery.
  • Define and improve standards for Infrastructure as Code, configuration management, CI/CD, release execution, rollback, and environment consistency, primarily using Terraform and Git-based delivery practices.
  • Lead the consolidation and improvement of observability across infrastructure and data workloads, linking alerts to operational evidence from Argo, dbt, Snowflake, and supporting runbooks.
  • Improve alert routing and escalation workflows across tools such as Splunk, Opsgenie, PagerDuty, Microsoft Teams, and data-specific observability platforms.
  • Lead or support major incident response, root-cause analysis, post-incident reviews, and corrective actions, with clear communication to technical and service stakeholders.
  • Design and implement reliability improvements such as selective auto-remediation, dependency-aware alert correlation, impact analysis, and automation of repetitive operational work.
  • Track and contribute to service metrics including availability, SLA compliance, alert volumes, workflow reliability, deployment outcomes, and mean time to restore service.
  • Apply disciplined change-management, access-control, secrets-management, auditability, and documentation practices appropriate for a HIPAA-, GDPR-, and FDA-regulated environment.
  • Create and maintain runbooks, operating procedures, architecture context, recovery procedures, and knowledge-transfer materials.
  • Mentor Middle-level engineers, review technical work, improve team practices, and promote consistent execution across the distributed team.
  • Participate in the Cloud / DevOps on-call rotation for critical incidents outside staffed service hours.

Skills

5+ years exp
Kubernetes / EKS
Argo Workflows
Terraform
CI/CD pipelines
Python / Shell scripting
English communication

Tools

Argo Workflows
Terraform
Kubernetes
Python

Job description

Job Description

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US

If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE

We are looking for a Senior Cloud/DevOps Engineer to operate and improve the cloud infrastructure and reliability layers behind an enterprise data platform in a regulated healthcare environment.

The mandatory requirements are 5+ years of experience in Cloud Engineering, DevOps, or Site Reliability Engineering, advanced hands-on experience with AWS EKS and Kubernetes, experience administering and troubleshooting Argo Workflows, and strong English communication skills.

MUST HAVES
  • 5+ years of professional experience in Cloud Engineering, DevOps or Site Reliability Engineering.
  • Strong hands-on experience operating AWS infrastructure in production environments.
  • Advanced experience with Kubernetes and Amazon EKS, including workload operations, troubleshooting, access, observability, capacity, and reliability.
  • Hands-on experience administering and troubleshooting Argo Workflows or comparable workflow orchestration platforms.
  • Strong Infrastructure as Code experience with Terraform and source-controlled infrastructure practices.
  • Experience building, hardening, and supporting CI/CD pipelines and production release processes.
  • Strong experience with monitoring, logging, alerting, and incident-routing tools such as Splunk, PagerDuty, Opsgenie, or comparable platforms.
  • Demonstrated ability to lead complex incident resolution, perform root-cause analysis, and translate findings into preventive improvements.
  • Proficiency in automation and scripting using Python, Shell, Bash, or similar languages.
  • Ability to make well-reasoned technical decisions, identify tradeoffs, estimate work, and drive improvements across a complex platform.
  • Experience mentoring engineers and collaborating effectively with Data Engineering, Security, Governance, Analytics, and business stakeholders.
  • Strong written and verbal English communication skills, with the ability to work directly with client stakeholders.
  • Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed on-call rotation.
NICE TO HAVES
  • Experience supporting data-platform infrastructure involving Snowflake, dbt, Fivetran, HVR, Tableau Cloud, or custom ingestion pipelines.
  • Familiarity with data-specific observability platforms such as SYNQ.
  • Experience modernizing or migrating legacy orchestration and ingestion solutions such as Boomi or AWS Data Pipeline.
  • Experience with service-management and change-control tools such as Freshservice and Jira.
  • Experience operating in healthcare, life sciences, financial services, or another regulated environment.
  • Familiarity with HIPAA, GDPR, FDA-related controls, least-privilege access, separation of duties, and audit-ready operational practices.
WHAT YOU WILL DO
  • Provide senior technical ownership for the Cloud / DevOps service tower during the LatAm coverage window, including day-to-day operations, complex troubleshooting, and L2/L3 escalation.
  • Operate, maintain, and improve AWS infrastructure supporting the Data Platform, including Amazon EKS, S3, EventBridge, SQS, API Gateway, Lambda, and related services.
  • Administer Kubernetes-hosted workloads and Argo Workflows, including deployment, scheduling, monitoring, troubleshooting, capacity management, resiliency, and recovery.
  • Define and improve standards for Infrastructure as Code, configuration management, CI/CD, release execution, rollback, and environment consistency, primarily using Terraform and Git-based delivery practices.
  • Lead the consolidation and improvement of observability across infrastructure and data workloads, linking alerts to operational evidence from Argo, dbt, Snowflake, and supporting runbooks.
  • Improve alert routing and escalation workflows across tools such as Splunk, Opsgenie, PagerDuty, Microsoft Teams, and data-specific observability platforms.
  • Lead or support major incident response, root-cause analysis, post-incident reviews, and corrective actions, with clear communication to technical and service stakeholders.
  • Design and implement reliability improvements such as selective auto-remediation, dependency-aware alert correlation, impact analysis, and automation of repetitive operational work.
  • Track and contribute to service metrics including availability, SLA compliance, alert volumes, workflow reliability, deployment outcomes, and mean time to restore service.
  • Apply disciplined change-management, access-control, secrets-management, auditability, and documentation practices appropriate for a HIPAA-, GDPR-, and FDA-regulated environment.
  • Create and maintain runbooks, operating procedures, architecture context, recovery procedures, and knowledge-transfer materials.
  • Mentor Middle-level engineers, review technical work, improve team practices, and promote consistent execution across the distributed team.
  • Participate in the Cloud / DevOps on-call rotation for critical incidents outside staffed service hours.
PERKS AND BENEFITS
  • Professional growth : Mentorship, TechTalks, and personalized growth roadmaps.
  • Competitive compensation : USD-based pay with education, fitness, and team activity budgets.
  • Exciting projects : Modern solutions with Fortune 500 and top product companies.
  • Flextime : Flexible schedule with remote and office options.
Requirements
  • 6+ years of professional experience in Cloud Engineering, Platform Engineering, DevOps or Site Reliability Engineering
  • Deep hands-on expertise with AWS production environments, including cloud architecture, security, networking fundamentals, access management, monitoring, capacity, and operational troubleshooting.
  • Advanced experience with Kubernetes and Amazon EKS, including workload deployment, cluster and application troubleshooting, observability, scaling, upgrades, access, and reliability.
  • Hands-on experience operating Argo Workflows or a comparable orchestration platform supporting production data workloads.
  • Strong experience with Infrastructure as Code, preferably Terraform, and with source-controlled configuration, CI/CD, release automation, and rollback practices.
  • Strong experience designing and operating observability, logging, monitoring, alerting, and incident-routing solutions for distributed production platforms.
  • Demonstrated ability to lead major incidents and cross-functional troubleshooting across infrastructure, applications, data pipelines, and analytics layers.
  • Experience operating production platforms with defined service levels, escalation paths, runbooks, change controls, release processes, and on-call responsibilities.
  • Strong working knowledge of modern data platforms, including Snowflake, S3-based data lakes, SQL, dbt, managed ingestion tools such as Fivetran or HVR, and custom data pipelines.
  • Understanding of data quality, freshness, lineage, schema evolution, pipeline dependencies, backfills, and recovery procedures.
  • Proficiency in Python, Shell, Bash, or comparable languages for automation and operational tooling.
  • Ability to make well-reasoned architecture and operational decisions, communicate tradeoffs, estimate work, identify risks, and guide teams through change.
  • Proven experience mentoring engineers, reviewing technical work, delegating ownership, and improving engineering processes across a distributed team.
  • Strong stakeholder-management and communication skills, including the ability to collect requirements, explain technical risks, and present recommendations to client leaders.
  • Strong written and verbal English communication skills.
  • Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed senior escalation and on-call rotation.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud/Devops Engineer Id92208
Senior Cloud/Devops Engineer Id92208

Agileengine • Mar del Plata

Hybrid
ARS 182,485,000 - 273,727,000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Devops Engineer – Remote (Aws, Kubernetes)
Senior Devops Engineer – Remote (Aws, Kubernetes)

Flux It • Rosario

Remote
ARS 167,117,000 - 258,272,000
Professional growth
Competitive USD-based pay
Exciting projects
+1
Remote Senior Devops Engineer: Azure, Kubernetes & Ci/Cd
Remote Senior Devops Engineer: Azure, Kubernetes & Ci/Cd

Encora Inc. • Partido de Quilmes

Hybrid
ARS 182,310,000 - 273,465,000
Professional growth
Competitive compensation
Exciting projects
+1
Cloud Devops Engineer For Ai-Scale Infra & Automation
Cloud Devops Engineer For Ai-Scale Infra & Automation

Lever, Inc. • Buenos Aires

Hybrid
ARS 182,310,000 - 273,465,000
Professional growth: Mentorship, TechT
USD‑based pay with education, fitness,
Exciting projects with Fortune 500
+1
Senior Devops Engineer: Aws, Kubernetes, Remote Latam
Senior Devops Engineer: Aws, Kubernetes, Remote Latam

Arize Ai, Inc • Buenos Aires

Hybrid
ARS 182,485,000 - 273,727,000
USD-based pay
Flex-time
Remote + Office options
+1
Senior Genesys Cloud Engineer — Remote & Global Impact
Senior Genesys Cloud Engineer — Remote & Global Impact

Miratech • Victoria

Hybrid
ARS 167,117,000 - 227,887,000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Genesys Cloud Engineer - Remote, Global Impact
Senior Genesys Cloud Engineer - Remote, Global Impact

Atmosera • Partido de Quilmes

Remote
ARS 182,771,000 - 274,156,000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Tech Lead .Net: Líder Técnico De Microservicios (Remoto)
Tech Lead .Net: Líder Técnico De Microservicios (Remoto)

Skydropx - Frenet • Rosario

Remote
ARS 258,272,000 - 319,042,000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Devops Engineer — Remote (Aws/Eks/Terraform)
Senior Devops Engineer — Remote (Aws/Eks/Terraform)

Flux It • Partido de Quilmes

Hybrid
ARS 182,310,000 - 273,465,000
Professional growth
Competitive pay
Exciting projects
+1
Senior Devsecops Engineer — Aws Cloud Platform & Sre
Senior Devsecops Engineer — Aws Cloud Platform & Sre

Talan • Ciudad de Mendoza

On-site
ARS 212,695,000 - 273,465,000
Professional growth
Competitive compensation
Exciting projects
+1