Cloud DevOps Engineer

Azumo

Santiago

A distancia

CLP 88.409.000 - 147.348.000

Jornada completa

Hace 13 días
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote-First LATAM
PTO
U.S. Holidays
AI Training
Mentored Growth
Profit Sharing
US Remuneration
Maternity Coverage

Descripción de la vacante

Azumo is seeking a Cloud DevOps Engineer to own production infra across client environments, including clusters, CI/CD pipelines, and monitoring to keep systems healthy.

You will manage infrastructure for AI workloads, optimize cost, and ensure security across Latin America time zones. This is a fully remote role with strong tooling and collaboration with AI engineers.

Formación

  • 5+ years of DevOps, SRE or production infra experience.
  • Kubernetes in production with provisioning/upgrades.
  • IaC using Terraform or equivalent with environment parity.
  • CI/CD pipelines with GitHub Actions or GitLab; container images ownership.
  • Cloud deployment on AWS/Azure/GCP with cost awareness.
  • Monitoring/incident response with Datadog or CloudWatch.
  • Security hardening of Linux, containers and Kubernetes.
  • Strong English communication.

Responsabilidades

  • Own clusters, deployments, and infra for AI workloads.
  • Build and maintain CI/CD pipelines and delivery tooling.
  • Ensure observability and proactive incident response.
  • Work across client environments and internal infra.
  • Optimize cost and performance; document improvements.

Conocimientos

DevOps
SRE
Linux
Networking
Git
Bash
Python
Go
Kubernetes
Terraform
CI/CD
GitHub Actions
GitLab
Datadog
CloudWatch
Istio
Linkerd
Helm
Kustomize
PostgreSQL
AWS
GCP
Azure

Educación

Bachelor's degree in Computer Science or related

Herramientas

Terraform
GitHub Actions
GitLab
Datadog
CloudWatch
Istio
Linkerd
Helm
Kustomize
PostgreSQL
AWS
GCP
Azure

Descripción del empleo

Cloud DevOps Engineer

Azumo builds and operates production AI systems for companies ranging from seed-stage startups to Meta. We are hiring a Cloud DevOps Engineer to own the infrastructure those systems run on: the clusters, the pipelines that deploy to them, and the monitoring that says whether any of it is healthy. The role is fully remote across Latin America, aligned to your client's working day.

You will not be handed a runbook. Azumo has shipped production infrastructure since 2016, and the work here is in systems that are already live: what breaks, what a cluster costs to run, and what nobody has automated yet because it was always easier to do by hand.

Where this role sits

At Azumo, ownership is split by what gets built: the pipelines, the method behind a decision, the product, and what an AI system does once it's live. This role owns what all of it runs on — clusters, deployment, and the infrastructure those systems depend on to stay up.

A system that is down, a slow query nobody escalates, a cluster that costs more than it should: they are all yours. Recovery first, root cause after, and the change that keeps it from happening again.

What you will build

  • Clusters that hold. Kubernetes in production — provisioning, upgrades, networking, and the workload configuration that decides whether a bad deploy takes one service down or all of them.
  • Infrastructure as code. Terraform or equivalent, with environments reproducible from the repository rather than from memory, and drift that gets detected instead of discovered.
  • Delivery pipelines. CI/CD that developers and QA can rely on, container images that are built the same way every time, and deployments to production that are unremarkable.
  • Observability and response. Monitoring and alerting that fires on what matters, plus the investigation, the root cause and the follow-up change when something does break.
  • Infrastructure for AI workloads. The clusters and inference services the AI Engineer lane deploys onto, whether that's hosted APIs or models we run ourselves, with the cost and scaling profile each one brings.
  • Cost and capacity. What the infrastructure costs to run, what it would cost at three times the traffic, and which of the two problems is worth solving now.
  • Hardening. Linux, Kubernetes, containers and the service mesh secured as a default rather than as a project, with secrets, access and audit trails that survive a client's review.
  • Tooling, documentation and support. The internal tooling your own team runs on, the documentation that makes any of it operable by someone else, and the developers and QA you unblock during the release cycle.
  • Work inside the client's environment, when that's the engagement. Some of this work runs on Azumo's own infrastructure and some inside a client's — their repositories, their cloud account, their change process. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.

How we work

Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an automated audit across the whole codebase on day one and every day after, grading security, cost, and architecture findings by severity with the exact file and line, so a small team can move quickly without quality drifting. We stay vendor-neutral across OpenAI, Anthropic, and open-weight models, and we run Valkyrie, our own production layer, when a single interface to any model is the right call.

About Azumo

Azumo is a San Francisco based software development company that has been building intelligent applications since 2016. We provide nearshore AI engineering teams to organizations that need production AI faster than they can hire for it: as an embedded engineering team, as AI staff augmentation alongside an existing team, or as a full project build. Our engineers work from Latin America, aligned to United States time zones, and have delivered for Twitter, Meta, Discovery Channel, Omnicom, UnitedHealth, and CENTEGIX.

We hire for seniority and test for it before anyone joins a client team. We support engineers in going deep on the modern AI stack, and we give time back to open-source work, community teaching, and philanthropy.

Basic qualifications
  • 5+ years as a DevOps, SRE or systems engineer running production infrastructure, with the fundamentals that go with it: Linux administration, networking, Git, and scripting in Bash plus Python or Go.
  • Kubernetes in production, at depth: not only deploying to a cluster someone else built, but provisioning, upgrading and debugging one when it misbehaves.
  • Infrastructure as code as your working practice: Terraform or equivalent, with state, modules and environment parity you can defend.
  • CI/CD pipelines you built and maintain — GitHub Actions, GitLab or equivalent — and container images you are responsible for.
  • Cloud deployment experience on AWS, Azure or GCP, including the managed Kubernetes service and the cost model that comes with it.
  • Monitoring and incident work: Datadog, CloudWatch or equivalent, and a track record of taking an incident from alert to root cause to a change that prevented the repeat.
  • Security hardening of Linux, containers and Kubernetes as part of how you build, not as a separate phase.
  • Working discipline around infrastructure cost and capacity. You can explain what a cluster costs to run and what you did about it.
  • Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in real delivery work.
  • Clear written and spoken English, C1 or above, and the confidence to explain a technical trade-off directly to a client.
  • Bachelor's degree in Computer Science, a related field, or equivalent professional experience.
Preferred qualifications
  • Service mesh and traffic management: Istio, Linkerd or equivalent, and the failure modes they introduce as well as the ones they solve.
  • Helm, Kustomize or an equivalent approach to templating and environment configuration.
  • Database operations at production scale: PostgreSQL, MongoDB, RDS, DynamoDB or equivalent.
  • Running or scaling inference workloads, and the cost and latency trade-offs against hosted APIs.
  • Delivery under a compliance regime such as SOC 2 or HIPAA.
  • Cloud certifications, or contributions to open-source infrastructure tooling and published technical writing.

100% remote-first culture (work anywhere in Latin America)

Paid time off (PTO)

U.S. Holidays

Solid AI Training and certification

Mentored career development

Profit sharing

$US remuneration

Maternity coverage

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Forward Deployed AI Engineer (Python) - Remote -Latin America
Forward Deployed AI Engineer (Python) - Remote -Latin America

FullStack • Valparaíso

A distancia
CLP 115.607.000 - 173.410.000
100% remote work
Continuing education opportunities
Grow your career with leading clients
Forward Deployed AI Engineer (Python) - Remote -Latin America
Forward Deployed AI Engineer (Python) - Remote -Latin America

FullStack • Concepcion

A distancia
CLP 134.875.000 - 183.044.000
100% remote work
Competitive pay
Continuing education opportunities
+1
Forward Deployed AI Engineer (TypeScript) - Remote - Latin America
Forward Deployed AI Engineer (TypeScript) - Remote - Latin America

FullStack • Concepcion

A distancia
CLP 144.509.000 - 183.044.000
Competitive pay
100% remote work
Continuing education opportunities
Senior Software Engineer (CRM+ Platform) - Remote - Latin America
Senior Software Engineer (CRM+ Platform) - Remote - Latin America

FullStack • Santiago

A distancia
CLP 115.607.000 - 173.410.000
100% remote work
Competitive pay
Continuing education opportunities
+1
DevOps Engineer (AWS) - Remote - Latin America
DevOps Engineer (AWS) - Remote - Latin America

FullStack • Concepcion

A distancia
CLP 86.042.000 - 124.283.000
Remote work
Competitive pay
Education opportunities
+1
RPA Developer (UiPath)- Latin America- Remote
RPA Developer (UiPath)- Latin America- Remote

Azumo • Santiago

Presencial
CLP 12.000.000 - 20.000.000
Paid time off
U.S. Holidays
Training
+4
AI Engineer, Automation and Developer Tooling Santiago, Chile · Remote
AI Engineer, Automation and Developer Tooling Santiago, Chile · Remote

Arcadia Power, Inc. • Santiago

A distancia
CLP 74.931.000 - 107.308.000
Remote first
Flexible PTO
11 holidays
+5
AI Developer
AI Developer

Brand & Bot • Chile

A distancia
CLP 89.021.000 - 128.586.000
Full-Stack Developer (AI Focused)
Full-Stack Developer (AI Focused)

Brand & Bot • Chile

Presencial
CLP 56.444.000 - 94.073.000
AI Machine Learning Engineer - Remote - Latin America
AI Machine Learning Engineer - Remote - Latin America

FullStack Labs, Inc. • Chile

A distancia
CLP 117.878.000 - 176.817.000
Competitive pay
Opportunities with leading startups &
Continuing education opportunities
+1