Senior Site Reliability Engineer (LATAM)

Reap

Argentina

A distancia

ARS 1.500.000 - 2.500.000

Jornada completa

Hace 4 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Transforma esta oferta en una entrevista: un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

Reap is seeking a senior Site Reliability Engineer to lead reliability for a payments platform built on AWS. You will drive IaC, define SLOs, and own incident response while pushing for self-service and governance across a distributed team.

You will implement observability, manage PCI DSS compliance, and collaborate with product teams to treat the platform as a product with measurable reliability. Expect a hands-on, high-impact role across multiple regions.

Formación

  • Real SRE practice with defined SLOs and error budgets.

Responsabilidades

  • Define SLIs/SLOs with product teams and set error budgets.
  • Participate in on-call rotations and build incident postmortems.
  • Drive IaC coverage across legacy infra, ensure no drift.
  • Consolidate Terraform into modular, governed codebase.
  • Automate account provisioning for new regions/services.
  • Build self-service platforms for deployment/observability.
  • Create ephemeral environments on demand with auto-cleanup.
  • Implement industry-standard observability and actionable alerts.
  • Own cloud operations for PCI DSS and regulated systems.
  • Embed security into infra: least-privilege IAM and segmentation.
  • Support AI assistants with safe, auditable actions.
  • Collaborate with product/engineering to treat reliability as a product.

Conocimientos

SRE practice
Linux fundamentals
Terraform
AWS multi-account
Kubernetes
CI/CD
Go/Python/Bash
Observability
Incident response
Security/NIST PCI DSS
Self-management
Communication

Herramientas

Terraform
CloudFormation
ECS/Fargate
Lambda
EKS
GitHub Actions
Argo CD
Kafka
SQS/EventBridge
New Relic/CloudWatch/Datadog/Prometheus

Descripción del empleo

About Reap

Reap is a global financial technology company headquartered in Hong Kong with employees across multiple countries. We enable financial connectivity and access for businesses worldwide by combining traditional finance with stablecoins for efficient money movement.

Through our stablecoin-powered corporate cards, payments, and expense management tools, we streamline financial operations and help businesses scale. Our APIs enable businesses to integrate stablecoin-enabled finance into their own products and services — from issuing Visa cards to facilitating cross-border payments.

Backed by leading investors including Index Ventures and HashKey Capital, Reap is building the future of borderless, stablecoin-enabled finance.

About the Role

Reap is building a Site Reliability Engineering practice, and this role is central to it.

We run card issuing, payouts, FX and stablecoin settlement across multiple AWS regions, under PCI DSS and financial regulation. The platform is growing quickly — into new markets, new products, and now agent-initiated payments — and the infrastructure underneath it needs to grow up with it. That means real service ownership, reliability measured in SLIs and SLOs rather than intuition, and a platform that engineering teams can serve themselves from instead of queueing for.

That work is largely still ahead of us, which is the appeal. You will help decide what reliability means at Reap, what the platform looks like, and what good engineering practice is in this domain — rather than inheriting someone else's answers.

This is a deeply hands‑on senior individual contributor role. We expect you to lead by example and drive technical excellence through the systems you build, the standards you set, and the way you help the engineers around you level up. The team is deliberately flat and distributed across time zones.

Technologies You'll Use
  • Cloud: AWS, multi-account across multiple regions — Transit Gateway, PrivateLink, site-to-site VPN, WAF, KMS

  • Compute: ECS/Fargate, Lambda and EKS

  • Infrastructure as Code: Terraform and CloudFormation

  • CI/CD: GitHub Actions, Argo CD

  • Databases: Aurora PostgreSQL, ElastiCache/Redis

  • Messaging and streaming: SQS, EventBridge, Kafka

  • Observability: New Relic, CloudWatch

  • Languages and scripting: Python, Go and Bash

What You'll Do

As a Senior Site Reliability Engineer at Reap, you will join a team of experienced engineers transitioning a DevOps organisation into Site Reliability Engineering. You will own the reliability of a payments platform while rebuilding the foundation it runs on — the lights stay on while the platform gets replaced underneath them. We treat repetitive manual work as a bug in the platform, not a chore for a human, and your job is to delete whole categories of it rather than absorb them faster.

  • Define what reliability means here: set SLIs and SLOs with product and engineering teams, introduce error budgets, and make "ship or stabilise?" a matter of arithmetic rather than argument

  • Take part in our on-call rotation, and help build the incident response and blameless postmortem practice around it

  • Drive Reap to full Infrastructure as Code coverage: bring the remaining legacy infrastructure under IaC, and get to no manual provisioning, no drift, and every resource defined, versioned and reproducible

  • Consolidate our Terraform estate into a coherent, well-structured codebase with module standards, governance and automated drift detection

  • Automate account provisioning and environment setup so that new regions and services can be stood up repeatably and consistently

  • Build the golden paths and self-service interfaces that let product and engineering teams provision, deploy and observe without filing a ticket, with escape hatches for the cases they do not cover

  • Design and implement an ephemeral environment platform so any developer, and any coding agent, can get an isolated production‑like environment on demand and have it cleaned up automatically

  • Implement industry-standard observability across logging, metrics and distributed tracing, with alerting that is actionable and trusted

  • Own cloud operations for PCI DSS and regulated financial systems — uptime, failover, capacity, disaster recovery and incident response

  • Embed security into the infrastructure layer: secrets management, least-privilege IAM, network segmentation and compliance controls

  • Build infrastructure that lets AI assistants and agents operate safely: sane blast radius, strong isolation, auditable actions

  • Partner with product and engineering teams so that reliability work is negotiated rather than imposed — we treat the platform as a product and our engineering teams as its customers

Skills We're Looking For
  • Real SRE practice, not just the vocabulary: you have defined SLOs, run error budgets, or built an incident and postmortem process that people actually used

  • Strong Linux and computer‑systems fundamentals: internals, networking, and administration. This work rests on them

  • Deep Terraform: state management, module design, and enforcing standards across a team. You have owned a full IaC migration or a greenfield buildout at scale

  • Strong AWS across multi‑account, multi‑region estates: Control Tower, IAM, networking, RDS, and cost management

  • Containers and orchestration: comfortable operating ECS and Fargate, and able to build and run production Kubernetes at scale — cluster lifecycle and upgrades, autoscaling, resource limits and multi‑tenant workload isolation

  • Serverless and event‑driven systems: Lambda, SQS and EventBridge in production, including the failure modes that only show up at scale — retries, poison messages, ordering and idempotency

  • GitOps and delivery: Argo CD or equivalent, and CI/CD pipelines with GitHub Actions that engineers actually enjoy using

  • Observability in practice: logging, metrics and tracing with tools such as New Relic, CloudWatch, Datadog or Prometheus — and getting standards adopted, not just published

  • Python, Go or Bash for automation and tooling

  • Full‑lifecycle ownership: you own a service across its whole life — from understanding the need, through design and delivery, to observability, incident response, capacity, upgrades and security patching

  • Experience operating in a regulated environment, and comfort with the constraints that come with PCI DSS and financial compliance

  • Self‑management and cross‑team influence: you can run your own projects and build alignment without authority. Nobody will sequence your work for you

  • Communication as a first‑class skill, weighted equally with technical depth. You can take a problem, explain your approach in plain language, and decompose it into work — for a person or for an agent

  • AI‑assisted engineering: you already use AI tools seriously for code, review, automation, incident analysis and documentation, and you have opinions about where they work and where they do not

  • Genuine comfort with production incidents. In payments, high blast‑radius incidents preempt everything

Bonus Skills
  • Significant experience in SRE, DevOps or infrastructure engineering, including time in an organisation with a mature SRE practice — defined SLOs, self‑service deploys, and real on‑call

  • Experience in fintech, payments, card issuing or another regulated environment

  • Hands‑on PCI DSS work: designing cardholder‑data environments, or reducing scope

  • Kubernetes or EKS inside a PCI DSS environment: isolating the cardholder‑data environment through namespace and network‑policy segmentation, dedicated node groups, admission control and pod security standards — and the audit logging that makes it provable

  • Having designed and operated ephemer​​al or on‑demand environment platforms, including the hard parts — database seeding, service dependencies and secrets in short‑lived environments

  • Having built an internal developer platform, account factory or landing zone from scratch

  • Configuration management with Ansible or similar

  • AWS certifications

  • Curiosity about stablecoins and the Web2 / Web3 intersection

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior SRE: Payments Platform Reliability & IaC
Senior SRE: Payments Platform Reliability & IaC

Reap • Argentina

A distancia
ARS 1.500.000 - 2.500.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Strategic Staffing Solutions • Argentina

Presencial
ARS 3.000.000 - 7.000.000
Full time employment
100% remote
Competitive salary in ARS
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

EPAM Systems • Argentina

Presencial
ARS 63.416.000 - 102.674.000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+2
Senior Site Reliability Engineer, AI-DNA, $100k/year USD
Senior Site Reliability Engineer, AI-DNA, $100k/year USD

IgniteTech • Argentina

Presencial
ARS 2.000.000 - 4.000.000
Site Reliability Engineer - Senior Associate (Troubleshooting & Python)
Site Reliability Engineer - Senior Associate (Troubleshooting & Python)

JPMorgan Chase & Co. • Buenos Aires

Presencial
ARS 900.000 - 1.500.000
Senior Site Reliability Engineer/Platform Engineer
Senior Site Reliability Engineer/Platform Engineer

Techunting • Córdoba

Presencial
ARS 3.500.000 - 6.000.000
Software Engineer (DevOps) - Site Reliability
Software Engineer (DevOps) - Site Reliability

Revolut • Argentina

Presencial
Senior Site Reliability Engineer IRC302878
Senior Site Reliability Engineer IRC302878

GlobalLogic • Buenos Aires

Híbrido
ARS 8.928.000 - 13.392.000
Exciting projects
Collaborative environment
Work-life balance
+2
Site Reliability Engineer
Site Reliability Engineer

AgileEngine • Argentina

Híbrido
ARS 2.000.000 - 4.000.000
Professional growth
Competitive USD-based compensation
Flextime
+1
Senior Site Reliability Engineer IRC302878
Senior Site Reliability Engineer IRC302878

GlobalLogic • Argentina

Presencial
ARS 1.200.000 - 1.800.000
Exciting Projects
Collaborative Environment
Work-Life Balance
+2