Principal Associate (Site Reliability Engineering)

Capital One

Ciudad de México

Híbrido

MXN 900.000 - 1.300.000

Jornada completa

hace 45 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Transforma esta oferta en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

Capital One is establishing a Site Reliability Engineering center in Mexico City with Principal Associate SREs to support payment-critical systems across Discover Network, Diners Club, and PULSE. You will join the first CDMX cohort to build reliability and automation that keeps millions of transactions flowing daily.

You will operate across hybrid on‑prem and AWS, developing automation in Python/Java, tuning Datadog dashboards, and driving incident response with on-call rotations.

Formación

  • Bachelor's degree in a relevant field.
  • 4+ years in DevOps or reliability engineering.
  • Experience with Java, Python, or Go.
  • 2+ years in cloud native tech (AWS, Azure, GCP).
  • 2+ years in container orchestration (Docker or Kubernetes).
  • Unix/Linux system administration experience.
  • Experience with shell scripting.

Responsabilidades

  • Build and maintain reliability tooling, dashboards, alerts, runbooks, and remediation scripts.
  • Develop automation solutions using Python, Java, and shell scripts.
  • Troubleshoot complex production issues across on-prem and AWS.
  • Contribute to observability with Datadog and Observe dashboards.
  • Support incident response and participate in on-call rotations.
  • Leverage AI tools to accelerate engineering workflows.
  • Manage secrets and certificates through automated rotation.

Conocimientos

English fluency
SRE / reliability engineering
DevOps
Python
Java
Go
Shell scripting
Unix/Linux

Educación

Bachelor's degree

Herramientas

Docker
Kubernetes
OpenShift
Datadog
Observe
HashiCorp Vault
Claude Code

Descripción del empleo

We're building a Site Reliability Engineering center in Mexico City and hiring Principal Associate SREs to join one of our founding teams. You'll work on payment-critical systems across the Discover Network, Diners Club International, and PULSE - contributing to settlement reliability, alert quality, observability, and automation that directly impacts millions of transactions daily.This is a ground-floor opportunity. You'll be part of the first cohort of engineers in CDMX, working alongside experienced SRE leaders to build the operational muscle that allows Mexico City to own reliability outcomes independently. Depending on team placement, you'll focus on one of the following areas:* Settlement - ensuring batch settlement cycles complete accurately, on time, and in compliance with regulatory requirements across domestic credit/debit and international cross-border networks* Alert Signal & Observability - reducing alert noise, building automated severity classification, and creating customer impact dashboards that make incident response faster and more decisive* Reliability Automation & Platform Convergence - building automated runbooks, driving Capital One platform adoption, and developing AI-powered remediation workflowsWhat You'll Do* Build and maintain reliability tooling - observability dashboards, automated alerts, runbooks, and remediation scripts that reduce toil and improve mean time to recovery* Develop automation solutions - using Python, Java, and shell scripting to eliminate manual operational processes, from certificate rotation to compliance artifact generation* Troubleshoot and debug complex production issues - diagnose failures across distributed systems spanning on-prem data centers and AWS, identify root causes, and implement durable fixes* Contribute to observability - configure and tune monitoring in Datadog and Observe, build dashboards that surface actionable signals, and reduce unactionable alert volume* Support incident response - participate in on-call rotations, respond to production incidents, drive diagnosis, and contribute to blameless postmortems* Leverage AI tools to accelerate engineering - use agentic AI automation (Claude Code and others) to develop solutions, generate runbook drafts, and build automation agents* Manage secrets and certificates - automate rotation and provisioning, ensuring security posture without manual toil* Deliver through CI/CD pipelines - build, test, and deploy automation via continuous integration and API automation frameworksWhat Success Looks Like* Independently troubleshooting and resolving production issues within your domain without escalation* At least one operational process fully automated and running in production* Contributing measurably to team OKRs - whether that's alert noise reduction, MTTR improvement, or settlement cycle reliability* Producing or improving runbooks and dashboards that your teammates and partner teams actively useThe EnvironmentYou'll work across hybrid on-prem and cloud infrastructure supporting real-time and batch financial transaction systems at global scale. The tech stack includes Python, Java, shell scripting, AWS, Kubernetes, OpenShift, CI/CD pipelines, and API automation frameworks. Observability runs on Datadog and Observe with extensive dashboard configuration. Secret management uses HashiCorp Vault. You'll use agentic AI tools (Claude Code and others) to develop automation solutions and accelerate your engineering output. The systems span three on-prem data centers and AWS, with both modern cloud-native services and legacy payment platforms. Strong troubleshooting and debugging skills are essential.Basic Qualifications* Professional English fluency* Bachelor's degree* Background in SRE, production operations, or reliability engineering* At least 4 years of experience in DevOps Engineering (internship experience does not apply)* 4+ years of experience in at least one of the following: Java, Python, Go* At least 2 years of experience with Cloud Native technologies (Amazon Web Services, Microsoft Azure, Google Cloud Platform)* 2+ years of experience with container orchestration services including Docker or Kubernetes* Experience with Shell or Bash scripting* At least 2 years of Unix or Linux system administration experiencePreferred Qualifications* Experience developing automation solutions using agentic AI tools (Claude Code, Copilot CLI)* Troubleshooting and debugging skills across distributed systems* Familiarity with payments, financial services, or other regulated high-availability domains* Knowledge or experience of Networking concepts (TCP/DNS/TLS)
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Sr. Manager SRE (Individual Contributor)
Sr. Manager SRE (Individual Contributor)

Capital One • Ciudad de México

Presencial
MXN 1.400.000 - 2.100.000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Capital One Group • Ciudad de México

Híbrido
MXN 1.200.000 - 1.800.000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Capital One National Association • Ciudad de México

Presencial
MXN 1.200.000 - 2.000.000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Military Friendly Company • Ciudad de México

Híbrido
MXN 1.200.000 - 1.800.000
Founding SRE Principal — CDMX, Reliability & Automation
Founding SRE Principal — CDMX, Reliability & Automation

Capital One • Ciudad de México

Híbrido
MXN 900.000 - 1.300.000
Principal Associate SRE: Payments Reliability & Automation
Principal Associate SRE: Payments Reliability & Automation

Capital One • Ciudad de México

Híbrido
MXN 1.400.315 - 1.750.394
Sr. SRE
Sr. SRE

Turtle Trax S.A. • Región Centro

Híbrido
MXN 900.000 - 1.500.000
Senior SRE Lead - Settlement Systems Reliability
Senior SRE Lead - Settlement Systems Reliability

Capital One • Ciudad de México

Presencial
MXN 2.100.472 - 2.625.591
Data Engineer
Data Engineer

Time To Hire • Ciudad de México

Híbrido
MXN 700.000 - 1.100.000
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

Presencial
MXN 1.433.948 - 1.792.436