Junior SRE Engineer Guadalajara Mexico

ReHire, LLC

Región Centro

Presencial

MXN 1.080.000 - 1.621.000

Jornada completa

Hace 7 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha a medida para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Descripción de la vacante

ReHire, LLC is seeking a Junior SRE Engineer to join a reliability engineering team embedded with a regulated fintech client on AWS. You will contribute to AI-driven incident detection, runbooks, and end-to-end incident response in a multi-service environment.

You will work under senior engineers, build monitoring as code, and advance in a growth-focused, AI-enabled reliability program. Excellent English and cloud experience are required.

Formación

  • 1–3 years in SRE, DevOps, cloud infrastructure, or production support.
  • Advanced English required (oral and written).
  • Hands-on AWS across compute, networking, storage, and managed databases.
  • Observability tools: Datadog preferred; others are relevant.
  • Foundational SLI/SLO and error-budget concepts.
  • Containers/orchestration (Docker, ECS, Kubernetes) and serverless.
  • Infrastructure as Code (Terraform preferred).
  • Scripting in Python or Bash; REST APIs and JSON parsing.
  • Linux troubleshooting and networking basics.
  • Git/PR workflows and CI/CD tools.
  • Incident management and on-call concepts.
  • Fintech/enterprise software experience is a plus.
  • PCI-DSS/SOC 2/SOX/GLBA awareness is a plus.

Responsabilidades

  • Support incident response as shadow/secondary responder using AI-detected signals.
  • Build incident timelines from metrics/logs/traces; document postmortems.
  • Operate AI SRE sub-agents and flag anomalies for review.
  • Develop Datadog monitors, dashboards, and SLOs as code; review thresholds.
  • Reduce alert noise with AI/ML-based detection and monitor hygiene reviews.
  • Maintain AWS workloads (ECS/Fargate, EKS, Lambda, RDS/Aurora, etc.).
  • Identify capacity, cost anomalies; attribute telemetry to owning teams.
  • Write automation in Python/Bash against platform APIs; contribute Terraform modules.
  • Integrate reliability controls and AI checks into CI/CD pipelines.
  • Participate in architecture, reliability, and AI-risk reviews in regulated environments.

Conocimientos

SRE mindset
Python scripting
Bash scripting
Observability concepts
CI/CD workflows
Git & PR workflows
On-call concepts
English proficiency

Educación

Bachelor's degree in Computer Science, Engineering, or a related field

Herramientas

Datadog
Grafana
Prometheus
New Relic
CloudWatch
ELK/OpenSearch
Terraform
Ansible
CloudFormation
Docker
Kubernetes
AWS
GitHub Actions
Jenkins
ArgoCD
PagerDuty/Opsgenie

Descripción del empleo

Role Overview

At Rehire, we are partnering with a US-based data engineering and cloud technologies company to find a Junior SRE Engineer to join its Reliability Engineering team. You will be embedded with the SRE function of a financial services client, supporting a regulated consumer-lending platform on AWS with hundreds of microservices and event-driven pipelines, where reliability directly impacts customer trust and compliance.

This is a unique opportunity to grow in an environment where AI is already part of day-to-day operations, including an AI SRE co-pilot, purpose-built AI agents for incident triage and monitoring, and a formal AI governance program. Working under the guidance of senior engineers and architects, you will help operate, improve, and learn from this AI-driven reliability program.

Key Responsibilities
  • Support incident response as a shadow or secondary responder, using AI-driven detection, correlation, and root-cause analysis tools.
  • Help build incident timelines from metrics, logs, traces, and deploy history, and document postmortems and corrective actions through to closure.
  • Help operate and monitor AI SRE sub-agents (incident summarization, monitor-gap detection, usage attribution), flagging anomalies for senior review.
  • Build and maintain Datadog monitors, dashboards, and SLO definitions as code using Terraform, and support SLI/SLO and error-budget reviews for critical customer journeys.
  • Help reduce alert noise with AI/ML-assisted detection (anomaly, outlier, and forecast monitors, dynamic thresholds) and run recurring monitor-hygiene reviews.
  • Support the day-to-day reliability of AWS workloads (ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS, Step Functions).
  • Identify capacity, saturation, and cloud cost anomalies, and help attribute spend and telemetry volume to owning teams and services.
  • Write automation in Python and Bash against platform APIs (Datadog, AWS, GitHub, PagerDuty, Jira) and contribute Terraform modules through pull requests.
  • Help integrate reliability controls and AI-assisted checks into CI/CD pipelines, and create runbooks progressively automated toward self-healing.
  • Participate in architecture, reliability, and AI-risk reviews, learning how compliance frameworks (PCI-DSS, SOC 2, SOX, GLBA) apply in a regulated environment.
Requirements
  • Bachelor's degree in Computer Science, Engineering, or a related field.
  • 1–3 years of experience in SRE, DevOps, cloud infrastructure, platform, or production-support engineering.
  • Advanced English (oral and written). REQUIRED
  • Hands-on exposure to AWS (or an equivalent hyperscaler) across compute, networking, storage, and managed database services.
  • Exposure to at least one observability platform (Datadog preferred; Grafana/Prometheus, New Relic, CloudWatch, or ELK/OpenSearch also relevant).
  • Foundational understanding of SLI, SLO, and error-budget concepts.
  • Foundational knowledge of containers and orchestration (Docker, ECS, or Kubernetes) and serverless execution models.
  • Beginner-to-intermediate experience with Infrastructure as Code (Terraform preferred; Ansible or CloudFormation acceptable).
  • Scripting experience in Python, Bash, or similar, including consuming REST APIs and parsing JSON.
  • Basic Linux troubleshooting and networking fundamentals (DNS, TLS, load balancing, timeouts, and retries).
  • Comfort with Git, pull-request workflows, and CI/CD tools (GitHub Actions, Jenkins, GitLab CI, ArgoCD, or similar).
  • Familiarity with incident management and on-call concepts (severity models, escalation policies, PagerDuty or Opsgenie).
  • Experience in product engineering services, enterprise software, or fintech is a plus.
  • Awareness of compliance frameworks (PCI-DSS, SOC 2, SOX, GLBA) is a plus.
Key Competencies
  • Automation mindset: you would rather automate a task the second time you do it than the tenth.
  • Good judgment to escalat e early instead of sitting on an uncertain production signal.
  • Clear written communication: you can explain an incident, a metric, or a trade-off to someone who was not in the room.
  • Curiosity about LLM-based assistants and agents applied to operations, and about how to verify that their output is correct.
Preferred Certifications (not required)
  • AWS Certified Cloud Practitioner or an Associate-level AWS certification.
  • HashiCorp Certified: Terraform Associate.
  • Datadog Fundamentals or an equivalent observability certification.
  • Certified Kubernetes Administrator (CKA) or KCNA.
About the Position
  • Work Schedule: US shift presential at Guadalajara, Mexico.
  • Work Modality: Full-time contractor basis.
  • Competitive Salary Paid in USD.
  • Work Environment: Dynamic and collaborative.
  • Professional Growth: Hands-on learning in AI-driven SRE practices and opportunities for career advancement.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Junior SRE Engineer - AI-Driven Reliability (Guadalajara)
Junior SRE Engineer - AI-Driven Reliability (Guadalajara)

ReHire, LLC • Región Centro

Presencial
MXN 1.080.000 - 1.621.000
Lead SRE Engineer
Lead SRE Engineer

Cloudsufi • Región Centro

Presencial
MXN 900.000 - 1.400.000
Site Reliability Engineering (SRE) Lead - 2770
Site Reliability Engineering (SRE) Lead - 2770

Xideral • Región Centro

Presencial
MXN 700.000 - 900.000
Attractive Salary
Performance bonuses
SGMM Medical insurance
Senior SRE Lead: Reliability, Observability & AI Ops
Senior SRE Lead: Reliability, Observability & AI Ops

Cloudsufi • Región Centro

Presencial
MXN 900.000 - 1.400.000
Sr SRE Engineer
Sr SRE Engineer

Dresden Partners • Región Centro

Presencial
MXN 600.000 - 900.000
Senior AWS Site Reliability Engineers - 2850
Senior AWS Site Reliability Engineers - 2850

Xideral • Región Centro

Híbrido
MXN 1.200.000 - 1.500.000
Premium Benefits
Performance bonuses
SGMM Medical insurance
Site Reliability Engineer ID45689
Site Reliability Engineer ID45689

AgileEngine • Rosarito

Presencial
MXN 1.531.000 - 2.212.000
Mentorship and TechTalks
Competitive USD-based compensation
Work on modern solutions
+1
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

Presencial
MXN 1.433.948 - 1.792.436
Senior Site Reliability Engineer
Senior Site Reliability Engineer

adlytics GmbH • Región Centro

Presencial
MXN 900.000 - 1.300.000
Sueldo acorde a experiencia
Esquema 100% nómina
Prestaciones de Ley
+5
Senior SRE Engineer: Cloud Reliability & AI Ops Guadalajara
Senior SRE Engineer: Cloud Reliability & AI Ops Guadalajara

Dresden Partners • Región Centro

Presencial
MXN 600.000 - 900.000