Sr. Software Engineer

Petco

Santiago de Querétaro

Presencial

MXN 700.000 - 1.000.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Petco is seeking a senior Production Support Engineer to own and stabilize Python backend services, React Native apps, GraphQL APIs, PostgreSQL, and AWS infrastructure in a 24x7 environment.

You will lead incident investigations, perform root cause analyses, and guide on-call rotations while partnering with development teams to improve monitoring, automation, and reliability across distributed cloud-native systems.

Formación

  • Bachelor's degree or equivalent practical experience in CS/Engineering or related field.
  • 5+ years of experience supporting and troubleshooting large-scale production systems.
  • Hands-on experience with Python, GraphQL, PostgreSQL, React Native and AWS cloud services.
  • Proven track record diagnosing and resolving production incidents in distributed cloud-native environments.
  • Experience analyzing logs, metrics, traces, and monitoring data using Datadog, CloudWatch, or similar tools.
  • Advanced SQL skills to investigate data issues and optimize query performance.
  • Experience with AWS services such as Lambda, ECS, DynamoDB, API Gateway and S3.
  • Strong understanding of incident management, escalation, problem management, and RCA processes.
  • Experience delivering production fixes and automated solutions to reduce MTTR.
  • Ability to leverage AI-assisted tools to improve troubleshooting and automation.

Responsabilidades

  • Serve as a technical owner for production support across backend, mobile, and database services.
  • Lead investigation and resolution of production incidents and outages.
  • Analyze logs, metrics, traces, and DB activity to identify root causes and restore service.
  • Troubleshoot issues across app, infra, database, and network layers.
  • Perform log analysis with Datadog, CloudWatch and other observability tools.
  • Execute SQL queries to investigate data issues and DB performance.
  • Develop bug fixes, hotfixes, and permanent corrective actions.
  • Participate in on-call rotations and act as senior escalation point.
  • Coordinate incident response and communicate with stakeholders.
  • Conduct root cause analyses and document corrective actions and preventive measures.
  • Create and maintain runbooks, troubleshooting guides, and knowledge base articles.
  • Improve monitoring, alerting and observability to detect issues early.
  • Collaborate with development teams to prioritize defects and reliability improvements.
  • Support releases, deployments and change management activities.
  • Identify opportunities for automation to reduce manual work and MTTR.
  • Mentor engineers on best practices in production support.

Conocimientos

Python
GraphQL
PostgreSQL
React Native
AWS
Datadog
CloudWatch
SQL
Incident management
Root cause analysis

Educación

Bachelor's degree in Computer Science, Engineering, or related field

Herramientas

Lambda
ECS
DynamoDB
API Gateway
S3

Descripción del empleo

  • Serve as a technical owner for production support, ensuring the stability, availability, and performance of Python backend services, React Native mobile applications, GraphQL APIs, PostgreSQL databases, and AWS cloud infrastructure.
  • Lead the investigation and resolution of production incidents, service disruptions, performance degradation, and customer-reported issues.
  • Analyze application logs, system metrics, distributed traces, and database activity to quickly identify root causes and restore service.
  • Troubleshoot complex issues across application, infrastructure, database, network, and integration layers.
  • Perform detailed log analysis using Datadog, CloudWatch, and other observability tools to identify failure patterns, bottlenecks, and system anomalies.
  • Investigate data-related issues by executing SQL queries, analyzing database performance, validating data integrity, and troubleshooting PostgreSQL and DynamoDB-related problems.
  • Develop and deploy bug fixes, hotfixes, and permanent corrective actions to prevent incident recurrence.
  • Participate in on-call rotations and act as a senior escalation point for critical production incidents.
  • Coordinate incident response activities, provide timely stakeholder communication, and drive service restoration efforts.
  • Conduct root cause analysis (RCA) for major incidents and document findings, corrective actions, and preventive measures.
  • Create and maintain operational runbooks, troubleshooting guides, support documentation, and knowledge base articles.
  • Improve monitoring, alerting, logging, and observability capabilities to proactively detect issues before they impact customers.
  • Partner with development teams to prioritize production defects, technical debt remediation, and system reliability improvements.
  • Support application releases, production deployments, and change management activities, ensuring successful implementation and post-deployment validation.
  • Identify opportunities for automation to reduce manual support effort and improve operational efficiency.
  • Mentor engineers on production support best practices, troubleshooting techniques, and operational excellence.
Education / Experience
  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
  • 5+ years of experience supporting and troubleshooting large-scale production systems.
  • Strong hands-on experience with Python, GraphQL, PostgreSQL, React Native, and AWS cloud services.
  • Proven experience diagnosing and resolving production incidents in distributed cloud-native environments.
  • Strong experience analyzing application logs, metrics, traces, and monitoring data using Datadog, CloudWatch, or similar observability platforms.
  • Advanced SQL skills with the ability to investigate data issues, analyze query performance, and troubleshoot database-related incidents.
  • Experience supporting AWS services such as Lambda, ECS, DynamoDB, API Gateway, S3, and CloudWatch.
  • Strong understanding of incident management, escalation procedures, problem management, and root cause analysis processes.
  • Experience developing production bug fixes, hotfixes, and permanent corrective actions in a fast-paced environment.
  • Utilize AI-assisted development tools to improve troubleshooting efficiency, generate code fixes, automate repetitive support tasks, and reduce mean time to resolution (MTTR).
  • Excellent analytical, troubleshooting, and problem-solving skills with the ability to quickly isolate and resolve complex issues.
  • Strong verbal and written communication skills with the ability to communicate technical issues clearly to both technical and business stakeholders.
  • Experience working in a 24x7 production support environment and participating in on-call rotations.
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Full Stack Engineer
Full Stack Engineer

Luxoft • Estado de México

Presencial
MXN 893.000 - 1.251.000
Tier 3 Support Engineer
Tier 3 Support Engineer

Data2 • Americas

Presencial
MXN 1.121.000 - 1.467.000
Senior Production Reliability Engineer
Senior Production Reliability Engineer

Petco • Santiago de Querétaro

Presencial
MXN 700.000 - 1.000.000
Senior Full Stack Developer ID71007
Senior Full Stack Developer ID71007

AgileEngine, LLC. • Ciudad de México

Presencial
MXN 900.000 - 1.300.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
+3
Senior Full Stack Developer ID71007
Senior Full Stack Developer ID71007

AgileEngine, LLC. • Región Centro

Presencial
MXN 600.000 - 1.200.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
Senior Full Stack Engineer ID67824
Senior Full Stack Engineer ID67824

AgileEngine • Rosarito

Híbrido
MXN 1.041.000 - 1.562.000
Mentorship
Competitive compensation
Flexible schedule
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • Santiago de Querétaro

Híbrido
MXN 1.394.000 - 1.918.000
Professional growth programs
Competitive compensation
Exciting projects with top-tier clients
+1
Application Support Engineer Assoc Manager
Application Support Engineer Assoc Manager

Accenture México • Monterrey

Híbrido
MXN 1.200.000 - 1.800.000
Application Support Engineer Specialist
Application Support Engineer Specialist

Accenture México • Monterrey

Híbrido
MXN 800.000 - 1.200.000
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • Monterrey

Híbrido
MXN 1.568.000 - 2.091.000
Professional growth
Competitive compensation
Exciting projects
+1