Site Reliability Engineer

KI people

México

Híbrido

MXN 745.156 - 1.117.734

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Payroll
Direct hire by client
Multicultural teams
Permanent project

Descripción de la vacante

A leading company in the Human Resources Services sector is seeking a Site Reliability Engineer to work in hybrid mode from Guadalajara, Monterrey, or Mexico City. The role involves providing support for B2B applications, identifying issues, and implementing solutions within a multicultural team. A Bachelor's degree and 3+ years in IT operations are essential for this mid-senior level position.

Formación

  • 3+ years of experience in IT Operations SRE platform/Service Cloud operations.
  • Ability to investigate application code for debugging.
  • Familiarity with CI/CD pipelines and Cloud platforms.

Responsabilidades

  • Work with L2 support team on Root Cause Analyses (RCAs).
  • Identify and implement proactive health checks.
  • Collaborate on SRE orchestration framework.

Conocimientos

IT operations experience
Analytical skills
Proactive issue identification
Troubleshooting
Understanding of application architecture
Scripting proficiency

Educación

Bachelor’s degree in Computer Science

Herramientas

AppDynamics
ELK Stack
FullStory
Prometheus
Grafana

Descripción del empleo

18 hours ago Be among the first 25 applicants

Direct message the job poster from KI people

In Search of the Best Global IT & Digital Talent

We are looking for a Site Reliability Engineer to work on hybrid mode from GDL, MTY o CDMX for a multicultural project with stability and growth in the short, medium and long term.

Role Overview:

  • The SRE Operations specialist focuses on B2B applications support providing round the clock support to identify self healing automation and proactive health checks.
  • They need to be specialized in Site Reliability Engineering (SRE) mode of operations and help to onboard applications to any SRE Orchestration framework for higher business resiliency.
  • The resource needs to have strong IT operations experience, analytical skills and a mindset of proactive issue identification.
  • This resource champions Site Reliability Engineering and collaborates with the customer and the business to troubleshoot issues to identify the root cause and opportunities for automation/proactive health checks.
  • This resource should be able to investigate application code as needed.
  • The SRE Ops needs to have good understanding of different architecture types - legacy/modern app and their logging mechanisms and got exposure to observability tools like APPD, ELK, FullStory etc.
  • SRE Ops should be responsible as pro-active support engineer, diagnosing any anomalies and driving the necessary remediations across the teams involved.
  • SRE Ops resource will work with existing L2 support team, understand production issues, participate & contribute to RCA.
  • SRE Ops will identify gaps in proactive health checks, automate and implement self healing mechanism wherever needed and work with SRE orchestration team to bring readiness to on board SRE orchestration framework.

Qualifications Basic

  • Bachelor’s degree in computer science or related field.
  • 3+ years related experience in IT Operations SRE platform/Service Cloud operations.

Responsibilities:

  • Work with the existing L2 support team to understand production issues and actively participate in and contribute to Root Cause Analyses (RCAs).
  • Identify gaps in proactive health checks and implement new checks to detect potential issues before they impact production.
  • Automate and implement self-healing mechanisms wherever needed to minimize manual intervention and improve system resilience.
  • Collaborate with the SRE orchestration team to onboard and operationalize the SRE orchestration framework.
  • Diagnose anomalies in production environments and drive the necessary remediations across the teams involved.

Mandatory Skills:

  • Proven IT operations experience, with a focus on production support.
  • Strong analytical and problem-solving skills, with the ability to troubleshoot complex issues to identify root causes.
  • A mindset of proactive issue identification and prevention.
  • Ability to investigate application code (e.g., debugging, log analysis) to understand system behavior.
  • Understanding of different application architecture types (legacy and modern) and their logging mechanisms.
  • Exposure to observability tools such as AppDynamics (APPD), ELK Stack (Elasticsearch, Logstash, Kibana), and FullStory.
  • IT operations experience, analytical skills. A mindset of proactive issue identification.
  • Troubleshoot issues to identify the root cause and opportunities for automation/proactive health checks. Able to investigate application code as needed.
  • Understanding of different architecture types - legacy/modern app and their logging mechanisms
  • Exposure to observability tools like APPD, ELK, FullStory, Prometheus, Grafana
  • Responsible as pro-active support engineer, diagnosing any anomalies and driving the necessary remediations across the teams involved.
  • Proficiency in scripting

Nice-to-Have Skills:

  • Knowledge in Cloud platform –Azure/GCP
  • Knowledge in SQL
  • Exposure to CI/CD pipelines
  • Networking concepts to diagnose the issue
  • Experience with SRE (Site Reliability Engineering) principles and practices.
  • Experience with SRE orchestration frameworks.
  • Knowledge of scripting languages (e.g., Python, Bash, PowerShell) for automation.
  • Experience with containerization technologies (e.g., Docker, Kubernetes).
  • Knowledge of infrastructure-as-code tools (e.g., Terraform, Ansible).
  • Experience with CI/CD pipelines.
  • Excellent communication and collaboration skills.

Other Relevant Experience

  • Experience working as part of a SRE Operations team practicing SRE orchestration framework.
  • Experience and desire to work in a Global delivery environment
  • Ability to work in team in diverse/ multiple stakeholder environment

Offer:

  • Payroll
  • Direct hire by client
  • Multicultural teams
  • Perm project

If you are looking for a new professional challenge, this is a good opportunity, let's talk about your next professional experience.

Seniority level
  • Seniority level
    Mid-Senior level
Employment type
  • Employment type
    Full-time
Job function
  • Job function
    Engineering and Information Technology
  • Industries
    Human Resources Services

Referrals increase your chances of interviewing at KI people by 2x

Sign in to set job alerts for “Site Reliability Engineer” roles.
Site Reliability Engineer - Remote Work | REF#180173
Senior Site Reliability / Gitops Engineer
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics
Software Engineer (Python/Linux/Packaging)
Junior Software Development Engineer in Test / R+D - Remote Work | REF#271058
Python and Kubernetes Software Engineer - Data, Workflows, AI/ML & Analytics
Software Development Engineer in Test - Remote Work | REF#254437
Python Software Engineer - Ubuntu Hardware Certification Team
Software Engineer - Solutions Engineering
Golang System Software Engineer - Containers / Virtualisation
Software Engineer for AI Training (Code Quality & Debugging Focus)
Distributed Systems Software Engineer, Python / Go
Junior Software Engineer - Cross-platform C++ - Multipass
Software Engineer, Ceph & Distributed Storage

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Python and Kubernetes Software Engineer - Data, AI/ML & Analytics
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics

Canonical • Santiago de Querétaro

A distancia
MXN 557.517 - 929.195
Personal learning and development budget of USD 2,000 per year
Annual compensation review
Maternity and paternity leave
+2
Distributed Systems Software Engineer, Python / Go
Distributed Systems Software Engineer, Python / Go

Canonical • Ciudad Juárez

A distancia
MXN 1.115.000 - 1.859.000
Personal learning and development budget of USD 2,000 per year
Annual compensation review
Recognition rewards
+3
Polarion Software Developer/ 100% Remote in Mexico
Polarion Software Developer/ 100% Remote in Mexico

Pyramid Consulting, Inc • México

A distancia
MXN 1.115.000 - 1.487.000
Distributed Systems Software Engineer, Python / Go
Distributed Systems Software Engineer, Python / Go

Canonical • Santiago de Querétaro

A distancia
MXN 743.000 - 1.301.000
Personal learning and development budget of USD 2,000 per year
Annual compensation review
Maternity and paternity leave
+2
Distributed Systems Software Engineer, Python / Go
Distributed Systems Software Engineer, Python / Go

Canonical • Monterrey

A distancia
MXN 1.115.000 - 1.859.000
Personal learning and development budget of USD 2,000 per year
Annual compensation review
Recognition rewards
+4
Python Engineer (Middle) ID39656
Python Engineer (Middle) ID39656

AgileEngine • Región Centro

A distancia
MXN 932.000 - 1.306.000
Professional growth opportunities
Competitive compensation
Flexible working hours
Junior Python Engineer - Remote Work | REF#283520
Junior Python Engineer - Remote Work | REF#283520

BairesDev • Región Centro

A distancia
100% remote work
Excellent compensation in USD
Hardware and software setup
+3
Distributed Systems Software Engineer, Python / Go
Distributed Systems Software Engineer, Python / Go

Canonical • Tijuana

A distancia
MXN 929.000 - 1.487.000
Personal learning and development budget of USD 2,000 per year
Annual compensation review
Maternity and paternity leave
+2
Data Engineer Snowflake / Databricks
Data Engineer Snowflake / Databricks

Perficient • México

Presencial
MXN 60.000 - 90.000
Law benefits
Certifications
Internal & External Courses
+5
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics

Canonical • Tijuana

A distancia
MXN 650.436 - 1.115.034
Distributed work environment
Personal learning budget of $2,000
Performance-driven annual bonus
+4