Senior Site Reliability Engineer IRC304301

GlobalLogic

Argentina

Híbrido

ARS 1.800.000 - 3.000.000

Jornada completa

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Competitive salary
Family medical insurance
Extended paternity leave
Annual performance bonuses
Referral bonuses

Descripción de la vacante

GlobalLogic is seeking an experienced SRE/Platform Engineer to lead reliability initiatives for AI-first workloads. You will drive observability, CI/CD improvements, and platform patterns across multiple teams and cloud regions.

The role emphasizes mentorship, hands-on leadership, and strategic reliability initiatives in a fast-paced engineering environment with opportunities to work across global centers.

Formación

  • 6-8 years in SRE, Platform Engineering, or DevOps in senior or technical leadership roles.
  • Deep expertise in AWS (EKS, Lambda, CloudWatch) and multi-region architecture.
  • Proficiency in Terraform and GitOps practices.
  • Advanced operational proficiency with Datadog (dashboards, tracing, and SLO management).
  • Solid experience with Docker and Kubernetes.
  • Experience with CI/CD tools (Bitbucket Pipelines or GitHub Actions) and internal developer platforms like Backstage or Compass.
  • Proficiency in Python and/or Bash scripting.
  • Ability to independently lead strategic reliability initiatives.
  • A 'first principles' approach to resolving complex technical issues.
  • Strong commitment to mentoring junior and intermediate engineers.
  • Excellent verbal and written communication skills to engage stakeholders across teams.

Responsabilidades

  • AI Reliability & Operations: Manage SLOs, SLIs, and error budgets; design patterns for LLM observability and recovery.
  • Engineering Enablement: Serve as SRE liaison for software and AI teams; optimize CI/CD and promote standard templates.
  • Observability & Service Catalog: Maintain Datadog telemetry and evolve internal service catalog for visibility.
  • FinOps & IaC: Oversee Terraform/GitOps deployments and optimize AWS cloud costs.
  • Strategic Growth: Lead cross-team reliability initiatives and mentor team members.

Conocimientos

SRE experience
Platform engineering
DevOps
AWS
EKS
Lambda
CloudWatch
Terraform
GitOps
Datadog
Docker
Kubernetes
CI/CD
Backstage
Compass
Python
Bash scripting
Leadership
Mentoring
Communication

Herramientas

AWS
Terraform
GitOps
Datadog
Docker
Kubernetes
Backstage
Compass
Python
Bash

Descripción del empleo

Description

The project is building the reliability and AI operations foundation for its next chapter-an AI-first intelligence platform that runs the most demanding semiconductor intelligence workflows in the world. The SRE team operates as a technical leader within our engineering organization, responsible for defining reliability patterns for AI agent pipelines, architecting observability, and building an Internal Developer Platform (IDP). We provide a fast-scaling environment where reliability is prioritized from day one, offering the opportunity to apply deep SRE expertise to cutting-edge AI workloads and agentic systems.

The project is building the reliability and AI operations foundation for its next chapter-an AI-first intelligence platform that runs the most demanding semiconductor intelligence workflows in the world. The SRE team operates as a technical leader within our engineering organization, responsible for defining reliability patterns for AI agent pipelines, architecting observability, and building an Internal Developer Platform (IDP). We provide a fast-scaling environment where reliability is prioritized from day one, offering the opportunity to apply deep SRE expertise to cutting-edge AI workloads and agentic systems.

Requirements
  • Experience: 6-8 years in SRE, Platform Engineering, or DevOps in senior or technical leadership roles.
  • Cloud & Infrastructure: Deep expertise in AWS (EKS, Lambda, CloudWatch) and multi-region architecture.
  • Infrastructure as Code: Proficiency in Terraform and GitOps practices.
  • Observability: Advanced operational proficiency with Datadog (dashboards, tracing, and SLO management).
  • Containerization: Solid experience with Docker and Kubernetes.
  • CI/CD & IDP: Experience with CI/CD tools (Bitbucket Pipelines or GitHub Actions) and internal developer platforms like Backstage or Compass.
  • Automation: Demonstrated proficiency in Python and/or Bash scripting.
  • Leadership: Ability to independently lead strategic reliability initiatives.
  • Problem Solving: A "first principles" approach to resolving complex technical issues.
  • Mentorship: A strong commitment to mentoring junior and intermediate engineers.
  • Communication: Excellent verbal and written communication skills to engage stakeholders across engineering, product, and leadership teams.
Job responsibilities
  • AI Reliability & Operations: Manage SLOs, SLIs, and error budgets. Design and implement patterns for LLM observability and recovery to ensure the reliability of AI-agent pipelines.
  • Engineering Enablement: Act as the primary SRE liaison for software and AI teams. Optimize CI/CD pipelines and drive the adoption of "golden path" templates and SRE best practices across the organization.
  • Observability & Service Catalog: Maintain Datadog for telemetry and performance tracing of pipeline services. Manage and evolve the internal service catalog to provide clear visibility into architecture.
  • FinOps & Infrastructure Management: Oversee infrastructure-as-code deployments (Terraform/GitOps) while actively managing and optimizing cloud infrastructure costs in AWS.
  • Strategic Growth: Lead cross-team reliability initiatives, mentor team members, and ensure the engineering culture prioritizes platform robustness and technical excellence.
What we offer

Exciting Projects: Come take your place at the forefront of digital transformation! With clients across all industries and sectors, we offer an opportunity to work on market-defining products using the latest technologies.

Collaborative Environment: Expand your skills by collaborating with a diverse team of highly talented people in an open, laidback environment - or even abroad in one of our global centers or client facilities!

Work-Life Balance:GlobalLogic prioritizes work-life balance, which is why we offer flexible work schedules.We offer you the best quality of work life so that you exceed the expectations of our clients, while achieving your professional and personal ambitions.

Professional Development:Our dedicated Learning & Development team regularly organizes English classes, professional certifications, and technical and soft skill trainings. We also offer the chance to travel internationally

Excellent Benefits: We provide our employees with competitive salaries, family medical insurance, extended paternity leave, annual performance bonuses, and referral bonuses.

About GlobalLogic

GlobalLogic, a Hitachi Group Company, is a trusted digital engineering partner to the world’s largest and most forward-thinking companies. Since 2000, we’ve been at the forefront of the digital revolution - helping create some of the most innovative and widely used digital products and experiences. Today we continue to collaborate with clients in transforming businesses and redefining industries through intelligent products, platforms, and services.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer IRC304301
Senior Site Reliability Engineer IRC304301

t2s - Group International . your partner in executive search • Argentina

Presencial
ARS 2.400.000 - 4.200.000
Principal DevOps AI IRC303057
Principal DevOps AI IRC303057

GlobalLogic • Municipio de Rincón de los Sauces

Presencial
ARS 3.000.000 - 6.000.000
Competitive salary
Global exposure and travel
Professional development
+1
Senior SRE: AI Ops & Platform Reliability Lead
Senior SRE: AI Ops & Platform Reliability Lead

GlobalLogic • Argentina

Híbrido
ARS 1.800.000 - 3.000.000
Competitive salary
Family medical insurance
Extended paternity leave
+2
Senior SRE - AI Reliability & Platform Engineer
Senior SRE - AI Reliability & Platform Engineer

t2s - Group International . your partner in executive search • Argentina

A distancia
ARS 2.400.000 - 4.200.000
SR Advisory Consultant
SR Advisory Consultant

Globallogic • Buenos Aires

Híbrido
ARS 135.820.000 - 226.367.000
Competitive salaries
Family medical insurance
Extended paternity leave
+2
.Net Developer
.Net Developer

GlobalLogic • Argentina

Presencial
ARS 1.800.000 - 3.000.000
SSR CRA Consultant
SSR CRA Consultant

Hitachi, Ltd. • Buenos Aires

Presencial
ARS 106.626.000 - 167.555.000
Competitive salaries
Family medical insurance
Extended paternity leave
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Argentina

Presencial
ARS 135.892.000 - 196.289.000
Flexible work format
Education reimbursement
Mentorship program
+2
SAS Sys Admin IRC304136
SAS Sys Admin IRC304136

GlobalLogic Inc. • Argentina

Presencial
ARS 1.800.000 - 2.800.000
Family medical insurance
Annual performance bonuses
Referral bonuses
+2
Fullstack .NET/Phyton IRC303631
Fullstack .NET/Phyton IRC303631

GlobalLogic Inc. • Argentina

Presencial
ARS 90.697.000 - 181.395.000
Exciting projects
Collaborative environment
Work-life balance
+2