Lead Operational Intelligence Engineer

EPAM Systems

Argentina

Presencial

ARS 1.200.000 - 2.400.000

Jornada completa

hace 6 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Transforma esta oferta en una entrevista: un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Connectivity Bonus
Medicina Prepaga
Paternity Leave
Discounts card
English Training
Training Program
Marriage bonus
Referral Program
External Agreements and Discounts
Vacations: 14 calendar days a year

Descripción de la vacante

EPAM Systems is seeking a Lead Operational Intelligence Engineer to own the development and maintenance of an Elastic & Observability Platform across cloud environments. This role drives strategic platform initiatives, guides a high-performing team, and ensures platform reliability while enabling self-service capabilities for consumers.

The position includes on-call responsibilities, collaboration with multiple teams and vendors, and ongoing modernization through AI tooling and scalable

Formación

  • 5+ years in Operational Intelligence with leadership in large-scale observability platforms.
  • Architect and manage Elastic clusters in multi-cloud environments.
  • Strong infrastructure-as-code experience (Terraform/Ansible).
  • Advanced Python scripting for automation and data processing.

Responsabilidades

  • Oversee availability, performance, and security of observability platforms to meet SLAs.
  • Provide leadership during incidents and drive rapid resolutions during on-call periods.
  • Develop platform documentation and SOPs for knowledge sharing.
  • Mentor team members and collaborate with cross-functional teams and vendors.

Conocimientos

Elastic Stack
Terraform
Ansible
Python
PagerDuty
Incident Management
Leadership
English Proficiency

Herramientas

Kibana
Elasticsearch
Logstash

Descripción del empleo

We are looking for a highly experienced and dynamic Lead Operational Intelligence Engineer to join our team. In this role, you will take ownership of leading the development, maintenance, and enhancement of our Elastic & Observability Platform deployed across cloud environments. You will drive strategic initiatives, guide a high-performing technical team, and ensure platform reliability while fostering innovation and enabling self-service capabilities for platform consumers. This position also involves participating in an on-call rotation to oversee platform health and functionality.

Responsibilities
  • Oversee the availability, functionality, performance, and security of observability and search platforms to exceed business SLAs
  • Provide technical leadership during complex incidents and expedite resolutions promptly during on-call periods
  • Develop and maintain comprehensive platform documentation, standard operating procedures, and knowledge-sharing resources
  • Collaborate with cross-functional teams, stakeholders, and vendors to oversee operational requirements, drive strategic initiatives, and manage installations, troubleshooting, and upgrades
  • Lead the enhancement of platform features and self-service capabilities, including advanced Elastic Synthetics and chargeback automation
  • Architect and implement proof-of-concepts for platform innovation, such as AI-driven observability, advanced data processing models, or Kubernetes-based platform migration
  • Supervise the building, deployment, and maintenance of Elastic clusters using Infrastructure-as-Code tools like Terraform and Ansible, while mentoring team members on best practices
  • Oversee platform lifecycle management activities, including component upgrades, capacity planning, cost optimization, and evolving compliance requirements
  • Continuously assess and fine-tune ELK stack performance, including ingestion, indexing, and query optimization for large-scale environments
  • Establish and enhance comprehensive alerting and incident management workflows, integrating sophisticated monitoring tools such as Kibana Rules, Watchers, and PagerDuty
  • Supervise the ingestion, enrichment, backup, and restoration of large-scale platform data while optimizing data workflows
  • Lead and plan critical operational events such as SSL certificate rotations, cluster migrations, or scalability optimization projects
Requirements
  • 5+ years of experience in Operational Intelligence, with a proven track record of leadership and technical expertise in managing large-scale observability platforms
  • Demonstrated ability to architect and manage Elastic clusters in complex, multi-cloud environments
  • In-depth knowledge of Elastic Stack components, including advanced configurations of Elasticsearch, Kibana, and Logstash
  • Advanced proficiency in Infrastructure-as-Code tools like Terraform and Ansible, with demonstrated flexibility in adapting other tools like Jenkins CI or GitOps frameworks
  • Advanced Python scripting skills for automation, data processing, and extending platform interoperability
  • Deep understanding of incident management frameworks and workflows with tools like PagerDuty, Uptrends, and other enterprise monitoring solutions
  • Proven expertise in troubleshooting and resolving complex platform challenges under tight SLAs
  • Strong capability in managing and scaling fault‑tolerant platforms while ensuring performance, security, and compliance across large distributed systems
  • Demonstrated ability to mentor and grow team members, manage priorities, and act as a bridge between technical and non-technical teams
  • Excellent command of English (B2+ level), both written and spoken, with a strong emphasis on technical communication skills
Nice to have
  • Expertise in scripting with Groovy or experience in advanced Linux administration to optimize platform processes
  • Track record of optimizing observability workflows with additional integrations or customizations in tools like Uptrends, PagerDuty, or Elastic features
  • Hands‑on experience with advanced Elastic Synthetics setups for robust monitoring and custom synthetic testing frameworks
  • Experience driving strategic initiatives such as modernization through AI tooling, cloud‑native transitions, or cost-saving observability optimizations
We offer
  • Connectivity Bonus (25,000 ARS are paid with a salary receipt at the end of each month as a non-wages concept).
  • Medicina Prepaga (It covers the collaborator and direct family group).
  • Paternity Leave (Two additional days are added to what is established by law, total of 4 days).
  • Discounts card.
  • English Training (English lessons, twice per week).
  • Training Program (Access to multiple customized training plans according to the needs of each role within the company).
  • Marriage bonus (The company doubles the allowance established by law that ANSES offers).
  • Referral Program (Referral bonus is paid when the referral of a collaborator joins the Company).
  • External Agreements and Discounts.
  • Vacations: 14 calendar days a year
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting‑edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead Data DevOps Engineer
Lead Data DevOps Engineer

EPAM Systems • Argentina

Presencial
ARS 2.000.000 - 3.200.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Lead AI Engineer
Lead AI Engineer

EPAM Systems • Argentina

Presencial
ARS 3.000.000 - 5.000.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Lead Data Analyst
Lead Data Analyst

EPAM Systems • Argentina

Presencial
ARS 1.200.000 - 1.800.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Lead Data Software Engineer (Python+AWS)
Lead Data Software Engineer (Python+AWS)

EPAM Systems • Argentina

Presencial
ARS 2.000.000 - 4.000.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Senior GenAI Engineering Lead - AI Delivery & Intelligent Systems
Senior GenAI Engineering Lead - AI Delivery & Intelligent Systems

EPAM Systems • Argentina

Presencial
ARS 4.500.000 - 7.500.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Senior Python Engineer
Senior Python Engineer

EPAM Systems • Argentina

Presencial
ARS 4.464.000 - 7.812.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Senior/Lead AI Engineer
Senior/Lead AI Engineer

EPAM Systems • Argentina

Presencial
ARS 1.000.000 - 2.000.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7
Lead Data Software Engineer (Python+Azure)
Lead Data Software Engineer (Python+Azure)

EPAM Systems • Argentina

Presencial
ARS 16.740.000 - 29.016.000
Connectivity Bonus (25,000 ARS paid at
Medicina Prepaga
Paternity Leave (Two additional days,
+7
Lead Data Software Engineer (Java+AWS)
Lead Data Software Engineer (Java+AWS)

EPAM Systems • Argentina

Presencial
ARS 1.500.000 - 2.500.000
Connectivity Bonus (ARS 25,000 monthly
Medicina Prepaga
Paternity Leave
+7
Lead Python Full Stack Engineer
Lead Python Full Stack Engineer

EPAM Systems • Argentina

Presencial
ARS 3.500.000 - 7.000.000
Connectivity Bonus
Medicina Prepaga
Paternity Leave
+7