Lead Operational Intelligence Engineer

EPAM Systems

Brasil

Presencial

BRL 300 000 - 600 000

Tempo integral

Há 7 dias
Torna-te num dos primeiros candidatos

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

EPAM Systems is seeking a Lead Operational Intelligence Engineer to own the development, maintenance, and enhancement of our Elastic & Observability Platform across cloud environments. You will drive strategic initiatives, guide a high-performing team, and ensure platform reliability with self-service capabilities for users.

You will participate in on-call rotations to monitor health, improve performance, and drive incident response while mentoring engineers and collaborating with stakeholders

Qualificações

  • 5+ years in Operational Intelligence with leadership in observability platforms.
  • Experience architecting Elastic clusters in multi-cloud environments.
  • Deep knowledge of Elasticsearch, Kibana, Logstash.
  • Proficient Infra-as-Code (Terraform, Ansible), Jenkins or GitOps.
  • Advanced Python scripting for automation and interoperability.
  • Strong incident management experience with PagerDuty, Uptrends, etc.
  • Fluency in English (B2+).

Responsabilidades

  • Oversee availability, functionality, performance, and security of observability platforms to meet SLAs.
  • Lead complex incident response during on-call rotations.
  • Document platform processes and provide knowledge sharing resources.
  • Collaborate with cross-functional teams, vendors, and stakeholders on requirements and upgrades.
  • Advance platform features and self-service capabilities, including AI tooling and data processing models.
  • Prototype platform innovations such as Kubernetes migrations and AI-driven observability.

Conhecimentos

Team leadership
Technical leadership
Python scripting
Incident management
English communication

Ferramentas

Terraform
Ansible
Kubernetes
ELK Stack

Descrição da oferta de emprego

We are looking for a highly experienced and dynamic Lead Operational Intelligence Engineer to join our team.

In this role, you will take ownership of leading the development, maintenance, and enhancement of our Elastic & Observability Platform deployed across cloud environments. You will drive strategic initiatives, guide a high-performing technical team, and ensure platform reliability while fostering innovation and enabling self-service capabilities for platform consumers. This position also involves participating in an on-call rotation to oversee platform health and functionality.

Responsibilities
  • Oversee the availability, functionality, performance, and security of observability and search platforms to exceed business SLAs
  • Provide technical leadership during complex incidents and escalated resolutions promptly during on-call periods
  • Develop and maintain comprehensive platform documentation, standard operating procedures, and knowledge-sharing resources
  • Collaborate with cross-functional teams, stakeholders, and vendors to oversee operational requirements, drive strategic initiatives, and manage installations, troubleshooting, and upgrades
  • Lead the enhancement of platform features and self-service capabilities, including advanced Elastic Synthetics and chargeback automation
  • Architect and implement proof-of-concepts for platform innovation, such as AI-driven observability, advanced data processing models, or Kubernetes-based platform migration
  • Supervise the building, deployment, and maintenance of Elastic clusters using Infrastructure-as-Code tools like Terraform and Ansible, while mentoring team members on best practices
  • Oversee platform lifecycle management activities, including component upgrades, capacity planning, cost optimization, and evolving compliance requirements
  • Continuously assess and fine-tune ELK stack performance, including ingestion, indexing, and query optimization for large-scale environments
  • Establish and enhance comprehensive alerting and incident management workflows, integrating sophisticated monitoring tools such as Kibana Rules, Watchers, and PagerDuty
  • Supervise the ingestion, enrichment, backup, and restoration of large-scale platform data while optimizing data workflows
  • Lead and plan critical operational events such as SSL certificate rotations, cluster migrations, or scalability optimization projects
Requirements
  • 5+ years of experience in Operational Intelligence, with a proven track record of leadership and technical expertise in managing large-scale observability platforms
  • Demonstrated ability to architect and manage Elastic clusters in complex, multi-cloud environments
  • In-depth knowledge of Elastic Stack components, including advanced configurations of Elasticsearch, Kibana, and Logstash
  • Advanced proficiency in Infrastructure-as-Code tools like Terraform and Ansible, with demonstrated flexibility in adapting other tools like Jenkins CI or GitOps frameworks
  • Advanced Python scripting skills for automation, data processing, and extending platform interoperability
  • Deep understanding of incident management frameworks and workflows with tools like PagerDuty, Uptrends, and other enterprise monitoring solutions
  • Proven expertise in troubleshooting and resolving complex platform challenges under tight SLAs
  • Strong capability in managing and scaling fault-tolerant platforms while ensuring performance, security, and compliance across large distributed systems
  • Demonstrated ability to mentor and grow team members, manage priorities, and act as a bridge between technical and non-technical teams
  • Excellent command of English (B2+ level), both written and spoken, with a strong emphasis on technical communication skills
Nice to have
  • Expertise in scripting with Groovy or experience in advanced Linux administration to optimize platform processes
  • Track record of optimizing observability workflows with additional integrations or customizations in tools like Uptrends, PagerDuty, or Elastic features
  • Hands-on experience with advanced Elastic Synthetics setups for robust monitoring and custom synthetic testing frameworks
  • Experience driving strategic initiatives such as modernization through AI tooling, cloud-native transitions, or cost-saving observability optimizations

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Principal Search Consulting Architect
Principal Search Consulting Architect

Elastic • Brasil

Híbrido
BRL 450 000 - 750 000
Health coverage for you and family
Flexible locations and schedules
Volunteer time off
+1
Lead Data Platform Engineer
Lead Data Platform Engineer

EPAM Systems • Brasil

Presencial
BRL 260 000 - 380 000
Principal Search Consulting Architect
Principal Search Consulting Architect

PowerToFly • Brasil

Híbrido
BRL 509 000 - 714 000
Health coverage for you and your family
Flexible locations and schedules
Generous vacation days
Senior Reliability and Platform Engineer
Senior Reliability and Platform Engineer

Jobtailor • São Paulo

Presencial
BRL 180 000 - 300 000
Principal Solutions Architect, Security Specialist
Principal Solutions Architect, Security Specialist

WinsAbove • Brasil

Presencial
BRL 120 000 - 160 000
Competitive pay based on performance
Health coverage for you and your family
Flexible locations and schedules
+2
Lead PHP Software Engineer
Lead PHP Software Engineer

EPAM Systems • Brasil

Presencial
BRL 260 000 - 420 000
International projects with top brands
Global teams of diverse peers
Employee financial programs
+3
Principal Search Consulting Architect
Principal Search Consulting Architect

Elasticsearch B.V. • Brasil

Presencial
Competitive pay
Health coverage
Flexible locations and schedules
+3
Enterprise Account Executive - Sao Paulo
Enterprise Account Executive - Sao Paulo

Elastic • Brasil

Híbrido
BRL 200 000 - 280 000
Competitive pay
Health coverage for family
Flexible locations and schedules
+4
Senior DevOps Engineer - LATAM
Senior DevOps Engineer - LATAM

Neura Market • Brasil

Presencial
BRL 260 000 - 420 000
None
Cloud Operations Engineer
Cloud Operations Engineer

Tenarai - LATAM • São Paulo

Presencial
BRL 180 000 - 280 000