Lead DevOps Engineer

Jobtailor

São Paulo

Presencial

BRL 180 000 - 320 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Jobtailor seeks an experienced Infrastructure Engineer to lead the architecture and reliability of AI-powered systems at scale. You will own SLO/SLI design, observability, and resilient deployment patterns across cloud environments.

Responsibilities include mentoring engineers, enforcing standards, building secure distributed infra with Terraform and Kubernetes, and shaping the culture of reliability with a focus on drift, latency, and cost management.

Qualificações

  • Extensive infra engineering experience combining DevOps and SRE at scale.
  • Deep GCP expertise (AWS a strong plus) and relevant cloud certifications.
  • Production experience with SRE fundamentals: SLO/SLI design, error budgets, toil reduction.
  • Strong background in distributed systems failure modes and resilience patterns.
  • Expert‑level infrastructure‑as‑code (Terraform), container orchestration (Kubernetes), and CI/CD.
  • Hands‑on with modern observability stacks (OpenTelemetry, Sentry) and AI‑specific observability tooling.
  • Experience with API management platforms, particularly Apigee and Cloud Run.
  • Comfort working across Python, JavaScript, and Bash for infra tooling.
  • Strong spoken and written communication in English with teams and stakeholders.

Responsabilidades

  • Lead the architecture and maintenance of the infrastructure and reliability practices that keep AI‑powered systems performant, observable, and trustworthy under real production load, including redundancy, latency, and cost management.
  • Help define SLOs/SLIs for AI‑powered services, including latency and quality SLOs for LLM inference paths, and build the error‑budget discipline that lets product teams ship fast without breaking trust.
  • Design scalable, secure infrastructure for distributed AI services, event‑driven workloads, and multi‑LLM‑provider integrations.
  • Build metrics, tracing, and alerting that surface not just 'is it up' but 'is it behaving correctly' for LLM‑powered features (drift, regression, hallucination rates, tool‑call failures).
  • Define and enforce PRR‑style standards across teams launching new AI products and features.
  • Mentor engineers, drive architecture reviews, and shape the broader engineering culture around reliability.

Conhecimentos

DevOps
SRE
GCP
Terraform
Kubernetes
CI/CD
Observability
AI observability
Python
JavaScript
Bash
English fluency
APIs

Formação académica

Bachelor's degree in Computer Science

Ferramentas

Kubernetes
Terraform
Sentry
OpenTelemetry
Apigee
Cloud Run

Descrição da oferta de emprego

Responsibilities
  • Lead the architecture and maintenance of the infrastructure and reliability practices that keep AI‑powered systems performant, observable, and trustworthy under real production load, including redundancy, latency, and cost management.
  • Help define SLOs/SLIs for AI‑powered services, including latency and quality SLOs for LLM inference paths, and build the error‑budget discipline that lets product teams ship fast without breaking trust.
  • Design scalable, secure infrastructure for distributed AI services, event‑driven workloads, and multi‑LLM‑provider integrations.
  • Build metrics, tracing, and alerting that surface not just "is it up" but "is it behaving correctly" for LLM‑powered features (drift, regression, hallucination rates, tool‑call failures).
  • Define and enforce PRR‑style standards across teams launching new AI products and features.
  • Mentor engineers, drive architecture reviews, and shape the broader engineering culture around reliability.
Requirements
  • Significant infrastructure engineering experience combining DevOps and SRE disciplines at scale.
  • Deep GCP expertise (AWS a strong plus); relevant cloud certifications welcome.
  • Production experience with SRE fundamentals: SLO/SLI design, error budgets, toil reduction, blameless incident review.
  • Strong background in distributed systems failure modes and resilience patterns.
  • Expert‑level infrastructure‑as‑code (Terraform), container orchestration (Kubernetes), and CI/CD.
  • Hands‑on with modern observability stacks (i.e., OpenTelemetry, Sentry) and AI‑specific observability tooling (Arize, LangSmith, Braintrust, or similar).
  • Experience with API management platforms, particularly Apigee and Cloud Run.
  • Comfort working across Python, Javascript, and Bash for infra tooling.
  • Strong spoken and written communication in English with teams and stakeholders.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

AI/ML Ops Engineer
AI/ML Ops Engineer

99x • São Paulo

Presencial
BRL 120 000 - 190 000
Senior AI Engineer
Senior AI Engineer

Amplify IT • Brasil

Presencial
BRL 343 000 - 443 000
Lead AI Engineer
Lead AI Engineer

EPAM Systems • Brasil

Presencial
BRL 250 000 - 520 000
AI Engineer
AI Engineer

Jobtailor • Barueri

Presencial
BRL 180 000 - 260 000
AI Engineer
AI Engineer

Factspan • Brasil

Presencial
BRL 120 000 - 180 000
SRE Specialist, Agro Retail Community
SRE Specialist, Agro Retail Community

Jobtailor • São Paulo

Presencial
BRL 350 000 - 500 000
AI Engineer
AI Engineer

Jobtailor • Belo Horizonte

Presencial
BRL 180 000 - 280 000
Senior AI Engineer
Senior AI Engineer

Quartile • Brasil

Presencial
BRL 180 000 - 240 000
Senior GenAI Engineer
Senior GenAI Engineer

EPAM Systems • Brasil

Presencial
BRL 403 000 - 606 000
International projects with top brands
Paid time off and sick leave
Upskilling and certification courses
+2
Data Engineer - Oracle to PostgreSQL Re-Platform ( 102-08SENG-02 )
Data Engineer - Oracle to PostgreSQL Re-Platform ( 102-08SENG-02 )

Cloudary • Acre

Presencial
BRL 240 000 - 360 000