Staff Engineer (Core & MLOps)

Lever, Inc.

Brasil

Teletrabalho

BRL 300 000 - 480 000

Tempo integral

Há 2 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Não envies um currículo genérico — gera um currículo e uma carta de apresentação adaptados a esta função específica.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Fully remote, remote‑first
Flexible hours
Global collaboration

Resumo da oferta

Lever, Inc. seeks a Staff Engineer (Core & MLOps) in Brazil. You will shape foundational infrastructure powering large‑scale web data products and distributed teams, owning core planes, service contracts, and multi‑cloud deployments. This role blends hands‑on architecture with leadership across squads.

You will guide reliability practices, mentor engineers, and drive platform strategy in a globally distributed, remote‑first environment, influencing project roadmaps and standards.

Qualificações

  • 10+ years building scalable distributed backends or core libraries.
  • Advanced Java with reactive frameworks and strong Python.
  • Deep gRPC/Protobuf experience and schema evolution.
  • Production Kubernetes at scale and Terraform usage.
  • Experience with telemetry pipelines and feature stores is a plus.
  • Strong reliability, SLO/SLI, and service contracts experience.
  • Excellent written communication for cross‑team alignment.
  • Remote‑first collaboration across time zones.

Responsabilidades

  • Architect and evolve core control and context planes for multi‑language clients.
  • Own service chassis, deployment pipelines, and standardized workloads.
  • Govern inter‑service contracts, including gRPC and API gateway policies.
  • Maintain platform infra across Kubernetes, Terraform, and Kafka.
  • Lead architectural strategy with RFDs for orchestration and routing.
  • Establish reliability practices, SLOs/SLIs, and fault tolerance standards.
  • Participate in on‑call rotations and post‑mortems, enabling platform improvements.
  • Mentor engineers and promote consistent, reliable software development.

Conhecimentos

Java
Python
gRPC
Kafka
Kubernetes
Terraform
Protobuf
Distributed systems
Reliability engineering
Technical writing

Ferramentas

Vert.x / Netty
Protocol Buffers
K8s tooling
CI/CD pipelines

Descrição da oferta de emprego

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in Brazil.

This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you’ll tackle complex distributed systems challenges at production scale. You’ll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands‑on architecture with technical leadership, mentoring, and cross‑functional alignment. In a globally distributed, remote‑first environment, you’ll have significant autonomy to solve challenging infrastructure problems and influence long‑term platform strategy.

Accountabilities:
  • Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health‑aware routing, automated canary releases, and operational feedback loops.
  • Own the service chassis and golden path, maintaining and improving multi‑language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
  • Define and govern inter‑service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
  • Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real‑time billing pipelines, Valkey, and database modernization initiatives.
  • Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi‑cluster routing, automated failover, and other critical platform initiatives.
  • Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
  • Participate in shared infrastructure on‑call rotations, lead incident post‑mortems, and convert operational insights into platform improvements.
  • Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
Requirements:
  • 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
  • Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
  • Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission‑critical systems.
  • Hands‑on production experience with Kubernetes at scale, Terraform, and event‑streaming platforms such as Kafka.
  • Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
  • Strong reliability engineering background, including SLO/SLI definition, blast‑radius analysis, fault tolerance, and rigorous service contracts.
  • Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
  • Strong written and interpersonal communication skills suited to a globally distributed, remote‑first environment.
  • A curious, continuous‑learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
  • Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
  • MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
  • Familiarity with zero‑trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
  • Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
  • Experience with large-scale web scraping or crawling, or contributions to distributed‑systems and data‑extraction open‑source projects, is advantageous.
Benefits:
  • Fully remote, remote‑first working environment with flexible working hours.
  • Freedom and flexibility to work from the location where you are most productive.
  • Opportunity to work on core infrastructure supporting large‑scale web data pipelines and distributed systems.
  • Exposure to cutting‑edge open‑source technologies, tools, and evolving AI and web data infrastructure.
  • Opportunities to attend conferences and connect with colleagues across the globe.
  • Collaboration with a diverse, multicultural, and globally distributed engineering community.
  • High level of autonomy and organizational trust.
  • Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.
Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

MLOps Field Engineer
MLOps Field Engineer

Jobgether SRL • Brasil

Presencial
BRL 180 000 - 240 000
Learning budget
Travel opportunities
Global team sprints
Senior DevOps & Infrastructure Engineer
Senior DevOps & Infrastructure Engineer

Jobgether SRL • Brasil

Teletrabalho
BRL 240 000 - 360 000
Fully remote opportunity for Brazil
Exposure to AI infrastructure
Cutting-edge tech projects
+4
Sr. Platform Engineer
Sr. Platform Engineer

Jobgether SRL • Brasil

Teletrabalho
BRL 250 000 - 450 000
Fully remote (Brazil)
Full-time
Competitive compensation
+7
Platform Engineer ID90030
Platform Engineer ID90030

AgileEngine, LLC. • Campinas

Presencial
BRL 622 000 - 829 000
Professional growth
Competitive USD compensation
A selection of exciting projects
+1
Platform Engineer ID90030
Platform Engineer ID90030

AgileEngine, LLC. • Curitiba

Presencial
BRL 622 000 - 933 000
Professional growth
Competitive compensation
A selection of exciting projects
+1
Platform Engineer ID90030
Platform Engineer ID90030

AgileEngine, LLC. • Brasília

Presencial
BRL 180 000 - 300 000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Platform Engineer ID90030
Platform Engineer ID90030

AgileEngine, LLC. • São Bernardo do Campo

Presencial
BRL 180 000 - 300 000
Flextime
Professional growth
Competitive compensation (USD-based)
Platform Engineer ID90030
Platform Engineer ID90030

AgileEngine, LLC. • Recife

Presencial
BRL 622 000 - 933 000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
Platform Engineer ID90030
Platform Engineer ID90030

AgileEngine, LLC. • Sorocaba

Presencial
BRL 622 000 - 933 000
Professional growth
Competitive compensation
A selection of exciting projects
+1
Platform Engineer ID90030
Platform Engineer ID90030

AgileEngine, LLC. • Porto Alegre

Presencial
BRL 622 000 - 933 000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1