Lead Data Engineer ID71008

AgileEngine

Recife

Teletrabalho

BRL 728 000 - 988 000

Tempo integral

Há 5 dias
Torna-te num dos primeiros candidatos

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Vantagens oferecidas por esta oferta de emprego

Growth opportunities
Competitive compensation
Fully remote work
Modern projects
Collaborative culture
Well-being support

Resumo da oferta

AgileEngine is seeking a Lead Data Engineer to own the data pipeline and analytical architecture for a large-volume marketing analytics platform. You will decide on partitioning, schema, and near-real-time processing for OLAP workloads on an S3-backed data lake, define DAG-based orchestration with Airflow, and lead senior developers while upholding quality standards.

The role emphasizes autonomy, architecture-driven problem solving, and deep AWS integration, including S3, Athena, and EKS, with

Qualificações

  • 7+ years of engineering experience designing and implementing ETL pipelines for large-volume data systems.
  • Hands-on experience with OLAP-style architecture and engines like Athena/Trino/Presto/BigQuery/Snowflake/Spark SQL.
  • Experience designing data lakes on object storage (S3 or equivalent) with partitioning and Parquet/ORC formats.
  • Deep familiarity with DAG-style workflow orchestration (Airflow) or comparable tools (Dagster, Prefect, Luigi, Step Functions).
  • Proficiency in AWS data stack components (S3, Athena, EKS/Kubernetes) and collaborating with DevOps.
  • Backend proficiency in Python (FastAPI or Flask) and REST/GraphQL APIs.
  • PostgreSQL for transactional/app layer and Docker for containerization.
  • Comfort with Mac/Linux terminals and AI-assisted development tools with responsible usage.

Responsabilidades

  • Design and own ETL pipelines for data extraction, transformation, and validation at scale.
  • Make architectural calls on partitioning, file formats, and near-real-time processing for large OLAP systems built on a data lake.
  • Design and manage Airflow DAGs for batch workflows; collaborate with DevOps to align requirements.
  • Lead and mentor senior developers; enforce code quality and engineering practices.
  • Drive AWS data stack adoption (S3, Athena, EKS) and ensure scalable data architecture.

Conhecimentos

ETL pipelines
Airflow
AWS data stack
Python
REST API
GraphQL
PostgreSQL
Docker
Linux/macOS
AI tooling

Ferramentas

Athena
Trino/Presto
BigQuery
Snowflake
Spark SQL
ClickHouse
Airflow

Descrição da oferta de emprego

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US

If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE

We are looking for a Lead Data Engineer to own the data pipeline and analytical architecture layer for a large-volume marketing analytics platform. You will make architectural decisions around partitioning strategy, file formats, schema design, and near-real-time processing for OLAP-oriented workloads built on an S3-backed data lake. You will design and govern ETL pipelines, define DAG-based orchestration strategies using Airflow, drive the AWS data stack including Athena and EKS, and lead a team of senior developers while enforcing code quality standards. The role requires a high degree of autonomy: you will often work on ad hoc or underspecified problems, defining the problem, gathering context, identifying constraints, and shaping the right technical approach before implementation.

WHAT YOU WILL DO
  • Design and own ETL pipelines that extract, transform, and validate data from internal databases and external APIs at scale.
  • Make architectural calls on partitioning, file formats, schema/data-type strategy, and near-real-time processing for large-volume, OLAP-oriented data systems built on an object-storage data lake.
  • Own the design of scheduled batch workflows (DAGs) on the Airflow setup — defining pipeline structure, dependencies, and triggering strategy, and driving architectural conversations about them. Not responsible for administering Airflow itself.
  • Drive use of the AWS data stack (S3-backed data lake, Athena, EKS/Kubernetes), and partner directly with the DevOps team to clarify functional and non-functional requirements.
  • Review PRs and enforce code quality standards.
  • Guide senior developers and ensure alignment with established engineering practices.
MUST HAVES
  • 7+ years of engineering experience , with a proven track record designing and implementing ETL pipelines and making architectural decisions for large-volume data systems.
  • Hands-on experience with OLAP-style analytical data architecture — comparable experience with Athena, Trino/Presto, BigQuery, Snowflake, Spark SQL, ClickHouse, or similar is acceptable; a specific stack isn't mandatory as long as the OLAP depth is real.
  • Object-storage-backed data lakes : hands-on experience designing against a data lake sitting on object storage (S3 or equivalent) queried via a serverless engine — including partitioning strategy, file formats (Parquet/ORC), and the cost/performance tradeoffs that come with them. Athena specifically is a plus, not a requirement.
  • Task orchestration : Deep familiarity with DAG-style workflow definition and triggering. Most batch processing is orchestrated through Airflow, so this role needs either substantial prior Airflow experience they can draw on to drive architectural conversations, or enough depth in a comparable orchestrator (Dagster, Prefect, Luigi, Step Functions) to ramp on Airflow quickly and lead those conversations from day one. Managing the Airflow deployment itself is out of scope.
  • AWS Ecosystem : practical comfort across the AWS data stack — S3-backed data lake, serverless query engines (Athena or equivalent), and EKS/Kubernetes — with the ability to drive infrastructure conversations with DevOps.
  • Backend proficiency in Python (FastAPI or Flask).
  • Comfortable with REST and GraphQL .
  • Docker and PostgreSQL for the transactional/application layer.
  • Highly comfortable in Mac/Linux terminal-centric environments .
  • Practical, hands-on use of AI-assisted development tools (e.g., Claude Code), paired with the critical judgment to challenge AI output when it compromises long-term maintainability — including the leadership presence to set the standard for how the team uses AI tooling responsibly (e.g., flagging risky AI-driven shortcuts during PR review).
  • Strong soft skills: the ability to hold and defend a technical opinion — challenging a stakeholder's or a tool's proposed "quick fix" with sound reasoning in pursuit of a solution that scales and is maintainable long-term, while still being pragmatic enough to ship.
  • Comfort with ambiguity (mandatory) : work is frequently ad hoc and underspecified. This role requires defining the problem — gathering context, identifying constraints, and framing the work — before solving it, rather than waiting for a specification. Experience limited to well-specified work executed through agent workflows is not a fit.
  • Upper-Intermediate English level.
NICE TO HAVES
  • Direct production experience with Athena specifically.
  • Working knowledge of TypeScript/React — enough to guide integration and review frontend-adjacent PRs, even if not the primary focus.
  • Production AI features using AWS Bedrock, LangChain, Pydantic AI, or similar.
  • Monorepo tooling (Nx) or modern package managers (Poetry, UV, Yarn).
  • Redis/caching layers, SageMaker.
  • Experience with marketing data structures, campaign management APIs, or digital advertising metrics.
PERKS AND BENEFITS
  • Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
  • Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews
  • Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
  • Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands
  • Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
  • Well-being & support : access local well-being programs and people-focused support tailored to your location
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Lead Data Engineer
Lead Data Engineer

AgileEngine, LLC • Brasil

Presencial
BRL 260 000 - 520 000
100% remote
Flexible hours
Annual learning budget
+2
Tech Lead Data Engineer
Tech Lead Data Engineer

AgileEngine • Brasil

Híbrido
BRL 180 000 - 300 000
Flextime
Remote work options
Mentorship
+3
Technical Lead ID71009
Technical Lead ID71009

AgileEngine • Rio de Janeiro

Presencial
BRL 180 000 - 300 000
Growth without limits
Competitive compensation
Flexibility: 100% remote
+3
Technical Lead
Technical Lead

AgileEngine • Brasil

Presencial
BRL 614 000 - 922 000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Data Engineer ID71670
Senior Data Engineer ID71670

AgileEngine • Salvador

Híbrido
BRL 624 000 - 936 000
Professional growth
Competitive compensation
Exciting projects
+1
Data Engineer (Lead) ID52236
Data Engineer (Lead) ID52236

AgileEngine • Brasília

Híbrido
BRL 624 000 - 936 000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1
Data Engineer (Lead) ID52236
Data Engineer (Lead) ID52236

AgileEngine • Curitiba

Híbrido
BRL 624 000 - 936 000
Professional growth: Mentorship and TechTalks
Competitive compensation with budgets for education and fitness
Exciting projects with Fortune 500 companies
+1
Data Engineer (Lead) ID52236
Data Engineer (Lead) ID52236

AgileEngine • Fortaleza

Híbrido
BRL 624 000 - 936 000
Professional growth: Mentorship and personalized roadmaps
Competitive compensation: USD-based pay with multiple budgets
Exciting projects: Work with Fortune 500 companies
+1
Data Engineer (Lead) ID52236
Data Engineer (Lead) ID52236

AgileEngine • São José dos Campos

Híbrido
BRL 240 000 - 420 000
Mentorship and TechTalks
Competitive USD-based compensation
Flexible schedule with remote options
+1
Data Engineer (Lead) ID52236
Data Engineer (Lead) ID52236

AgileEngine • Belo Horizonte

Híbrido
BRL 421 000 - 580 000
Mentorship and personalized growth roadmaps
Competitive compensation with budgets for fitness and education
Exciting projects with Fortune 500 brands
+1