Senior Data Engineer

Jobtailor

Campinas

Presencial

BRL 180 000 - 280 000

Tempo integral

Há 3 dias
Torna-te num dos primeiros candidatos

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Jobtailor is seeking a senior data engineer to build and optimize data ingestion and transformation pipelines on AWS. The role focuses on designing data lake architectures, leveraging PySpark, Glue, and Redshift to deliver scalable analytics.

You will ensure governance and cataloging, implement IaC with Terraform, and integrate data from SQL Server sources into the lake while maintaining reliability and performance across platforms.

Qualificações

  • Strong experience in cloud data engineering on AWS.
  • Advanced Python applied to data pipelines.
  • PySpark for distributed processing (high-volume ETL/ELT).
  • AWS Glue (serverless ETL jobs and Glue Data Catalog).
  • Amazon Athena (serverless queries on S3).
  • Terraform (Infrastructure as Code for environment provisioning).
  • Knowledge of data lake modeling and architecture.
  • Certifications like AWS Data Engineer Associate or Solutions Architect are a plus.

Responsabilidades

  • Build and evolve data ingestion and transformation pipelines (batch and near-real-time).
  • Structure and maintain data lake layers in S3 using Parquet.
  • Develop and optimize PySpark and AWS Glue jobs.
  • Model, load, and optimize tables and queries in Redshift and Athena.
  • Implement and maintain governance, cataloging, and permissions through Lake Formation and Glue Data Catalog.
  • Provision and version data infrastructure with Terraform.
  • Integrate data from transactional sources, such as SQL Server, into the data lake.
  • Monitor pipelines, respond to incidents, and ensure data quality and reliability.
  • Document architecture, data flows, and processes.

Conhecimentos

Cloud Data Engineering on AWS
Advanced Python for Data Pipelines
PySpark for Distributed Processing
AWS Glue for ETL Jobs
Amazon Redshift Data Warehouse
Advanced SQL
Data Governance
Infrastructure as Code
Data Modeling
Data Integration
Data Lake Architecture

Formação académica

Bachelor's degree

Ferramentas

Apache Airflow
Docker
Datadog
AWS CloudWatch
GitHub Actions
Terraform
MWAA
dbt
Kinesis/Kafka
MSK

Descrição da oferta de emprego

  • Build and evolve data ingestion and transformation pipelines (batch and near-real-time)
  • Structure and maintain data lake layers in S3 using Parquet
  • Develop and optimize PySpark and AWS Glue jobs
  • Model, load, and optimize tables and queries in Redshift and Athena, considering partitioning, cost, and performance
  • Implement and maintain governance, cataloging, and permissions through Lake Formation and Glue Data Catalog
  • Provision and version data infrastructure with Terraform
  • Integrate data from transactional sources, such as SQL Server, into the data lake
  • Monitor pipelines, respond to incidents, and ensure data quality and reliability
  • Document architecture, data flows, and processes
  • Work at Mobato, a brand that provides digital solutions to connect customers with the automotive sector and integrate ERP and DMS systems
Requirements
  • Strong experience in cloud data engineering on AWS
  • Advanced Python applied to data pipelines
  • PySpark for distributed processing (high-volume ETL/ELT)
  • AWS Glue (serverless ETL jobs and Glue Data Catalog)
  • Amazon Athena (serverless queries on S3)
  • Amazon S3 organized into layers (raw/bronze silver gold) using the Parquet format
  • Amazon Redshift (data warehouse modeling, loading, and optimization)
  • Advanced SQL, including data extraction and integration from SQL Server
  • AWS Lake Formation (data lake governance, cataloging, and permissions management)
  • Terraform (Infrastructure as Code for environment provisioning and versioning)
  • Knowledge of data lake modeling and architecture
  • Nice to have: Apache Airflow, AWS Step Functions, or MWAA
  • Nice to have: Apache Iceberg, Delta Lake, or Hudi
  • Nice to have: dbt
  • Nice to have: Great Expectations, Monte Carlo, or similar tools
  • Nice to have: Kinesis, Kafka, or MSK
  • Nice to have: GitHub Actions or GitLab CI
  • Nice to have: Docker
  • Nice to have: Datadog and AWS CloudWatch
  • Nice to have: AWS certifications (Data Engineer Associate or Solutions Architect)
  • Nice to have: Experience with agile methodologies
  • Bachelor's degree
Core Competencies

Demonstrates expertise in building and optimizing data ingestion and transformation pipelines using AWS services, with a strong focus on data lake architecture and governance. Proficient in Python and SQL for data processing and integration, ensuring data quality and reliability.

Highest-signal resume keywords
  • Cloud Data Engineering on AWS
  • Advanced Python for Data Pipelines
  • PySpark for Distributed Processing
  • AWS Glue for ETL Jobs
  • Amazon Redshift Data Warehouse Optimization
Hard Skills
  • Data Ingestion Pipelines
  • Data Lake Architecture
  • Advanced SQL
  • Data Transformation
  • ETL/ELT Processes
  • Data Quality Assurance
  • Data Governance
  • Data Modeling
  • Infrastructure as Code
  • Data Integration
Certifications & Qualifications
  • AWS Data Engineer Associate
  • AWS Solutions Architect
Industry Keywords
  • Data Lake
  • Digital Solutions
  • Automotive Sector
  • ERP Systems
  • DMS Systems
Tools & Technologies
  • AWS Glue Data Catalog
  • Amazon S3
  • Terraform
  • Amazon Athena
  • AWS Lake Formation
  • Apache Airflow
  • Docker
  • Datadog
  • AWS CloudWatch
  • GitHub Actions
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Data Engineering Consultant
Data Engineering Consultant

Jobtailor • São Paulo

Presencial
BRL 180 000 - 320 000
Data Architect ID52062
Data Architect ID52062

AgileEngine • Campo Grande

Híbrido
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Data Architect ID52062
Data Architect ID52062

AgileEngine • São Luís

Presencial
Professional growth programs
Competitive compensation with benefits
Flexible schedule with remote options
Data Architect ID52062
Data Architect ID52062

AgileEngine • Campinas

Híbrido
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Senior Analytics Engineer (Databricks, Spark)
Senior Analytics Engineer (Databricks, Spark)

exadelinc • São Paulo

Presencial
BRL 300 000 - 420 000
Tech Lead Data Engineer
Tech Lead Data Engineer

AgileEngine • Brasil

Híbrido
BRL 180 000 - 300 000
Flextime
Remote work options
Mentorship
+3
Data Engineering Analyst III
Data Engineering Analyst III

Jobtailor • Fortaleza

Presencial
BRL 180 000 - 240 000
Data Engineer
Data Engineer

Jobtailor • Osasco

Presencial
BRL 180 000 - 320 000
Lead Data Engineer ID71008
Lead Data Engineer ID71008

AgileEngine, LLC. • Salvador

Presencial
BRL 260 000 - 420 000
Growth without limits
Competitive compensation
Flexibility: 100% remote
+3
ENGENHEIRO DE DADOS SÊNIOR
ENGENHEIRO DE DADOS SÊNIOR

Revise Group • Região Geográfica Imediata de Vitória

Presencial
BRL 180 000 - 260 000