Senior /Master Data Developer

Jobgether

Brasil

Presencial

BRL 180 000 - 240 000

Tempo integral

Há 4 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Health Coverage
Food Allowance
Family Support
Wellness Partnerships
Profit Sharing
Life Insurance
Continuous Learning
Online Learning
Language Learning
Employee Discounts
Wellbeing Resources
Family Development
Inclusive Workplace
Professional Growth

Resumo da oferta

Jobgether seeks a Senior/Master Data Developer based in Brazil to lead the modernization of large-scale data pipelines. You will refactor legacy Databricks workloads to a cloud-native architecture on Google Cloud, using Python, PySpark, SQL, and strong data quality practices across Raw/Trusted Core/Gold layers.

You will collaborate with Data Stewards and business stakeholders to validate migrations, manage dependencies, and ensure reliable data contracts in a collaborative engineering

Qualificações

  • Strong hands-on Python, PySpark, and SQL for large-scale data processing.
  • Experience migrating data lakes and refactoring legacy data-processing code.
  • Proficiency with Google Cloud (BigQuery, Dataform, Dataproc Serverless).
  • Knowledge of Medallion Architecture and star schema modelling.
  • Experience with GitOps, code reviews, and CI/CD pipelines.
  • Familiarity with CDC, event-driven ingestion, and data vault vs star schema.

Responsabilidades

  • Translate Databricks notebook logic into modern ELT patterns with SQL and Python.
  • Preserve data contracts and protect downstream systems during migrations.
  • Develop scalable ingestion pipelines via YAML-configured DAGs.
  • Implement data quality tests and validation within transformations.
  • Validate migrations with parity checks between legacy and new platform.
  • Follow GitOps practices with branches, PRs, and automated deployments.
  • Collaborate with Data Stewards and tech teams on dependencies.
  • Contribute to cloud-native modernization of data processing architecture.

Conhecimentos

Python
PySpark
SQL
Google Cloud
BigQuery
Dataproc
Airflow
Data Quality
GitOps
Collaboration
Generative AI

Ferramentas

Databricks
Dataform
Dataproc Serverless
Composer

Descrição da oferta de emprego

Senior /Master Data Developer based in Brazil.

This senior role is central to a large-scale modernization of data pipelines, helping transform legacy Databricks workloads into a modern, cloud-native architecture on Google Cloud. You will work hands-on with complex data processing logic, refactoring legacy code into scalable ELT patterns while maintaining critical downstream integrations. The position combines advanced Python, PySpark, SQL, and Google Cloud data engineering with strong data quality and validation practices. You will build scalable ingestion pipelines and help maintain reliable data contracts across Raw, Trusted Core, and Gold layers. Close collaboration with Data Stewards and business stakeholders will be essential to validate migrations and manage dependencies. This is an opportunity to play a key technical role in a major data transformation initiative within a collaborative, engineering-focused environment.

Accountabilities:
  • Legacy Code Refactoring: Translate complex Databricks notebook logic into modern ELT patterns, using declarative SQL-based transformations for standard use cases and Python-based distributed processing for more intricate logic, including RDD-based operations.
  • Data Contract Preservation: Maintain backward compatibility throughout migrations by implementing trusted-core layers and reverse-view strategies that protect downstream systems, integrations, and production dashboards from breaking changes.
  • Ingestion Pipeline Development: Build scalable ingestion pipelines by parameterizing YAML configurations that feed automated DAG-generation processes, using workflow orchestration and standardized processing templates.
  • Data Quality & Testing: Implement synchronous and unit-level assertions within transformation layers to identify null values, enforce key uniqueness, and validate domain-specific data requirements.
  • Migration Validation: Conduct technical validation exercises by reconciling data between legacy environments and the new platform, performing record-level and field-level parity checks to confirm migration accuracy.
  • GitOps & CI/CD: Follow structured GitOps development practices, including feature branches, code reviews, pull requests, and automated deployment pipelines across development and production environments.
  • Stakeholder Collaboration: Work closely with business Data Stewards and technical teams to understand integration dependencies, negotiate refactoring scope, and validate migrated data and outcomes.
  • Architecture Modernization: Contribute to the evolution of data processing practices by applying cloud-native engineering principles and helping transition tightly coupled legacy workloads into modular, maintainable components.
Requirements:
  • Data Engineering Expertise: Strong hands-on technical expertise in Python, PySpark, and advanced SQL, particularly for large-scale and distributed data processing.
  • Migration & Refactoring: Proven experience migrating data lakes and refactoring legacy data-processing code, including the ability to reverse-engineer tightly coupled logic from platforms such as Databricks or Azure Data Factory.
  • Google Cloud: Practical mastery of the Google Cloud data engineering ecosystem, particularly BigQuery, Dataform, and Dataproc Serverless.
  • Data Architecture: Strong understanding of Medallion Architecture, including Raw/Bronze, Silver/Trusted Core, and Gold layers, as well as analytical modeling approaches such as Star Schema and One Big Table.
  • Orchestration & Automation: Hands‑on experience with modern GitOps development practices, including code reviews and pull requests, combined with orchestration using Airflow or Composer.
  • Data Quality: Mature understanding of data quality practices and experience implementing testing directly within data transformation pipelines.
  • Performance Optimization: Ability to analyze and optimise distributed data workloads, with strong attention to scalability, reliability, and processing efficiency.
  • Collaboration: Consultative and collaborative communication style, with the ability to align cross‑team dependencies, negotiate technical scope, and validate outcomes with business Data Stewards.
  • Generative AI: Experience using Generative AI tools such as Gemini, Vertex AI, or Claude to support PySpark‑to‑SQL refactoring, dependency analysis, documentation, or other data engineering automation is a plus.
  • Additional Data Engineering Expertise: Experience with Change Data Capture (CDC), event‑driven ingestion, BigQuery performance tuning, partitioning, execution cost optimisation, and Data Vault versus Star Schema modelling approaches is advantageous.
  • Location Requirement: Candidates residing in the Campinas Metropolitan Region are expected to work from the local offices in accordance with the applicable attendance policy.
Benefits:
  • Health Coverage: Health and dental insurance.
  • Food Allowance: Food and meal allowances.
  • Family Support: Childcare assistance and extended parental leave.
  • Wellness: Partnerships with gyms and health and wellness professionals through Wellhub (Gympass) and TotalPass.
  • Profit Sharing: Participation in a profit‑sharing and results programme (PLR).
  • Life Insurance: Life insurance coverage.
  • Continuous Learning: Access to a dedicated continuous learning platform and professional development resources.
  • Online Learning: Partnerships with online course platforms to support ongoing skill development.
  • Language Learning: Access to a dedicated language‑learning platform.
  • Discounts: Employee discount club with partner offers.
  • Wellbeing Resources: Free access to an online platform focused on physical health, mental well‑being, and overall wellness.
  • Family Development: Pregnancy and responsible parenting courses.
  • Inclusive Workplace: Dedicated health and wellbeing teams, inclusion specialists, and affinity groups providing support throughout the employee journey.
  • Professional Growth: Opportunities to work on large‑scale cloud and data modernisation initiatives while developing expertise across modern data engineering technologies.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Mid Data & Software Engineer
Mid Data & Software Engineer

Jobgether • Brasil

Presencial
BRL 150 000 - 230 000
Global collaboration
Professional development
Travel opportunities
+1
Senior Data Engineer
Senior Data Engineer

Jobgether • Brasil

Presencial
BRL 240 000 - 420 000
Performance-based bonus
Private pension plan
Meal allowance
+6
Mid-Level Data Scientist
Mid-Level Data Scientist

Jobgether • Brasil

Presencial
BRL 120 000 - 160 000
Health insurance
Meal allowance
Childcare assistance
+3
Teach Lead Data Engineer
Teach Lead Data Engineer

Jobgether • Brasil

Presencial
BRL 280 000 - 520 000
Career growth opportunities
Collaborative environment
Global data/tech exposure
IT Coordinator - Data & Analytics Delivery
IT Coordinator - Data & Analytics Delivery

Jobgether • Brasil

Híbrido
BRL 180 000 - 240 000
Permanent full-time employment
Flexible work model
Healthcare coverage
+3
Senior Data Engineer aa
Senior Data Engineer aa

Jobgether • Brasil

Presencial
BRL 300 000 - 480 000
Stock options
Health benefits
Remote within the Americas
+1
Senior Data Engineer, Campinas, Brazil (Hybrid)
Senior Data Engineer, Campinas, Brazil (Hybrid)

CI&T Software S.A. • Campinas

Híbrido
Health and dental insurance
Meal and food allowance
Extended paternity leave
+7
Senior Analytics Engineer (Databricks, Spark)
Senior Analytics Engineer (Databricks, Spark)

externaljobboards • São Paulo

Presencial
BRL 180 000 - 260 000
Senior Data Engineer (Databricks) - Remote Work | REF#303337
Senior Data Engineer (Databricks) - Remote Work | REF#303337

BairesDev • São Paulo

Presencial
BRL 614 000 - 922 000
100% remote
Competitive USD compensation
Home office setup provided
+3
Data Engineer ID43407
Data Engineer ID43407

AgileEngine • Belo Horizonte

Híbrido
BRL 624 000 - 936 000
Professional growth
Competitive compensation
A selection of exciting projects
+1