Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
People, Jobs, and News is seeking a Senior Data Engineer to design, build, and optimize our Ingest Factory and Data Processing Frameworks using Python, PySpark, and the Databricks Lakehouse ecosystem. You will implement metadata-driven pipelines and reusable data frameworks with strong software engineering discipline.
Collaborating in an Agile environment, you will own testing, governance, and orchestration, while tuning performance of large Spark workloads and ensuring scalable, maintainable
We are seeking a seasoned Senior Data Engineer with a strong software engineering mindset to design,
build, and optimize our next-generation Ingest Factory and Data Processing Frameworks. In this role,
you will go beyond traditional ETL scripting to build scalable, metadata-driven pipelines and reusable
data frameworks.
The ideal candidate possesses deep expertise in Python, PySpark, and the Databricks Lakehouse
ecosystem (including LakeFlow and Delta Lake), combined with rigorous software engineering
discipline (SOLID, CI/CD, and infrastructure as code). You will work both independently and
collaboratively within an Agile environment to build production-grade software that ensures data quality,
governance, and seamless orchestration.
leveraging Databricks LakeFlow, managed connectors, and declarative pipelines.
architecture, ensuring optimized storage, table design, and strict data lineage enforcement.
flows to automate ingestion across diverse source system patterns.
SOLID and DRY principles. Move beyond basic PySpark scripting to contribute to and publish
reusable internal packages (e.g., PyPI).
dramatically improve development efficiency across the data team.
frameworks. Lead test practices including Unit, Integration, and End-to-End (E2E) testing.
configurations, and SQL queries (efficient filtering, indexing, and joins) to reduce processing
costs.
performance bottlenecks, data skew, and serialization issues.
pipelines and provision infrastructure utilizing Terraform (IaC).
stories, vet architectures with the team, and deliver retro demos prior to production deployment.
enterprise-grade software.
optimization, Delta Lake, Connectors, and LakeFlow (jobs, tasks, flows).
package management. Clear understanding of distributed workloads (Spark vs. single-node
processing).
(PR reviews, branching strategies), and modern IDE features (Cursor/VSCode).
or GitLab CI).
data processing and model design.
purpose behind it.
during retro demos.
Skills: python,data engineer,pyspark,spark,metadata,metadata driven pipeline,deployment,customization,configuration,data packages,data products