A leading tech consulting firm in Fort Worth, Texas, is seeking a skilled data engineer to design and optimize scalable data pipelines using Azure Databricks and PySpark. The role requires a solid background in SQL and experience with data ingestion and transformation processes. The ideal candidate will collaborate closely with data analysts and stakeholders, ensuring efficient data solutions are delivered. This position offers an exciting opportunity to work in a dynamic tech environment focused on enterprise-level data processing.
Qualifications
5+ years' experience in SQL.
4+ years of experience in Azure Databricks with PySpark.
2+ years of experience in Python programming & package builds.
Strong understanding of ETL/ELT design patterns, data warehousing, and data lakehouse architectures.
Responsibilities
Design, develop, and optimize scalable data pipelines leveraging Databricks.
Write clean, maintainable, and efficient PySpark and Python code.
Integrate Azure Databricks with various Azure data services.
Skills
SQL
Azure Databricks
PySpark
Azure Cloud
Data ingestion
Data transformation
Python
ETL/ELT design patterns
Data warehouses
GitHub Actions
Job description
Job Summary
5+ years' experience in SQL (Expert)
4+ years of experience in Azure Databricks with PySpark
4+ years of experience in Azure Cloud platform
3+ years of experience in ADF (Azure Data Factory), ADLS Gen 2 and Azure SQL
2+ years of experience in Databricks workflow & Unity catalog
2+ years of experience in Python programming & package builds
Experience in data ingestion, cleansing, and transformation processes from various structured and unstructured data sources like Cassandra & Mark Logic and on-prem Mainframe sources using Databricks/PySpark supporting batch and near-real-time ingestion, transformation, and processing.
Manage job scheduling, orchestration, and monitoring (e.g., using Azure Data Factory, Airflow, or Databricks Workflows).
Ability to design modular, reusable workflows using tasks, triggers, and dependencies.
Skilled in using dynamic expressions, parameterized pipelines, custom activities, and triggers.
Familiarity with integration runtime configurations, pipeline performance tuning, and error handling strategies.
Strong understanding of ETL/ELT design patterns, data warehousing, and data lakehouse architectures.
Good to have Azure Entra/AD skills and GitHub Actions
Good to have experience working on event-driven architectures using Kafka, Azure Event Hub
Responsibilities
Design develop and optimize scalable data pipelines leveraging Databricks (Spark, PySpark, SQL, Delta Lake) to support enterprise level data processing and analytics
Write clean maintainable and efficient PySpark and Python code to support data ingestion transformation
Integrate Azure Databricks with various Azure data services to build robust and scalable data platforms
Implement and maintain ETL workflows for scalability cost effective and operational efficiency
Collaborate with data analysts and stakeholders to gather requirements and deliver scalable data solutions
Participate actively in design discussions code reviews and agile ceremonies to foster a collaborative and high performing team environment.