Data Engineer � Databricks SME.

PlanIT Group, LLC

Raleigh (NC)

On-site

USD 120,000 - 160,000

Full time

19 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

PlanIT Group, LLC is seeking a Senior Data Engineer to support data ingestion, de-duplication, and tagging for migrating a large-scale data environment into Databricks. The role focuses on end-to-end data pipeline management, data ingestion from diverse sources, and governance-enabled analytics workflows.

The candidate will work with Azure/Databricks, implement deduplication strategies, and develop tagging frameworks to support metadata, lineage, and compliance.

Qualifications

  • 5+ years designing and implementing data ingestion pipelines using Azure Data Factory, Kafka, NiFi, Spark or equivalents.

Responsibilities

  • Design, develop, and maintain scalable data ingestion pipelines for batch and streaming data into Azure/Databricks.

Skills

Data pipeline design
Python or scripting
Azure Cloud
Databricks / Spark
Data deduplication
Data tagging / metadata
Communication to CIO

Education

Bachelor's degree in a relevant field

Tools

Azure Data Factory
Apache Kafka
Apache NiFi
Databricks
Apache Spark

Job description

We are seeking a Senior Data Engineer to support our client with data ingestion, data deduplication and data tagging for migration of a large-scale data environment into Databricks.

Job Description

We are seeking a Senior Data Engineer to support our client with data ingestion, data tagging for migration of a large-scale data environment into Databricks. The ideal candidate will also bring hands-on expertise in end-to-end data pipeline management, including data ingestion from diverse sources, de-duplication of large-scale datasets, and data tagging to support downstream analytics, governance, and machine learning workflows.

Roles And Responsibilities (including But Not Limited To)
  • Design, develop, and maintain scalable data ingestion pipelines to onboard structured, semi-structured, and unstructured data from batch and streaming sources (e.g., APIs, databases, flat files, message queues) into the Azure/Databricks environment.
  • Implement de-duplication strategies across large-scale datasets using deterministic and probabilistic matching techniques to ensure data integrity and reduce redundancy within the Data Lake.
  • Develop and enforce data tagging frameworks to classify, label, and annotate datasets with appropriate metadata (e.g., sensitivity, source, domain, lineage) to support data governance, discoverability, and compliance requirements.
  • Assist with Operationalizing deployments and support of Cloud services for ETL Operations. This will include standardizing and automating processes and workflows, creating documentation/knowledge articles, and overall assisting Operations staff who have limited experience in Cloud.
  • Written and oral presentations to high-level CIO management on status of current efforts.
  • Possesses skills and experience related to business management, systems engineering, operations research, and management engineering. Typically has specialization in a particular technology or business application. Keeps abreast of technological developments and industry trends.
  • Assist with deployment, configuration, and management of Azure Cloud environment.
  • Assist with migration efforts of existing ETL jobs into Azure/Databricks cloud environment.
  • Ability to share optimization and efficiencies with the larger team and management.
  • Ability to automate solutions to repetitive problems/tasks.
Basic Qualifications
  • Must be eligible for a Position of Public Trust, including U.S. citizenship or permanent residency, five years of U.S. residency, and no more than six months of international travel in the past five years (excluding travel for U.S.-based work).
  • Bachelor's degree and 13 years of experience. A degree from an accredited College/University in the applicable field of services is preferred. Four additional years of relevant experience in lieu of a college degree is required. If Degree is not in the applicable field, then four additional years of related experience is required.
  • 5+ years demonstrated experience designing and implementing data ingestion pipelines using tools such as Azure Data Factory, Apache Kafka, Apache NiFi, Spark Structured Streaming, or equivalent technologies.
  • 5+ years of experience applying de-duplication techniques at scale, including record linkage, fuzzy matching, and entity resolution across structured and unstructured datasets.
  • 5+ Hands-on experience with data tagging and metadata management, including the use of tagging schemas, data catalogs (e.g., Azure Purview, Apache Atlas), and automated classification tools to support data governance and lineage tracking.
  • 5 + Demonstrated experience working with unstructured data.
  • 2+ years of experience in using Databricks or other Spark-based platforms.
  • Fluency in at least one scripting language (Python, Perl, Ruby, or equivalent).
Desired Skills
  • Integration of Git in continuous deployment and experience with DevOps monitoring tools.
  • Experience with one or more of the following products and technologies: SAS, Python, C++, Hadoop, SQL Database/Coding, Teradata, Oracle, Amazon S3, Apache Spark, Machine Learning, Natural Language Processing, and visualization tools such as Tableau, Strategy and QLIK.
  • Strong skills and experience in Cloud Operations support in Azure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DataBricks Data Engineer
DataBricks Data Engineer

Prodapt • Irving (TX)

On-site
USD 140,000 - 190,000
Databricks Engineer
Databricks Engineer

Tredence Inc. • Chicago (IL)

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

Compunnel, Inc. • Oakland (CA)

On-site
USD 120,000 - 160,000
Sr. Databricks Engineer - 258215
Sr. Databricks Engineer - 258215

Medix Technology • United States

On-site
USD 88,000 - 95,000
Senior Data Engineer on-site)
Senior Data Engineer on-site)

Ziosk • Dallas (TX)

On-site
USD 140,000 - 190,000
Azure Databricks Engineer (Dallas, TX)
Azure Databricks Engineer (Dallas, TX)

Cedent • Dallas (TX)

On-site
USD 120,000 - 150,000
Databricks Data Engineer
Databricks Data Engineer

Compunnel, Inc. • Spring (TX)

On-site
USD 110,000 - 140,000
Data Architect
Data Architect

Aroha Technologies Inc • Plano (TX)

On-site
USD 90,000 - 130,000
Azure Databricks Engineer
Azure Databricks Engineer

ZEUS SOLUTIONS INC • Houston (TX)

On-site
USD 140,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Prosum • Glendale (CA)

On-site
USD 150,000 - 210,000