Lead Data Engineer with Databricks

Univedge Consulting LLC

St. Louis (MO)

On-site

USD 120,000 - 180,000

Full time

32 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Univedge Consulting LLC is seeking a senior data engineer to design and implement scalable, secure AWS and Databricks data platforms. You will build ETL/ELT pipelines with PySpark, manage Delta Lake governance, and optimize queries and pipelines for performance and cost.

Experience with Terraform for infra-as-code, and CI/CD with Jenkins and GitHub Actions is required. The role involves integrating diverse data sources, ensuring data quality, lineage, security, and compliance while driving

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or related field.
  • AWS Solutions Architect - Professional certification.
  • Databricks Data Engineer Professional or Apache Spark certification.

Responsibilities

  • Design and architect scalable, secure AWS and Databricks solutions using AWS networking and PaaS services.
  • 2,

Skills

Advanced AWS networking
Real-time data streaming
SQL optimization
Distributed data processing
Python
PySpark
Data governance

Education

Bachelor's degree in Computer Science or related field
AWS Certified Solutions Architect - Professional
Databricks Certified Data Engineer Professional or Apache Spark certification

Tools

Databricks
Delta Lake
Unity Catalog
Delta Live Tables
Terraform
Jenkins
GitHub Actions
AWS Kinesis
AWS Redshift
AWS Elastic Cache
AWS Cognito
EKS

Job description

  • 1. Design and architect scalable, secure, and high-performance AWS and Databricks solutions, leveraging AWS Networking, Kinesis, Elastic Cache, Redshift, EKS, Cognito, and PaaS services
  • 2. Develop and maintain robust ETL/ELT pipelines using PySpark and Python within Databricks, including notebooks, jobs, Delta Lake tables, and Unity Catalog for governance
  • 3. Implement medallion architecture (bronze, silver, gold layers) and optimize Spark jobs for performance, cost, scalability, and reliability, addressing partitioning, skew, caching, and adaptive query execution
  • 4. Provision and manage cloud infrastructure using Terraform for Databricks workspaces, clusters, jobs, storage (ADLS/S3), networking, IAM roles and permissions, and related resources on Azure and AWS
  • 5. Write efficient SQL queries for data transformation, validation, and analytics within Databricks, ensuring data quality, lineage, security, and compliance
  • 6. Implement and maintain CI/CD pipelines using Jenkins and GitHub Actions for automated testing and deployment of notebooks, jobs, Delta Live Tables, and Terraform configurations
  • 7. Integrate data from diverse sources (databases, APIs, streaming, files) into cloud storage and processing layers
  • 8. Optimize cloud costs through auto-scaling clusters, spot instances, job scheduling, and efficient resource usage
  • 9. Monitor pipeline performance, troubleshoot failures, and implement alerting and observability using Databricks tools, cloud monitoring services, or third-party solutions
Required Skills:
  • 1. Advanced proficiency in AWS Networking and VPC design
  • 2. Expertise in AWS Kinesis for real-time data streaming and analytics
  • 3. Deep experience with AWS Elastic Cache for distributed caching
  • 4. Extensive experience with AWS Redshift for data warehousing
  • 5. Strong knowledge of AWS PaaS services including Lambda, S3, and Glue
  • 6. Hands-on experience with AWS EKS for container orchestration
  • 7. Proficiency in AWS Cognito for identity and access management
  • 8. Expertise in Databricks, including Delta Lake, Unity Catalog, and Delta Live Tables
  • 9. Strong Python and PySpark skills for distributed data processing
  • 10. Advanced SQL skills for complex querying and optimization
  • 11. Proficiency in Terraform for infrastructure automation and management
  • 12. Experience with CI/CD tools such as Jenkins and GitHub Actions
Preferred Skills:
  • 1. Knowledge of Azure Data Lake, Data Factory, Synapse, and Key Vault
  • 2. Understanding of big data modeling (star/snowflake, dimensional) and lakehouse architecture
  • 3. Experience with performance tuning in Spark and Databricks environments
  • 4. Familiarity with data governance, lineage, and compliance in cloud environments
Desired Qualifications:
  • 1. Bachelor's degree in Computer Science, Information Technology, or a closely related discipline
  • 2. AWS Certified Solutions Architect - Professional
  • 3. Databricks Certified Data Engineer Professional or Apache Spark certification
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Databricks Technical Lead
Databricks Technical Lead

Anblicks • Dallas (TX)

On-site
USD 120,000 - 160,000
Databricks Engineer
Databricks Engineer

iLink Digital • Milpitas (CA), Northern (KY)

On-site
USD 120,000 - 160,000
Databricks Technical Lead
Databricks Technical Lead

Anblicks Inc. • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 190,000
Lead Data Engineer- Databricks
Lead Data Engineer- Databricks

Phaxis • Beaverton (OR)

On-site
USD 120,000 - 180,000
Databricks Data Engineer
Databricks Data Engineer

Compunnel, Inc. • Spring (TX)

On-site
USD 110,000 - 140,000
Databricks Architect
Databricks Architect

Shrive Technologies • Town of Texas (WI)

On-site
USD 150,000 - 200,000
Databricks SME
Databricks SME

Scicominfra • Atlanta (GA)

On-site
USD 180,000 - 240,000
Principal Architect
Principal Architect

Xcede • United States

On-site
USD 150,000 - 210,000
Azure Databricks Engineer (Dallas, TX)
Azure Databricks Engineer (Dallas, TX)

Cedent • Dallas (TX)

On-site
USD 120,000 - 150,000
Databricks Architect
Databricks Architect

Shrive Technologies LLC • Town of Texas (WI)

On-site
USD 140,000 - 210,000