Senior PySpark & Databricks Data Engineer

InfoSmart Technologies, Inc

Atlanta (GA)

Hybrid

USD 120,000 - 150,000

Full time

39 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

InfoSmart Technologies, Inc. is seeking a Data Engineer to design, build, and maintain scalable data pipelines using PySpark and Databricks, with a focus on streaming data via Kafka and orchestrating workflows with Airflow.

The role requires hands-on AWS data warehousing on Redshift, building cloud-native solutions with S3, Glue, Lambda, EMR, IAM, and CloudWatch, and close collaboration with analysts to deliver high-quality data products.

Qualifications

  • Hands-on PySpark experience with Databricks.
  • Real-time data streaming with Kafka.
  • Proficiency in Airflow DAGs.
  • Experience with Redshift and AWS data services.
  • Strong SQL and data modeling skills.

Responsibilities

  • Design and develop scalable ETL/ELT pipelines using PySpark on Databricks
  • Build and maintain real-time data streaming pipelines using Apache Kafka
  • Orchestrate and schedule data workflows using Apache Airflow
  • Manage and optimize data models and queries in Amazon Redshift
  • Architect and deploy data solutions on AWS using services such as S3, Glue, Lambda, EMR, IAM, and CloudWatch
  • Collaborate with analysts and platform teams to deliver high-quality data products
  • Monitor pipeline performance and implement tuning strategies for large-scale data workloads
  • Implement data quality checks, observability, and alerting across pipelines
  • Participate in code reviews and contribute to engineering best practices
  • Document data flows, architecture decisions, and pipeline logic

Skills

PySpark
DataFrame API
Spark SQL
Databricks
Apache Kafka
Apache Airflow
Amazon Redshift
AWS
Python
Data modeling
Query tuning
Cloud architecture

Education

Bachelor's or Master's degree in CS/Engineering or related field

Tools

Databricks
Apache Kafka
Apache Airflow
Amazon Redshift
AWS
S3
Glue
Lambda
EMR
IAM
CloudWatch
Delta Lake
dbt
Terraform

Job description

Location: Atlanta, Georgia - Hybrid 3 days/onsite

Role Summary :

We are looking for a skilled Data Engineer with strong hands‑on experience in PySpark and Databricks to design, build, and maintain scalable data pipelines. The ideal candidate has deep expertise in stream

processing with Apache Kafka, workflow orchestration with Apache Airflow, data warehousing on Amazon Redshift, and building cloud‑native solutions on AWS.

Key Responsibilities
  • Design and develop scalable ETL/ELT pipelines using PySpark on Databricks
  • Build and maintain real‑time data streaming pipelines using Apache Kafka
  • Orchestrate and schedule data workflows using Apache Airflow (DAG development, monitoring, and troubleshooting)
  • Manage and optimize data models and queries in Amazon Redshift
  • Architect and deploy data solutions on AWS using services such as S3, Glue, Lambda, EMR, IAM, and CloudWatch
  • Collaborate with analysts and platform teams to deliver high‑quality data products
  • Monitor pipeline performance and implement tuning strategies for large‑scale data workloads
  • Implement data quality checks, observability, and alerting across pipelines
  • Participate in code reviews and contribute to engineering best practices
  • Document data flows, architecture decisions, and pipeline logic
Required Skills & Experience
  • PySpark — 3+ years
  • DataFrame API, Spark SQL, query optimizations and performance tuning
  • Databricks — 3+ years
  • Notebooks, Jobs, Delta Lake, Unity Catalog
  • Apache Kafka — 2+ years
  • Apache Airflow — 2+ years
  • DAG authoring, task dependencies, operators, scheduling
  • Amazon Redshift — 2+ years
  • Data modeling, query tuning, Redshift Spectrum
  • AWS — 3+ years
  • S3, Glue, Lambda, EMR, IAM, CloudWatch, VPC
  • Python — 4+ years
  • Strong scripting and engineering fundamentals
  • Complex queries, window functions, performance tuning
Nice to Have
  • Experience with Delta Lake
  • Familiarity with dbt for transformation layers
  • Knowledge of Databricks Workflows alongside Airflow
  • Exposure to Confluent Platform or AWS MSK (Managed Kafka)
  • Experience with Terraform or AWS CDK for infrastructure-as-code
  • Understanding of data governance, security best practices, and lake house architecture
Qualifications
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field (or equivalent practical experience)
  • 4–7 years of overall experience in data engineering roles
  • Strong problem-solving skills and ability to work independently in an agile environment
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Prosum • Glendale (CA)

On-site
USD 150,000 - 210,000
Databricks Data Engineer
Databricks Data Engineer

VOLTO Consulting • Irving (TX)

On-site
USD 120,000 - 150,000
Data Engineer – Databricks
Data Engineer – Databricks

XEqualTo • United States

Remote
USD 120,000 - 180,000
Data Engineer
Data Engineer

Akaasa Technologies • New York (NY)

On-site
USD 99,000 - 135,000
Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
Senior PySpark Data Engineer
Senior PySpark Data Engineer

Covetus • Irving (TX)

On-site
USD 100,000 - 130,000
Databricks Data Engineer
Databricks Data Engineer

Henderson Scott • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer - AWS/Databricks - Mid Level
Data Engineer - AWS/Databricks - Mid Level

Acuity, Inc. • Reston (VA)

On-site
USD 100,000 - 130,000
Azure Databricks Engineer
Azure Databricks Engineer

ZEUS SOLUTIONS INC • Houston (TX)

On-site
USD 140,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Acestack • Irvine (CA)

On-site
USD 140,000 - 190,000