Data Engineer (Kafka / PySpark / Hadoop)

REALIGN LLC

Toronto

On-site

CAD 90,000 - 130,000

Full time

37 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

REALIGN LLC in Toronto, ON is seeking a Data Engineer to design, build, and support scalable batch and real-time data pipelines using Kafka, PySpark, Python and Hadoop. You will develop data processing applications, implement ETL/ELT pipelines, optimize Spark jobs, ensure data quality, and collaborate with architects, developers, analysts and business teams in an Agile environment.

The role is onsite with a fast-paced data-driven culture and opportunities to work across enterprise data

Qualifications

  • Strong hands-on experience with Python for data engineering and automation.
  • Strong expertise in PySpark / Apache Spark.
  • Hands-on experience with Apache Kafka for real-time data ingestion and streaming.
  • Strong experience with the Hadoop ecosystem and distributed data processing.
  • Strong SQL skills and experience working with large datasets.
  • Experience developing and maintaining ETL/ELT data pipelines.
  • Strong understanding of distributed computing and data processing concepts.
  • Experience with data ingestion, transformation, cleansing, and integration.
  • Strong troubleshooting and performance optimization skills.

Responsibilities

  • Design, develop, and maintain scalable batch and real-time data pipelines.
  • Develop data processing applications using Python and PySpark/Apache Spark.
  • Build and support Kafka-based data ingestion and streaming pipelines.
  • Work with Hadoop and related technologies to process large volumes of data.
  • Develop and maintain ETL/ELT pipelines for data ingestion, transformation, cleansing, and integration.
  • Perform data validation, reconciliation, and quality checks.
  • Troubleshoot pipeline failures, data discrepancies, and performance issues.
  • Optimize Spark/PySpark jobs and SQL queries for performance and scalability.
  • Monitor data pipelines and resolve production issues.
  • Collaborate with data architects, developers, analysts, and business teams.
  • Participate in Agile development, testing, deployment, and production support activities.

Skills

Python
PySpark
Apache Spark
Apache Kafka
Hadoop ecosystem
SQL
ETL/ELT pipelines
Data pipelines
Distributed computing

Job description

  • Job Title: Data Engineer – Kafka / PySpark / Hadoop
  • Location: Toronto, ON
  • Work Model: Onsite
  • Job Type: Full Time (FTE)
Job Description

We are seeking an experienced Data Engineer with strong hands‑on expertise in Kafka, PySpark, Python, and Hadoop to design, develop, and support scalable batch and real‑time data pipelines. The ideal candidate will have strong experience working with large-scale distributed data processing environments and enterprise data integration solutions.

Key Responsibilities
  • Design, develop, and maintain scalable batch and real-time data pipelines.
  • Develop data processing applications using Python and PySpark/Apache Spark.
  • Build and support Kafka-based data ingestion and streaming pipelines.
  • Work with Hadoop and related technologies to process large volumes of data.
  • Develop and maintain ETL/ELT pipelines for data ingestion, transformation, cleansing, and integration.
  • Perform data validation, reconciliation, and quality checks.
  • Troubleshoot pipeline failures, data discrepancies, and performance issues.
  • Optimize Spark/PySpark jobs and SQL queries for performance and scalability.
  • Monitor data pipelines and resolve production issues.
  • Collaborate with data architects, developers, analysts, and business teams.
  • Participate in Agile development, testing, deployment, and production support activities.
Required Skills
  • Strong hands‑on experience with Python for data engineering and automation.
  • Strong expertise in PySpark / Apache Spark.
  • Hands‑on experience with Apache Kafka for real‑time data ingestion and streaming.
  • Strong experience with the Hadoop ecosystem and distributed data processing.
  • Strong SQL skills and experience working with large datasets.
  • Experience developing and maintaining ETL/ELT data pipelines.
  • Strong understanding of distributed computing and data processing concepts.
  • Experience with data ingestion, transformation, cleansing, and integration.
  • Strong troubleshooting and performance optimization skills.
Good to Have
  • Hive
  • Databricks
  • Git and CI/CD
  • Airflow or Autosys
  • Relational and NoSQL databases
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Cyber Space Technologies LLC • Montreal (administrative region)

Hybrid
CAD 90,000 - 140,000
Senior Data Engineer (Python, Spark, Snowflake) - Toronto, ON
Senior Data Engineer (Python, Spark, Snowflake) - Toronto, ON

Techedin • Toronto

Hybrid
Data Engineer
Data Engineer

JLI Consulting Talent Search • Vaughan

On-site
CAD 80,000 - 100,000
Senior Data Engineer
Senior Data Engineer

Techedin • Toronto

Hybrid
Data Engineer
Data Engineer

ALLTECH CONSULTING SVC INC • Mississauga

On-site
CAD 80,000 - 120,000
Intermediate Data Engineer ( SQL, Databricks/ Snowflakes, DBT, Airflow)
Intermediate Data Engineer ( SQL, Databricks/ Snowflakes, DBT, Airflow)

Talent To Hire Inc. • Toronto

On-site
CAD 80,000 - 100,000
Data Engineer - DataBricks
Data Engineer - DataBricks

Soar Consultants • Toronto

On-site
CAD 80,000 - 100,000
Data Engineer
Data Engineer

Accuro • Mississauga

On-site
CAD 80,000 - 120,000
Sr Data Engineer
Sr Data Engineer

Cyber Space Technologies LLC • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Technology Lead - PySpark Developer
Technology Lead - PySpark Developer

Infosys Limited • Mississauga

On-site
CAD 93,000 - 123,000