Big Data Developer

Galent

Toronto

On-site

CAD 120,000 - 180,000

Full time

25 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Galent in Toronto is seeking a Senior Backend Developer to design, build, and maintain scalable data pipelines and backend systems in an enterprise environment. You will lead Spark-Scala projects on Hadoop/CDP clusters, optimize ETL pipelines, and migrate Spark 2 apps to Spark 3 while ensuring data governance and security.

The role requires 5+ years of experience in big data engineering, API integration, and AI-assisted development, with hands-on work on Spark, Hive, HDFS, and CI/CD tooling.

Qualifications

  • 5+ years of backend or data engineering experience in enterprise environments.
  • Hands-on with Spark, Hadoop, and data lake ecosystems.
  • Experience migrating Spark 2 to Spark 3 and building scalable ETL pipelines.

Responsibilities

  • Design and develop Spark-Scala apps for large-scale data processing on Hadoop/CDP clusters.
  • Build and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL.
  • Tune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization).
  • Migrate Spark 2 apps to Spark 3 on Cloudera CDP platforms.
  • Work with Parquet, ORC, Avro on HDFS.
  • Write complex HiveQL / Spark SQL queries with window functions, CTEs, subqueries and aggregations.
  • Design and maintain Hive external/managed tables and partitioned datasets.
  • Optimize slow-running queries and resolve correlated subquery issues.
  • Work with HDFS encryption zones and data governance requirements.
  • Develop and maintain bash scripts for job orchestration and automation.
  • Handle error management, logging and alerting in shell scripts.
  • Manage HDFS operations (hdfs dfs commands), file transfers, and data validation.
  • Build scripts and pipelines to extract data from REST APIs using curl and Python.
  • Parse and process JSON API responses and load into HDFS/Hive.
  • Manage pagination, error handling and retry logic for API calls.
  • Work with enterprise API gateways and URL parameter construction.
  • Leverage GitHub Copilot / AI coding assistants to accelerate development.
  • Use AI tools for code review, SQL generation, script debugging and documentation.
  • Contribute to AI-assisted data quality and anomaly detection pipelines.
  • Explore and implement LLM-based automation for repetitive data engineering tasks.

Skills

Big data engineering
API integration
AI-assisted development

Tools

Spark
Hadoop
Hive
HDFS
CDP
Spark SQL
Parquet/ORC/Avro

Job description

We are looking for a Senior Backend Developer with 5+ years of experience in big data engineering, API integration, and AI-assisted development. The ideal candidate will design, build, and maintain scalable data pipelines and backend systems in a enterprise environment.

Key Responsibilities
  • Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clusters
  • Build and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL
  • Tune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)
  • Migrate Spark 2 applications to Spark 3 on Cloudera CDP platforms
  • Work with Parquet, ORC, Avro file formats on HDFS
  • Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations
  • Design and maintain Hive external/managed tables and partitioned datasets
  • Optimize slow-running queries and resolve correlated subquery issues
  • Work with HDFS encryption zones and data governance requirements
Unix / Shell Scripting
  • Develop and maintain bash shell scripts for job orchestration and automation
  • Handle error management, return codes, logging and alerting in shell scripts
  • Manage HDFS operations (hdfs dfs commands), file transfers, and data validation
API Extraction & Integration
  • Build scripts and pipelines to extract data from REST APIs using curl and Python
  • Parse and process JSON API responses and load into HDFS/Hive
  • Manage pagination, error handling and retry logic for API calls
  • Work with enterprise API gateways and URL parameter construction
  • Leverage GitHub Copilot / AI coding assistants to accelerate development
  • Use AI tools for code review, SQL generation, script debugging and documentation
  • Contribute to AI-assisted data quality and anomaly detection pipelines
  • Explore and implement LLM-based automation for repetitive data engineering tasks
Scheduling & Orchestration
  • Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cron
  • Build and maintain Ansible playbooks for automated deployments
  • Manage deployment pipelines including artifact versioning, Vault secret injection and environment-specific configuration
  • Monitor job health, handle failures and implement alerting
Nice to Have
  • Experience with Cloudera CDP (7.x) and migration from HDP
  • Knowledge of Kerberos, Vault, HDFS encryption zones
  • Familiarity with CI/CD pipelines (Helios, GitHub Actions)
  • Experience with MSSQL / JDBC connectivity from Spark
  • Understanding of AML / Financial regulatory data domains.

We are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, citizenship status, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law. https://www.e-verify.gov/sites/default/files/everify/posters/IER_RighttoWorkPoster.pdf

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (Big Data)
Data Engineer (Big Data)

Pacer Group • Toronto

On-site
CAD 90,000 - 130,000
Data Engineer
Data Engineer

Soroc Technology • Toronto

On-site
CAD 110,000 - 150,000
Big Data Developer – Scala/Spark, Java
Big Data Developer – Scala/Spark, Java

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

On-site
CAD 90,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Capgemini • Mississauga

On-site
CAD 110,000 - 170,000
Senior Big Data Engineer: Spark, Hadoop & API Pipelines
Senior Big Data Engineer: Spark, Hadoop & API Pipelines

Galent • Toronto

On-site
CAD 120,000 - 180,000
Senior Java Spark Developer
Senior Java Spark Developer

Capgemini • Montreal (administrative region)

On-site
CAD 38,000 - 59,000
Medical benefits
Dental benefits
Vision benefits
+1
Data Engineer
Data Engineer

ALLTECH CONSULTING SVC INC • Mississauga

On-site
CAD 80,000 - 120,000
Sr Data Engineer
Sr Data Engineer

Tailored Brands, Inc. • Cambridge

On-site
CAD 120,000 - 180,000
Sr Databricks Data Engineer
Sr Databricks Data Engineer

Pacer Group • Toronto

On-site
CAD 110,000 - 170,000
Big Data Developer
Big Data Developer

Cogency • Toronto

Hybrid
CAD 120,000 - 150,000