Remote Data Engineer: Scale Lakehouse Pipelines & Products

Pashtek • Salesforce Partner | Data & AI

United States

Remote

USD 83,000 - 179,000

Full time

21 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Pashtek • Salesforce Partner | Data & AI is seeking a Data Engineer to build and operate pipelines, tables, and services powering a hybrid data platform across on-premises and AWS. You will implement lakehouse patterns, productionize batch and streaming workloads, and collaborate with security, platform, analytics, and application teams to deliver governed, high-performance data products at scale.

You will develop ELT/ETL jobs using Spark/SQL/dbt/Airflow/Glue, create Iceberg/Delta/Hudi tables,

Qualifications

  • 5+ years of data engineering delivering production pipelines on at least two large-scale platforms.
  • Hands-on with AWS data services (S3, Glue/EMR, Lake Formation, IAM) and at least one data warehouse (Snowflake or Redshift).
  • Deep experience with Apache Spark and related platforms (Databricks, EMR Spark, Snowpark, Dremio, Starburst/Trino).
  • Experience with open table formats (Iceberg preferred), Delta Lake, or Apache Hudi; knowledge of metadata and schema evolution.
  • On-prem stacks experience (Hadoop/Hive, Spark on Kubernetes, SQL Server/SSIS, Oracle/Exadata); Netezza/Teradata a plus.
  • Proven data modeling (3NF, dimensional, Data Vault), ELT/ETL design, and SQL performance tuning.
  • Security/governance: RBAC/ABAC, row/column security, masking/tokenization, KMS.
  • IaC & CI/CD for data workloads.
  • Excellent communicator with stakeholders to meet SLAs and roadmap goals.

Responsibilities

  • Build data pipelines and data products using Spark/SQL/dbt/Airflow/Glue.
  • Create and maintain lakehouse tables with Iceberg/Delta/Hudi.
  • Operate compute engines (Spark, Trino/Starburst, Dremio, Snowflake).
  • Model data for analytics and document contracts/SLAs.
  • Apply governance, lineage, and security controls; support audit readiness.
  • Tune performance and control costs; monitor clusters and workload isolation.
  • Execute migrations from legacy on-prem to lakehouse patterns.
  • Implement streaming/CDC pipelines with Kafka/MSK/Kinesis; Debezium/DMS/Fivetran.
  • Ensure quality and observability with tests, data contracts, and SLO dashboards.
  • Contribute to platform guardrails and write runbooks/docs.
  • Practice DevOps for data with Terraform/CloudFormation and CI/CD.

Skills

Data modeling
SQL performance tuning
Communication
Security/governance
IaC/CI/CD

Tools

S3
Glue/EMR
Lake Formation
IAM
Snowflake/Redshift
Apache Spark
Iceberg/Delta/Hudi
KMS
Kafka/MSK/Kinesis

Job description

Pashtek • Salesforce Partner | Data & AI is seeking a Data Engineer to build and operate pipelines, tables, and services powering a hybrid data platform across on-premises and AWS. You will implement lakehouse patterns, productionize batch and streaming workloads, and collaborate with security, platform, analytics, and application teams to deliver governed, high-performance data products at scale.

You will develop ELT/ETL jobs using Spark/SQL/dbt/Airflow/Glue, create Iceberg/Delta/Hudi tables,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer: Scale Pipelines & Lakehouse Leadership
Senior Data Engineer: Scale Pipelines & Lakehouse Leadership

Survey Sampling International Hyderabad Private Ltd. (India) • Westport (CT)

On-site
USD 130,000 - 150,000
Medical benefits
Discretionary incentive program
Senior Data Engineer - Lakehouse & Pipelines Expert
Senior Data Engineer - Lakehouse & Pipelines Expert

Taskus India • United States

Remote
USD 140,000 - 190,000
Senior Data Engineer: Lead Lakehouse Pipelines & ETL
Senior Data Engineer: Lead Lakehouse Pipelines & ETL

Dynata • United States

On-site
USD 130,000 - 150,000
Medical benefits
Discretionary incentive program
Data Engineer: Build Scalable Pipelines & Lakehouse
Data Engineer: Build Scalable Pipelines & Lakehouse

Sonatype Inc • Abbeyville (CO)

Hybrid
USD 90,000 - 130,000
Parental leave
Diversity and inclusion working groups
Flexible working practices
Senior PySpark Engineer: Cloud Data Pipelines & Lakehouse
Senior PySpark Engineer: Cloud Data Pipelines & Lakehouse

Resource Informatics Group, Inc • United States

On-site
USD 140,000 - 200,000
Remote AWS Lakehouse Data Engineer for AI/ML Pipelines
Remote AWS Lakehouse Data Engineer for AI/ML Pipelines

Delan Associates, Inc • United States

Remote
USD 140,000 - 180,000
AWS Lakehouse Data Engineer: Scalable AI/ML Pipelines
AWS Lakehouse Data Engineer: Scalable AI/ML Pipelines

Guidehouse • Washington

On-site
USD 113,000 - 188,000
Medical Insurance
401(k) Retirement Plan
Paid Holidays
Senior Data Engineer — AI-Driven Lakehouse & Pipelines
Senior Data Engineer — AI-Driven Lakehouse & Pipelines

TDIndustries, Inc. • Dallas (TX)

On-site
USD 108,000 - 132,000
Lead Data Engineer: Lakehouse & Pipelines
Lead Data Engineer: Lakehouse & Pipelines

Doctronic • New York (NY), Northern (KY)

Hybrid
USD 200,000 - 275,000
Data Engineer I - Cloud Lakehouse Pipelines (Remote)
Data Engineer I - Cloud Lakehouse Pipelines (Remote)

Vālenz Health • Northern (KY)

Hybrid
USD 75,000 - 105,000
Healthcare benefits
401(k) with match
Remote-first environment