Data Engineer

Pashtek • Salesforce Partner | Data & AI

United States

Remote

USD 83,000 - 179,000

Full time

10 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Pashtek • Salesforce Partner | Data & AI is seeking a Data Engineer to build and operate pipelines, tables, and services powering a hybrid data platform across on-premises and AWS. You will implement lakehouse patterns, productionize batch and streaming workloads, and collaborate with security, platform, analytics, and application teams to deliver governed, high-performance data products at scale.

You will develop ELT/ETL jobs using Spark/SQL/dbt/Airflow/Glue, create Iceberg/Delta/Hudi tables,

Qualifications

  • 5+ years of data engineering delivering production pipelines on at least two large-scale platforms.
  • Hands-on with AWS data services (S3, Glue/EMR, Lake Formation, IAM) and at least one data warehouse (Snowflake or Redshift).
  • Deep experience with Apache Spark and related platforms (Databricks, EMR Spark, Snowpark, Dremio, Starburst/Trino).
  • Experience with open table formats (Iceberg preferred), Delta Lake, or Apache Hudi; knowledge of metadata and schema evolution.
  • On-prem stacks experience (Hadoop/Hive, Spark on Kubernetes, SQL Server/SSIS, Oracle/Exadata); Netezza/Teradata a plus.
  • Proven data modeling (3NF, dimensional, Data Vault), ELT/ETL design, and SQL performance tuning.
  • Security/governance: RBAC/ABAC, row/column security, masking/tokenization, KMS.
  • IaC & CI/CD for data workloads.
  • Excellent communicator with stakeholders to meet SLAs and roadmap goals.

Responsibilities

  • Build data pipelines and data products using Spark/SQL/dbt/Airflow/Glue.
  • Create and maintain lakehouse tables with Iceberg/Delta/Hudi.
  • Operate compute engines (Spark, Trino/Starburst, Dremio, Snowflake).
  • Model data for analytics and document contracts/SLAs.
  • Apply governance, lineage, and security controls; support audit readiness.
  • Tune performance and control costs; monitor clusters and workload isolation.
  • Execute migrations from legacy on-prem to lakehouse patterns.
  • Implement streaming/CDC pipelines with Kafka/MSK/Kinesis; Debezium/DMS/Fivetran.
  • Ensure quality and observability with tests, data contracts, and SLO dashboards.
  • Contribute to platform guardrails and write runbooks/docs.
  • Practice DevOps for data with Terraform/CloudFormation and CI/CD.

Skills

Data modeling
SQL performance tuning
Communication
Security/governance
IaC/CI/CD

Tools

S3
Glue/EMR
Lake Formation
IAM
Snowflake/Redshift
Apache Spark
Iceberg/Delta/Hudi
KMS
Kafka/MSK/Kinesis

Job description

As a Data Engineer, you’ll build and operate the pipelines, tables, and services that power our hybrid data platform across on-premises and AWS. You’ll implement lakehouse patterns, productionize batch and streaming workloads, and partner with security, platform, analytics, and application teams to deliver governed, high-performance data products at scale.

What You’ll Do

Location: Remote (United States)

Employment Type: contract

About the Role

As a Data Engineer, you’ll build and operate the pipelines, tables, and services that power our hybrid data platform across on-premises and AWS. You’ll implement lakehouse patterns, productionize batch and streaming workloads, and partner with security, platform, analytics, and application teams to deliver governed, high-performance data products at scale.

What You’ll Do
  • Build data pipelines: Develop reliable ELT/ETL jobs in Spark/SQL/dbt/Airflow/Glue to ingest from on-prem and cloud sources into S3-backed lakes and warehouses.
  • Implement lakehouse tables: Create and maintain Iceberg (or Delta/Hudi) tables using the appropriate catalog (AWS Glue, Hive Metastore, Polaris/REST) with ACID, time travel, and schema evolution.
  • Operate compute engines: Run and tune Spark (Databricks/EMR), Trino/Starburst, Dremio, and Snowflake workloads; leverage pushdown and query acceleration where applicable.
  • Model data for analytics: Deliver dimensional models, semantic layers, and domain-oriented data products; document contracts and SLAs with consumers.
  • Governance & security: Apply data cataloging, lineage, PII classification, Lake Formation permissions, IAM roles, and row/column-level security; contribute to audit readiness.
  • Performance & cost tuning: Optimize partitioning, clustering/Z-order, predicate pushdown, file sizing/compaction, caching, and workload isolation; monitor and right-size clusters.
  • Migrations: Execute migration workstreams from legacy/on-prem EDW (Informatica/SSIS/SAP BW, SQL Server/Oracle, Hadoop) to lakehouse patterns (Spark/dbt/ELT), including dual-run cutovers and reconciliation.
  • Streaming & CDC: Build real-time and near-real-time pipelines using Kafka/MSK/Kinesis and Spark Structured Streaming; implement CDC with Debezium, DMS, or Fivetran.
  • Quality & observability: Add unit/integration tests, expectations/rules, data contracts, lineage, alerting, and SLO dashboards; participate in on‑call rotations.
  • Platform guardrails: Contribute to standards for naming, zones, schemas, S3 layout, encryption, backup/DR, and multi-region replication; write clear runbooks and docs.
  • DevOps for data: Use Terraform/CloudFormation and CI/CD (GitHub Actions/GitLab/Azure DevOps) to version, test, and deploy data assets.
Required Experience
  • 5+ years in data engineering (or equivalent), delivering production pipelines and tables on at least two large-scale platforms.
  • Hands‑on with AWS data services: S3, Glue/EMR, Lake Formation, IAM, and at least one warehouse (Snowflake or Redshift).
  • Deep experience with Apache Spark and at least one of: Databricks, EMR Spark, Snowflake Snowpark, Dremio, or Starburst/Trino.
  • Production experience with open table formats: Apache Iceberg (preferred), Delta Lake, or Apache Hudi; strong grasp of metadata/manifest, compaction, and schema evolution.
  • Comfortable with on-prem stacks: Hadoop/Hive, Spark on Kubernetes, SQL Server/SSIS and/or Oracle/Exadata; Netezza/Teradata a plus.
  • Proven data modeling (3NF, dimensional, Data Vault), ELT/ETL design, and SQL performance tuning.
  • Security/governance: RBAC/ABAC, row/column‑level security, masking/tokenization, KMS/key management.
  • IaC & CI/CD for data workloads.
  • Excellent communicator who can collaborate with platform engineers, analysts, and stakeholders to meet SLAs and roadmap goals.
Nice to Have
  • dbt for ELT, Airflow orchestration, Great Expectations/Deequ for quality, OpenLineage/Marquez for lineage.
  • Streaming experience with Kafka/MSK, Kinesis, or Flink.
  • Catalogs/semantic: AWS Glue Data Catalog, Unity Catalog, Amundsen/DataHub, Atlan/Collibra.
  • BI/serving: DuckDB, Athena, QuickSight, Tableau/Power BI/Looker.
  • Compliance: SOC 2, HIPAA/PHI, GDPR, PCI; SSO/OIDC with Okta.
  • Multi‑tenant platforms or federated governance exposure.
Location & Work Style

Remote with core hours in PST/CST;

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer – AWS Lakehouse (Mandarin Required)
Data Engineer – AWS Lakehouse (Mandarin Required)

Bitus Labs • Irvine (CA)

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

Deltek • United States

Remote
USD 140,000 - 230,000
Data Engineer
Data Engineer

The Value Maximizer • South Carolina

On-site
USD 90,000 - 120,000
AWS Lakehouse Data Engineer
AWS Lakehouse Data Engineer

engineeringjobs.net, Inc. • Atlanta (GA)

Remote
USD 140,000 - 200,000
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Data Engineer
Data Engineer

Compunnel, Inc. • Norfolk (VA)

On-site
USD 110,000 - 150,000
Senior Data Engineer - Full Time Only - Remote
Senior Data Engineer - Full Time Only - Remote

GD Resources LLC • United States

Remote
USD 126,000 - 154,000
Remote work
Data Engineer
Data Engineer

7Seventy • Northern (KY)

On-site
USD 90,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Madison-Davis, LLC • Chicago (IL)

On-site
USD 130,000 - 180,000
Data Engineer
Data Engineer

Compunnel, Inc. • Westlake (OH)

On-site
USD 90,000 - 120,000