Senior Data Engineer | Eraytec | Remote

Eraytec

United States

Remote

USD 140,000 - 180,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Eraytec is urgently seeking a Senior Data Engineer to architect, build, and optimize large-scale data pipelines in a fully remote contract role. The candidate will lead the design and implementation of scalable cloud data solutions using Python, PySpark, SQL, and AWS services, collaborating with cross-functional teams to deliver enterprise-grade analytics.

The role requires deep expertise in distributed data systems, Airflow orchestration, and data governance to support BI and advanced analytics

Qualifications

  • 6+ years in big data engineering, data platform development, and distributed data systems.
  • Advanced Python for complex data processing, pipeline automation, and ETL operations.
  • Proven PySpark experience for distributed data transformation and large-scale data cleansing.
  • Strong SQL, data modeling, database indexing, and query optimization for big data workloads.
  • Extensive experience with AWS data services: S3, Redshift, Lambda, EMR, Glue.
  • Experience designing and managing Apache Airflow DAGs for automated workflows.
  • Proven ability to work independently in a fully remote contract setting.

Responsibilities

  • Design, build, scale, and maintain high-performance batch and real-time data pipelines using Python, PySpark, and AWS services.
  • Develop complex distributed data transformations on Amazon EMR clusters.
  • Write, tune, and optimize SQL queries and data models across large relational/analytical databases.
  • Architect and manage cloud data lakehouse storage with S3, Glue Data Catalog, and Redshift.
  • Implement and orchestrate automated, fault-tolerant DAGs with Apache Airflow.
  • Deploy event-driven, serverless data ingestion pipelines using AWS Lambda.
  • Identify and resolve performance bottlenecks and pipeline failures in distributed environments.
  • Implement automated data validation and quality checks with comprehensive monitoring and alerts.

Skills

Python
PySpark
SQL
Airflow
AWS
Data modeling
ETL
Distributed systems

Tools

dbt
Delta Lake
Iceberg
Terraform
Docker

Job description

Position Summary:

Eraytec is urgently seeking a highly skilled, results-driven Senior Data Engineer to join our dynamic technology practice in a 100% remote contract capacity. In this high-impact engineering role, you will take ownership of architecting, building, scaling, and optimizing mission-critical big data processing pipelines and cloud analytics infrastructure. Leveraging over six years of specialized experience in advanced Python scripting, PySpark distributed computing, and AWS data ecosystems, you will transform massive volumes of complex data into high-performance analytical assets. This position offers an outstanding remote opportunity for an ambitious data engineering professional to deliver enterprise-scale solutions, optimize data workflows, and drive impactful business intelligence capabilities across distributed cloud environments.

Detailed Job Description:

As a Senior Data Engineer at Eraytec working remotely, you will lead the technical design, development, and ongoing maintenance of high-throughput distributed data pipelines. Collaborating closely with cross-functional agile teams including data architects, business intelligence specialists, machine learning engineers, and cloud infrastructure operations, you will translate complex analytical and enterprise reporting requirements into resilient, automated, and scalable cloud data solutions.

Your core technical mandate centers on engineering petabyte-scale data pipelines using Python and PySpark, writing highly optimized SQL queries to query complex relational and columnar data structures, and orchestrating serverless and managed workflows across Amazon Web Services. You will manage and streamline end-to-end data pipelines leveraging AWS S3 for scalable data lakes, Amazon Redshift for high-performance data warehousing, AWS Lambda for event-driven serverless ingestion, AWS EMR for distributed computing clusters, and AWS Glue for automated metadata cataloging and ETL jobs. Additionally, you will build and govern robust data orchestration DAGs using Apache Airflow, enforce strict data governance and automated quality checks, and diagnose performance bottlenecks across distributed environments. This contract engagement demands sharp diagnostic problem-solving capabilities, deep expertise in cloud data engineering internals, and the self-discipline to excel within a fast-paced remote work model.

Key Responsibilities:
  • Design, build, scale, and maintain high-performance batch and real-time big data pipelines utilizing Python, PySpark, and AWS cloud data services.
  • Develop complex, highly performant distributed data transformations and processing routines on Amazon EMR clusters.
  • Write, tune, and optimize advanced SQL queries, data models, and aggregation schemas across large-scale relational and analytical databases.
  • Architect and manage enterprise cloud data lakehouse storage structures utilizing Amazon S3, AWS Glue Data Catalog, and Amazon Redshift.
  • Implement and orchestrate automated, fault-tolerant workflow pipelines and directed acyclic graphs (DAGs) using Apache Airflow.
  • Deploy event-driven, serverless data ingestion pipelines and micro-batch processors utilizing AWS Lambda functions.
  • Identify, troubleshoot, and resolve performance bottlenecks, pipeline failures, and distributed computing memory issues across big data environments.
  • Implement automated data validation rules, data quality frameworks, comprehensive monitoring, and end-to-end pipeline alerting mechanisms.
Required Qualifications & Skills:
  • Minimum 6+ years of dedicated professional experience in big data engineering, data platform development, and distributed data systems.
  • Advanced programming and scripting expertise in Python specifically tailored for complex data processing, pipeline automation, and ETL operations.
  • Proven hands-on technical proficiency with PySpark for distributed data transformation, DataFrame manipulations, and large-scale data cleansing.
  • Demonstrated mastery of SQL, data modeling, database indexing, and query optimization for high-volume big data workloads.
  • Extensive hands-on architectural experience across native AWS data services: Amazon S3, Amazon Redshift, AWS Lambda, AWS EMR, and AWS Glue.
  • Substantial experience designing, deploying, and managing automated workflow orchestration DAGs utilizing Apache Airflow.
  • Proven track record of designing resilient data architectures capable of handling large-scale structured and semi-structured datasets.
  • Excellent analytical problem-solving skills, proactive communication abilities, and proven capability to work independently in a fully remote contract setting.
Nice-to-Have Skills:
  • Recognized AWS cloud certifications such as AWS Certified Data Engineer - Associate or AWS Certified Solutions Architect.
  • Experience implementing streaming data ingestion pipelines utilizing Apache Kafka, Amazon Kinesis, or Spark Streaming.
  • Hands-on experience with modern data build tools (dbt) or delta lake storage frameworks (Apache Iceberg, Delta Lake).
  • Familiarity with Infrastructure as Code (IaC) tools such as Terraform or AWS CloudFormation for provisioning data platform resources.
  • Working knowledge of containerized deployments utilizing Docker and CI/CD pipelines for automated data code releases.
Application Information:

Employer / Recruiting Firm: Eraytec

Contact Person: Roshini D

Position Title: Senior Data Engineer

Work Location: Remote role

Employment Type: Contract

Hiring Priority: Urgent Hiring

Application Email: roshini.d@eraytec.com

Required Core Tech Stack: Python Scripting, PySpark, SQL, Big Data Processing, AWS (S3, Redshift, Lambda, EMR, Glue), and Apache Airflow

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer - Remote Cloud Data Pipelines
Senior Data Engineer - Remote Cloud Data Pipelines

Eraytec • United States

Remote
USD 140,000 - 180,000
Remote | Data Engineer — $140,000–$180,000/year
Remote | Data Engineer — $140,000–$180,000/year

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 140,000 - 180,000
Data Engineer
Data Engineer

7Seventy • Northern (KY)

On-site
USD 90,000 - 130,000
Remote - Senior PySpark Developer
Remote - Senior PySpark Developer

Resource Informatics Group, Inc • United States

On-site
USD 140,000 - 200,000
Senior Data Engineer
Senior Data Engineer

Knowledge Services • United States

On-site
USD 120,000 - 150,000
Senior Data Engineer
Senior Data Engineer

Cala Sourcing Solutions LLC • Phoenix (AZ)

Remote
USD 140,000 - 170,000
Senior Data Engineer
Senior Data Engineer

7Th Sky Tech • Dallas (TX)

Hybrid
USD 120,000 - 170,000
Senior Data Engineer
Senior Data Engineer

7Th Sky Tech • Palo Alto (CA)

Hybrid
USD 170,000 - 210,000
Senior Data Engineer
Senior Data Engineer

7Th Sky Tech • Charlotte (NC)

Hybrid
USD 120,000 - 180,000
Senior Data Engineer
Senior Data Engineer

7Th Sky Tech • Chicago (IL)

Hybrid
USD 140,000 - 200,000