Hybrid Data Engineer: Cloud Pipelines & ETL Analytics

The Vanguard Group

East Whiteland Township (PA)

Hybrid

USD 90,000 - 130,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

The Vanguard Group is seeking a Data Engineer, Specialist to design, develop, and maintain scalable data pipelines and infrastructure that support analytics and business intelligence. This role involves building robust ETL processes, managing databases and optimizing cloud-based data platforms to ensure efficient and seamless integration of data from multiple sources.

Collaborate with data scientists, analysts and business stakeholders to understand their data needs, develop data solutions and

Qualifications

  • Minimum of three years data engineering, programming, database administration, or data management experience.
  • Undergraduate degree or equivalent combination of training and experience.
  • Strong proficiency in SQL, Python for data manipulation, automation and pipeline development.
  • Experience working with big data processing frameworks such as Apache Spark, Hadoop, or Kafka.

Responsibilities

  • Develop and maintain scalable ETL (Extract, Transform, Load) processes to efficiently extract data from diverse sources, transform it as required and load it into data warehouses or analytical systems.
  • Design and optimize database architectures and data pipelines to ensure high performance, availability and security while supporting structured and unstructured data.
  • Integrate data from multiple sources, including APIs, third-party services and on-prem/cloud databases to create unified and consistent datasets for analytics and reporting.
  • Collaborate with data scientists, analysts and business stakeholders to understand their data needs, develop data solutions and enable self-service analytics.
  • Develop automated workflows and data processing scripts using Python, Spark, SQL, or other relevant technologies to streamline data ingestion and transformation.
  • Optimize data storage and retrieval strategies in cloud-based data warehouses such as AWS Redshift, Google Big Query, or Azure Synapse, ensuring scalability and cost-efficiency.
  • Maintain and improve data quality by implementing validation frameworks, anomaly detection mechanisms and data cleansing processes.
  • Thoroughly tests code to ensure accuracy and alignment with its intended purpose.
  • Reviews the final product with end users to confirm clarity and understanding, providing data analysis guidance as needed.
  • Offers tool and data support to business users and team members, ensuring seamless functionality and accessibility.
  • Conducts regression testing for new software releases, identifying issues and collaborating with vendors to resolve them and successfully deploy the software into production.

Skills

SQL
Python
ETL
Data pipelines

Education

Undergraduate degree or equivalent

Tools

Apache Spark
Hadoop
Kafka
AWS Glue
AWS S3
AWS Lambda
Databricks
Google BigQuery
Azure Synapse
Azure Data Factory

Job description

The Vanguard Group is seeking a Data Engineer, Specialist to design, develop, and maintain scalable data pipelines and infrastructure that support analytics and business intelligence. This role involves building robust ETL processes, managing databases and optimizing cloud-based data platforms to ensure efficient and seamless integration of data from multiple sources.

Collaborate with data scientists, analysts and business stakeholders to understand their data needs, develop data solutions and

Get your free, confidential resume review.
or drag and drop your file here.