ETL/DB/Data Extraction Engineer

WebDataGuru

Indiana (PA)

On-site

USD 110,000 - 170,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

WebDataGuru is seeking a Senior Data Engineer to lead ETL development and large-scale data extraction initiatives. The role focuses on building robust data pipelines, implementing web scraping strategies, and delivering scalable data solutions across cloud platforms.

You will optimize ETL processes, manage data warehouses, and collaborate cross-functionally to ensure data quality from source to analytics. A strong emphasis on security, performance, and CI/CD is required.

Qualifications

  • ETL pipeline development and optimization for large datasets.
  • Bulk Data Extraction using web scraping or crawling techniques.
  • Parallelism, concurrency, and threading to improve performance.
  • Python-based libraries: Numpy, Pandas, Selenium, Scrapy.
  • Postgres, MySQL, and NoSQL databases.
  • AWS, GCP, or Azure with Lambda, EC2, S3.
  • REST APIs development and integration.
  • SQL procedures, complex queries, and data pipelines.
  • Git and SVN version control.
  • SDLC, testing methodologies, and CI/CD processes.
  • Strong analytical and problem-solving abilities for scalable data solutions.

Responsibilities

  • ETL Pipeline Management across production, QA, and UAT.
  • Develop, optimize, and troubleshoot ETL pipelines for large datasets.
  • Handle large datasets and manage server infrastructure.
  • Web Scraping and Crawling to extract data from hundreds of websites.
  • Apply parallelism/concurrency/ threading to improve extraction performance.
  • Leverage Selenium and Scrapy to build scalable scraping solutions.
  • Work with Postgres, MySQL, and NoSQL databases.
  • Implement cloud best practices with AWS, GCP, or Azure (Lambda/EC2/S3).
  • Create and optimize SQL procedures and data pipelines for integration.
  • Collaborate with teams to deliver end-to-end data solutions.
  • Ensure data quality and proper mappings across environments.

Skills

ETL Development
Bulk Data Extraction
Parallel Processing
REST API
SQL Proficiency
Version Control
SDLC Knowledge
Problem-Solving Skills

Tools

Selenium
Scrapy
Pandas
NumPy
Requests
xpath
Urllib
lxml
Postgres
MySQL
NoSQL
AWS
GCP
Azure
Lambda
EC2
S3

Job description

WebDataGuru is hiring a ETL/DB/Data Extraction Engineer

10.10.24

We are seeking an exceptionally skilled Senior Data Engineer with extensive experience in ETL development and Bulk Data Extraction through Web Scraping and Crawling technologies. The ideal candidate will have a solid background in building, optimizing, and maintaining ETL pipelines, with strong expertise in data management and cloud technologies. This is a challenging role designed for a detail-oriented professional who is passionate about delivering high-quality data solutions.

You will be responsible for managing large-scale data projects, ensuring best practices in security, scalability, and performance. If you have a proven track record of delivering comprehensive data engineering solutions across diverse industries, we would love to hear from you.

Roles & Responsibilities

As a Senior Data Engineer, You Will:

  • ETL Pipeline Management:
  • Develop, optimize, and troubleshoot ETL pipelines across production, QA, and UAT environments.
  • Handle large datasets and manage server infrastructure effectively.
  • Web Scraping and Crawling:
  • Implement advanced web scraping techniques for rapid and efficient data extraction from large-scale sources across hundreds of websites.
  • Performance Optimization:
  • Apply techniques such as parallelism, concurrency, and threading to improve data extraction and processing efficiency.
  • Library and Framework Expertise:
  • Utilize relevant tools such as requests, xpath, ActionChain, Urllib, Numpy, Pandas, and lxml.
  • Work with frameworks like Selenium and Scrapy to build scalable scraping solutions.
  • Database Management:
  • Work with both relational databases (Postgres, MySQL) and NoSQL databases.
  • Optimize database performance, including schemas, triggers, and queries.
  • Cloud and API Integration:
  • Implement cloud best practices for security, scalability, and performance using AWS, GCP, or Azure.
  • Utilize AWS services such as Lambda, EC2, and S3 alongside external APIs.
  • Data Engineering:
  • Participate in data warehousing, architecture, and ETL pipeline development at an enterprise scale.
  • Create and optimize SQL procedures and data pipelines for seamless integration.
  • Collaboration and Project Delivery:
  • Work cross-functionally with teams to deliver end-to-end data solutions.
  • Ensure source data quality, develop mapping and transformations, and validate jobs across various environments.
Skills & Requirements

To be successful in this role, you should possess:

  • ETL Development Expertise: Proficiency in ETL pipeline development and optimization for large datasets.
  • Bulk Data Extraction: Experience with Web Scraping and Crawling technologies.
  • Parallel Processing: Strong understanding of parallelism, concurrency, and threading.
  • Technical Stack: Expertise in Python-based libraries like Numpy, Pandas, Selenium, and Scrapy.
  • Database Management: Proficient in Postgres, MySQL, and NoSQL databases.
  • Cloud Technologies: Experience in AWS, GCP, and Azure, with hands-on knowledge of AWS Lambda, EC2, and S3.
  • REST API: Experience with developing and integrating REST APIs.
  • SQL Proficiency: Skilled in SQL procedures, query optimization, and pipeline creation.
  • Version Control: Familiarity with code versioning tools like Git and SVN.
  • SDLC Knowledge: Solid understanding of the Software Development Life Cycle (SDLC), testing methodologies, and CI/CD processes.
  • Problem-Solving Skills: Exceptional analytical and problem-solving abilities with a strong focus on delivering scalable data solutions.
Experience
  • Minimum 4-5 years of experience in ETL development, including experience with Bulk Data Extraction through web scraping and crawling technologies.
  • Proven track record of delivering large-scale data engineering solutions.
  • Prior experience with cloud environments like AWS, GCP, or Azure is highly preferred.
Why Join Us?
  • Opportunity to work on high-impact, large-scale data projects.
  • Collaborative, growth-oriented environment where innovation is encouraged.
  • Competitive compensation and benefits package.

We are seeking an exceptionally skilled Senior Data Engineer with extensive experience in ETL development and Bulk Data Extraction through Web Scraping and Crawling technologies. The ideal candidate will have a solid background in building, optimizing, and maintaining ETL pipelines, with strong expertise in data management and cloud technologies. This is a challenging role designed for a detail-oriented professional who is passionate about delivering high-quality data solutions.

You will be responsible for managing large-scale data projects, ensuring best practices in security, scalability, and performance. If you have a proven track record of delivering comprehensive data engineering solutions across diverse industries, we would love to hear from you.

Roles & Responsibilities

As a Senior Data Engineer, You Will:

  • ETL Pipeline Management:
  • Develop, optimize, and troubleshoot ETL pipelines across production, QA, and UAT environments.
  • Handle large datasets and manage server infrastructure effectively.
  • Web Scraping and Crawling:
  • Implement advanced web scraping techniques for rapid and efficient data extraction from large-scale sources across hundreds of websites.
  • Performance Optimization:
  • Apply techniques such as parallelism, concurrency, and threading to improve data extraction and processing efficiency.
  • Library and Framework Expertise:
  • Utilize relevant tools such as requests, xpath, ActionChain, Urllib, Numpy, Pandas, and lxml.
  • Work with frameworks like Selenium and Scrapy to build scalable scraping solutions.
  • Database Management:
  • Work with both relational databases (Postgres, MySQL) and NoSQL databases.
  • Optimize database performance, including schemas, triggers, and queries.
  • Cloud and API Integration:
  • Implement cloud best practices for security, scalability, and performance using AWS, GCP, or Azure.
  • Utilize AWS services such as Lambda, EC2, and S3 alongside external APIs.
  • Data Engineering:
  • Participate in data warehousing, architecture, and ETL pipeline development at an enterprise scale.
  • Create and optimize SQL procedures and data pipelines for seamless integration.
  • Collaboration and Project Delivery:
  • Work cross-functionally with teams to deliver end-to-end data solutions.
  • Ensure source data quality, develop mapping and transformations, and validate jobs across various environments.
To Be Successful In This Role, You Should Possess:
  • ETL Development Expertise: Proficiency in ETL pipeline development and optimization for large datasets.
  • Bulk Data Extraction: Experience with Web Scraping and Crawling technologies.
  • Parallel Processing: Strong understanding of parallelism, concurrency, and threading.
  • Technical Stack: Expertise in Python-based libraries like Numpy, Pandas, Selenium, and Scrapy.
  • Database Management: Proficient in Postgres, MySQL, and NoSQL databases.
  • Cloud Technologies: Experience in AWS, GCP, and Azure, with hands-on knowledge of AWS Lambda, EC2, and S3.
  • REST API: Experience with developing and integrating REST APIs.
  • SQL Proficiency: Skilled in SQL procedures, query optimization, and pipeline creation.
  • Version Control: Familiarity with code versioning tools like Git and SVN.
  • SDLC Knowledge: Solid understanding of the Software Development Life Cycle (SDLC), testing methodologies, and CI/CD processes.
  • Problem-Solving Skills: Exceptional analytical and problem-solving abilities with a strong focus on delivering scalable data solutions.
Experience
  • Minimum 4-5 years of experience in ETL development, including experience with Bulk Data Extraction through web scraping and crawling technologies.
  • Proven track record of delivering large-scale data engineering solutions.
  • Prior experience with cloud environments like AWS, GCP, or Azure is highly preferred.
Why Join Us?
  • Opportunity to work on high-impact, large-scale data projects.
  • Collaborative, growth-oriented environment where innovation is encouraged.
  • Competitive compensation and benefits package.

Custom Data Extraction

Price Intelligence

Data Analytics

Market Research

Industries

Retail

Automobile

Industrial Supply

Manufacturing

Delivery Methods

Data-as-a-Service

Software-as-a-Service

API

Company About Us

Media

Our Partners

Career

Our Brands

Customers

Contact Us

Resources

Blog

Case Study

Client Testimonial/Video

Terms & Conditions | Privacy Policy

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

WebDataGuru • Indiana (PA)

On-site
USD 100,000 - 150,000
Senior Data Engineer: ETL, Web Scraping, & Cloud DB
Senior Data Engineer: ETL, Web Scraping, & Cloud DB

WebDataGuru • Indiana (PA)

On-site
USD 110,000 - 170,000
ETL Developer
ETL Developer

Infinite Computer Solutions • Maryland

On-site
USD 110,000 - 160,000
Senior Data Engineer
Senior Data Engineer

AgileEngine, LLC • United States

On-site
USD 140,000 - 200,000
100% remote work with flexible hours
Data Engineering (Python + SQL with ETL)
Data Engineering (Python + SQL with ETL)

Cloud Analytics Technologies, LLC • Alpharetta (GA)

On-site
Data Engineer
Data Engineer

CEI • Richmond (VA)

Hybrid
USD 100,000 - 120,000
Data Engineer
Data Engineer

AgileEngine • United States

Hybrid
USD 110,000 - 140,000
Professional growth
Competitive compensation
Exciting project selection
+1
Data Integration Engineer
Data Integration Engineer

Business & Legal Resources • Dallas (TX)

On-site
USD 90,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
ETL Developer/Data Engineer
ETL Developer/Data Engineer

Meritore Technologies • Greenwood Village (CO)

On-site
USD 100,000 - 130,000