Data Engineer

Tata Consultancy Services

Irving (TX)

On-site

USD 125,000 - 140,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tata Consultancy Services in Irving, TX is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Spark, PySpark, and Hive within Cloudera-like environments. Strong programming in Python, Spark SQL, and familiarity with Parquet/ORC formats are required, as is experience with Airflow or similar schedulers.

This role offers a competitive salary and opportunities to work on cutting-edge data platforms.

Qualifications

  • Proficient in Spark architecture, drivers, executors, and DAGs.
  • Strong Python and PySpark programming skills for complex transformations.
  • Strong HiveQL/ANSI SQL with partitioning and schema design.
  • Experience with Parquet/ORC/Avro storage formats.
  • Solid foundation in dimensional modeling and data lake concepts.

Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines using PySpark and Spark SQL, Hive.
  • Optimize data warehouse layouts, partitioning, and indexing for performance.
  • Tune Spark jobs, monitor via Spark UI, and address memory issues and skew.
  • Ingest high-volume structured and unstructured data into the data ecosystem.
  • Automate data workflows with Airflow or platform schedulers for reliable delivery.
  • Collaborate with data scientists and analysts to translate business requirements into data solutions.

Skills

Spark Architecture
Python / PySpark
HiveQL / ANSI SQL
Parquet/ORC/Avro
Dimensional Modeling

Education

Bachelor's degree in Computer Science

Tools

Git
Jenkins
Ansible
AWS EMR
Databricks

Job description

Job Description: We are seeking a highly skilled and motivated Data Engineer to play a pivotal role in designing, building, and optimizing our next-generation scalable data pipelines. This position requires expertise in processing massive datasets using cutting-edge technologies like Apache Spark, PySpark, and Hive within Cloudera Platform. Your primary objective will be to ensure the utmost data reliability, speed, and efficiency, providing a robust foundation for downstream business intelligence and advanced analytics initiatives.

  • Data Pipeline Development & Maintenance: Design, build, and maintain highly scalable and efficient ETL/ELT data pipelines utilizing PySpark and Spark SQL , Hive for complex data transformations.
  • Data Warehousing & Storage Optimization: Strategically manage data layout, partitioning, and indexing within Apache Hive and various cloud data lake solutions to optimize performance and accessibility.
  • Performance Tuning & Optimization: Proactively identify and resolve performance bottlenecks in Spark jobs, leveraging Spark UI for in-depth analysis, effectively managing data skewness, and optimizing memory utilization.
  • Diverse Data Integration: Develop robust solutions for ingesting high-volume and diverse datasets from both structured relational databases and unstructured flat files into our data ecosystem.
  • Automated Workflow Orchestration: Implement and manage automated data workflows using industry-standard scheduling tools like Apache Airflow or platform-native schedulers, ensuring timely and reliable data delivery.
  • Strategic Collaboration: Partner closely with data scientists, business analysts, and cross-functional enterprise teams to translate complex business requirements into technically sound and efficient data solutions.
Job Description

We are seeking a highly skilled and motivated Data Engineer to play a pivotal role in designing, building, and optimizing our next-generation scalable data pipelines. This position requires expertise in processing massive datasets using cutting-edge technologies like Apache Spark, PySpark, and Hive within Cloudera Platform. Your primary objective will be to ensure the utmost data reliability, speed, and efficiency, providing a robust foundation for downstream business intelligence and advanced analytics initiatives.

Roles & Responsibilities

Job Title: Data Engineer

  • Data Pipeline Development & Maintenance: Design, build, and maintain highly scalable and efficient ETL/ELT data pipelines utilizing PySpark and Spark SQL , Hive for complex data transformations.
  • Data Warehousing & Storage Optimization: Strategically manage data layout, partitioning, and indexing within Apache Hive and various cloud data lake solutions to optimize performance and accessibility.
  • Performance Tuning & Optimization: Proactively identify and resolve performance bottlenecks in Spark jobs, leveraging Spark UI for in-depth analysis, effectively managing data skewness, and optimizing memory utilization.
  • Diverse Data Integration: Develop robust solutions for ingesting high-volume and diverse datasets from both structured relational databases and unstructured flat files into our data ecosystem.
  • Automated Workflow Orchestration: Implement and manage automated data workflows using industry-standard scheduling tools like Apache Airflow or platform-native schedulers, ensuring timely and reliable data delivery.
  • Strategic Collaboration: Partner closely with data scientists, business analysts, and cross-functional enterprise teams to translate complex business requirements into technically sound and efficient data solutions.
Qualifications
  • Big Data Frameworks Expertise: Demonstrated high proficiency in Apache Spark architecture, including a deep understanding of drivers, executors, and Directed Acyclic Graphs (DAGs).
  • Advanced Programming: Exceptional coding skills in Python and extensive experience with the PySpark API for developing intricate data transformations and processing logic.
  • Querying & Schema Management: Strong command of HiveQL and ANSI SQL, coupled with expertise in data partitioning techniques and effective schema definition.
  • Optimized Storage Formats: In-depth understanding and practical experience with optimized big data storage file formats such as Parquet, ORC, and Avro.
  • Data Warehousing Fundamentals: Solid foundation in Dimensional Data Modeling, including Star and Snowflake schemas, and practical experience with Data Lakes concepts and implementation.
Preferred Qualifications
  • CI/CD & DevOps Automation: Experience with Continuous Integration/Continuous Deployment (CI/CD) practices and automation tools like Git, Jenkins, or Ansible.
  • Cloud Ecosyste m Development: Experience in development experience utilizing cloud-native big data utilities (e.g., AWS EMR, AWS Databricks) within major cloud platforms.
  • NoSQL Database Integration: Exposure to and experience with NoSQL databases such as HBase, Cassandra, or MongoDB.
  • Professional Certifications: Relevant professional certifications on Spark or Data Engineer are highly valued

Salary Range: $125,000 to $140,000 per year

Qualifications: BACHELOR OF COMPUTER SCIENCE

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Siri InfoSolutions Inc • Town of Texas (WI)

On-site
USD 110,000 - 160,000
Developer
Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Engineer
Engineer

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
Developer
Developer

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 100,000 - 130,000
Pyspark Developer
Pyspark Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 120,000 - 170,000
Data Engineer
Data Engineer

Jobtailor • Kentucky

On-site
USD 110,000 - 140,000
Data scientist
Data scientist

OVA.Work • Alpharetta (GA)

On-site
USD 80,000 - 110,000
Data Engineer
Data Engineer

Aptdata Solutions Inc. • Farmington Hills (MI)

On-site
USD 90,000 - 115,000
Engineer
Engineer

Tata Consultancy Services • Dallas (TX)

On-site
USD 100,000 - 105,000