Senior PySpark Data Developer

Tata Consultancy Services

Irving (TX)

On-site

USD 90,000 - 150,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
Auto & Home Insurance Options
Certification & Training Reimbursement
Vacation & Holidays
401K Plan

Job summary

Tata Consultancy Services is seeking a Data Engineer to design, build, and optimize scalable data pipelines in a cloud environment. You will work with PySpark, Spark SQL, and HiveQL to ensure data reliability and performance for BI and analytics initiatives.

Ideal candidates will have hands-on experience with AWS/Azure/GCP, data warehousing fundamentals, and a strong focus on performance tuning and data quality.

Qualifications

  • Strong experience building scalable data pipelines and ETL/ELT processes.
  • Experience with Spark, PySpark, HiveQL, and SQL in cloud environments.
  • Knowledge of data warehousing concepts, data partitioning, and schema design.

Responsibilities

  • Design, build, and optimize scalable ETL/ELT pipelines using PySpark and Spark SQL.
  • Manage and scale cloud data infrastructure (AWS, Azure, GCP).
  • Optimize data storage with Hive, Parquet/ORC/Avro formats.
  • Tune performance and resolve bottlenecks in Spark jobs.

Skills

Apache Spark
Python/PySpark
HiveQL/SQL
Data Modeling
Cloud Platforms
Spark Tuning
Data Pipelines

Education

Bachelor's degree in Computer Science or related field

Tools

Airflow
ETL/ELT Tools
Parquet/ORC/Avro
Spark UI

Job description

Job Description

We are seeking a highly skilled and motivated Data Engineer to play a pivotal role in designing, building, and optimizing our next-generation scalable data pipelines. This position requires expertise in processing massive datasets using cutting-edge technologies like Apache Spark, PySpark, and Hive within a dynamic cloud environment. Your primary objective will be to ensure the utmost data reliability, speed, and efficiency, providing a robust foundation for downstream business intelligence and advanced analytics initiatives.

Roles & Responsibilities
  • Data Pipeline Development & Maintenance: Design, build, and maintain highly scalable and efficient ETL/ELT data pipelines utilizing PySpark and Spark SQL for complex data transformations.
  • Cloud Data Infrastructure Management: Deploy, manage, and scale critical data infrastructure components on leading cloud platforms such as Amazon Web Services (AWS) (e.g., EMR, Glue), Microsoft Azure (e.g., Databricks, Synapse), or Google Cloud Platform (GCP).
  • Data Warehousing & Storage Optimization: Strategically manage data layout, partitioning, and indexing within Apache Hive and various cloud data lake solutions to optimize performance and accessibility.
  • Performance Tuning & Optimization: Proactively identify and resolve performance bottlenecks in Spark jobs, leveraging Spark UI for in-depth analysis, effectively managing data skewness, and optimizing memory utilization.
  • Diverse Data Integration: Develop robust solutions for ingesting high-volume and diverse datasets from both structured relational databases and unstructured flat files into our data ecosystem.
  • Automated Workflow Orchestration: Implement and manage automated data workflows using industry-standard scheduling tools like Apache Airflow or platform-native schedulers, ensuring timely and reliable data delivery.
  • Strategic Collaboration: Partner closely with data scientists, business analysts, and cross-functional enterprise teams to translate complex business requirements into technically sound and efficient data solutions.
Qualifications
  • Big Data Frameworks Expertise: Demonstrated high proficiency in Apache Spark architecture, including a deep understanding of drivers, executors, and Directed Acyclic Graphs (DAGs).
  • Advanced Programming: Exceptional coding skills in Python and extensive experience with the PySpark API for developing intricate data transformations and processing logic.
  • Querying & Schema Management: Strong command of HiveQL and ANSI SQL, coupled with expertise in data partitioning techniques and effective schema definition.
  • Optimized Storage Formats: In-depth understanding and practical experience with optimized big data storage file formats such as Parquet, ORC, and Avro.
  • Cloud Ecosystem Development: Hands-on development experience utilizing cloud-native big data utilities (e.g., AWS EMR, Azure Databricks) with in major cloud platforms.
  • Data Warehousing Fundamentals: Solid foundation in Dimensional Data Modeling, including Star and Sno wflake schemas, and practical experience with Data Lakes concepts and implementation.
Preferred Qualifications
  • CI/CD & DevOps Automation: Experience with Continuous Integration/Continuous Deployment (CI/CD) practices and automation tools like Git, Jenkins, or Ansible.
  • NoSQL Database Integration: Exposure to and experience with NoSQL databases such as HBase, Cassandra, or MongoDB.
  • Professional Cloud Certifications: Relevant professional cloud certifications (e.g., AWS Certified Data Engineer, Microsoft Certified: Azure Data Engineer Associate) are highly valued
TCS Employee Benefits Summary
  • Discretionary Annual Incentive.
  • Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
  • Family Support: Maternal & Parental Leaves.
  • Insurance Options: Auto & Home Insurance, Identity Theft Protection.
  • Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
  • Time Off: Vacation, Time Off, Sick Leave & Holidays.
  • Legal & Financial Assistance: Legal Assistance, 401K Pl
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Pyspark Developer
Pyspark Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Developer
Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Developer
Developer

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 100,000 - 130,000
Engineer
Engineer

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
Data Engineer
Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Data & Software Engineer
Data & Software Engineer

Avalore, LLC • Chantilly (VA)

On-site
USD 120,000 - 190,000
Health insurance
401(k) plan
Life insurance
+4
PySpark Developer
PySpark Developer

Inizio Partners Corp • Hartford (CT)

On-site
USD 90,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Tata Consultancy Services • Chicago (IL)

On-site
USD 100,000 - 120,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Maternal & Parental Leaves
+2
Senior Data Engineer
Senior Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 120,000 - 130,000
Discretionary Annual Incentive
Medical Coverage: Medical & Health, +
Family Leaves
+4
BI / Data Architect
BI / Data Architect

Tata Consultancy Services • Marlborough (MA)

On-site
USD 150,000 - 160,000
Discretionary annual incentive
Comprehensive medical coverage
Parental leave
+4