Hadoop Hive Python Developer

Tata Consultancy Services

Charlotte (NC)

On-site

USD 110,000 - 125,000

Full time

17 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Comprehensive Medical Coverage
401K Plan
Training Reimbursement
Vacation and Holidays

Job summary

Tata Consultancy Services in Charlotte, NC seeks a Senior Big Data Engineer with 9-14 years of experience to architect and deliver scalable data platforms for large financial systems. You will design end-to-end data pipelines, enforce governance, and collaborate with quants, risk teams, and product owners across cloud and on-prem environments.

Key focus areas include Databricks Lakehouse, Delta Lake, Bronze/Silver/Gold modeling, and real-time ingestion using Kafka.

Qualifications

  • 9-14 years of hands-on Big Data engineering experience.
  • Strong expertise in Databricks Lakehouse and Delta Lake optimizations.
  • Experience designing data solutions for large-scale financial systems.
  • Proficiency with SQL on massive TB/PB datasets and CI/CD practices.

Responsibilities

  • Design, develop, and optimize PySpark ETL pipelines on on-prem and cloud.
  • Build high-volume ingestion frameworks using Kafka for real-time data.
  • Develop and tune Hadoop ecosystem components (HDFS, YARN, MapReduce, Tez, Oozie/Airflow).
  • Create high-performance Hive data models for regulatory and risk processing.

Skills

PySpark
Kafka
Hadoop
Hive
SQL
Databricks Lakehouse
Delta Lake
Data Modeling
ETL Pipelines
Real-time Ingestion
CI/CD
Cloud & On-Prem

Education

Bachelor of Computer Science

Tools

Git
Jenkins
Bitbucket
Airflow/Oozie
Unity Catalog

Job description

Job Description

Must Have Technical/Functional Skills

Primary skills: Hadoop, Hive, Python, PySpark, Apache Kafka, Hadoop Ecosystem, Hive, Databricks Lakehouse Architecture, Delta Lake, Bronze/Silver/Gold Data Modeling, Big Data ETL Pipeline Development, SQL, Real-time Data Ingestion Frameworks, Data Governance & Cataloging, CI/CD Tools – Git, Jenkins, Bitbucket, Workflow Orchestration, and Cloud & On-Prem Big Data Platforms.

Experience: Minimum 9+ years

Roles & Responsibilities

Seeking a Senior Big Data Engineer with 9-14 years of experience specializing in Hadoop, Python, Hive PySpark, Kafka, and strong experience designing data solutions for large-scale financial systems.

In addition, the candidate must possess advanced expertise in Databricks Lakehouse architecture, particularly around Bronze/Silver/Gold layer data modeling, Delta Lake optimizations, and building reliable, scalable pipelines for regulatory, risk, trading, and analytics workloads.

This role focuses on delivering highly performant, well-governed data platforms that support the bank’s mission-critical global markets functions.

Key Responsibilities
  • Design, develop, and optimize PySpark-based ETL pipelines running on on-prem Hadoop clusters and cloud environments.
  • Build high-volume ingestion frameworks using Kafka for real-time and near-real-time trading and market data.
  • Develop, tune, and manage Hadoop ecosystem components—HDFS, YARN, MapReduce, Tez, Oozie/Airflow.
  • Build high-performance, optimized Hive data models for regulatory reporting, trade lifecycle, and market risk processing.
Databricks Lakehouse & Delta Framework
  • Architect and implement Bronze/Silver/Gold layer modeling patterns within the Databricks Lakehouse.
  • Apply Delta Lake best practices including:
  • optimized file management
  • Z-Ordering
  • Delta Change Data Feed (CDF)
  • schema evolution & enforcement
  • ACID transaction handling
  • Build reusable frameworks for ingestion, cleansing, transformation, and consumption of data across Lakehouse layers.
  • Enable governance, lineage, and auditability using Unity Catalog or equivalent cataloging tools.
Collaboration, Leadership & Delivery
  • Collaborate closely with quants, product owners, architects, risk tech, and business users.
  • Participate in agile ceremonies — sprint planning, refinement, design reviews.
  • Mentor junior engineers and contribute to building strong engineering practices across tech teams.
Required Skills & Experience
  • 9-14 years of hands-on experience in Big Data engineering.
  • Expert skills in:
  • PySpark — dataframe optimizations, partitioning, broadcast strategies, distributed computing.
  • Kafka — producer/consumer design, schema registry, streaming ETLs.
  • Hadoop ecosystem — HDFS, YARN, MapReduce/Tez, Oozie/Airflow.
  • Hive — advanced query tuning, TEZ optimization, partition/bucket management.
  • Extensive hands-on experience with Databricks Lakehouse, including:
  • Bronze/Silver/Gold layer modeling
  • Delta Lake optimizations
  • Data quality frameworks on Lakehouse
  • Structured & unstructured data handling
  • Experience in Global Markets, Risk, Treasury, Trade Surveillance, or Regulatory Reporting.
  • Strong SQL knowledge with experience working on massive datasets (TB/PB scale).
  • Experience with CI/CD practices — Git, Jenkins, Bitbucket, build pipelines.
TCS Employee Benefits Summary
  • Discretionary Annual Incentive.
  • Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
  • Family Support: Maternal & Parental Leaves.
  • Insurance Options: Auto & Home Insurance, Identity Theft Protection.
  • Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
  • Time Off: Vacation, Time Off, Sick Leave & Holidays.
  • Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.

Salary Range: $110,000-$125,000 a year

Qualifications

BACHELOR OF COMPUTER SCIENCE

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ETL Developer
ETL Developer

Tata Consultancy Services • Cleveland (OH)

On-site
USD 80,000 - 140,000
Discretionary Annual Incentive
Medical Coverage
Family Leaves
+4
ETL Developer
ETL Developer

Tata Consultancy Services • Plano (TX)

On-site
USD 80,000 - 140,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Dental & Vision Coverage
+1
Python Hadoop Engineer
Python Hadoop Engineer

Tata Consultancy Services • Charlotte (NC)

On-site
USD 100,000 - 110,000
Discretionary annual incentive
Comprehensive medical coverage
Family leaves
+4
Hadoop developer
Hadoop developer

Tata Consultancy Services • Charlotte (NC)

On-site
USD 95,000 - 115,000
Annual incentive
Medical coverage
Parental leaves
+2
Engineer
Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 110,000
Discretionary Annual Incentive.
Comprehensive Medical Coverage
401K Plan
+1
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Senior PySpark Data Developer
Senior PySpark Data Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+4
Senior Developer
Senior Developer

Tata Consultancy Services • Charlotte (NC)

On-site
USD 110,000 - 120,000
Discretionary Annual Incentive.
Comprehensive Medical Coverage
Family Support: Parental Leaves
+3
Data Engineer
Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Data Engineer
Data Engineer

Infinite Computer Solutions • Town of Texas (WI)

On-site
USD 120,000 - 170,000