Get more replies from employers
Send a job-specific resume in minutes.
Tata Consultancy Services in Malvern, PA is seeking a skilled data engineer to design and optimize end-to-end data pipelines using Spark, Delta Lake, and Databricks on AWS. You will build Bronze-Silver-Gold medallion lakehouse architectures, implement security with Unity Catalog and IAM, and integrate with AWS services like S3 and Redshift.
Collaborate with data scientists and analysts, drive CI/CD with Git, and apply MLflow for experiment tracking and GenAI initiatives.
Must Have Technical/Functional Skills
Apache Spark: Mastery of Spark SQL and the DataFrame API for distributed data processing, performance tuning, and optimizing execution plans.
Delta Lake & Lakehouse: Building medallion architectures (Bronze, Silver, Gold layers) and using optimizations like Z-Ordering and ACID transactions
Pipelines & Orchestration: Developing with Delta Live Tables (DLT) and Databricks Workflows to design resilient, idempotent ELT/ETL pipelines.
Unity Catalog: Configuring data governance, row/column-level security, and data lineage
Languages: Advanced proficiency in Python (PySpark) and SQL
Storage & Data Lakes: Expert in reading/writing to Amazon S3 and integrating with AWS data warehouses (e.g., Amazon Redshift).
Security & IAM: Implementing cross-account IAM roles, S3 bucket policies, and Customer-Managed Keys (CMK) via AWS KMS
Ecosystem Integration: Familiarity with integrating Databricks jobs alongside native services like AWS Glue, Amazon Kinesis, and AWS Step Functions
CI/CD & Version Control: Git-based workflows using Databricks Git Folders and managing deployments.
AI & MLOps: Using MLflow for experiment tracking and model registries.
Familiarity with LLMs, Vector Search, and GenAI integrations.
Data Pipeline Development: Extract, transform, and load (ETL) data from multiple sources, building batch and streaming pipelines.
Spark & Code Optimization: Write highly efficient PySpark, Scala, or SQL code. Troubleshoot distributed processing bottlenecks and optimize job performance.
Lakehouse Management: Work with the Medallion architecture to transition data through Bronze (raw), Silver (cleansed), and Gold (aggregated) layers in Delta Lake.
Orchestration & Automation: Schedule and automate workflows using Databricks Jobs, Workflows, or Delta Live Tables (DLT). Configure automated tests, retries, and error handling.
Governance & Security: Implement data privacy and access controls utilizing the Databricks Unity Catalog, ensuring compliance and secure data sharing.
Collaboration: Work alongside Data Scientists, Data Analysts, and ML Engineers to prepare data features and support BI reporting, ML training, and generative AI initiatives.
Discretionary Annual Incentive.
Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
Family Support: Maternal & Parental Leaves.
Insurance Options: Auto & Home Insurance, Identity Theft Protection.
Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
Time Off: Vacation, Time Off, Sick Leave & Holidays.
Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.
Salary Range: $120,000 - 135,000 a year
BACHELOR OF COMPUTER SCIENCE