Databricks Architect

Tata Consultancy Services

Marlborough (MA)

On-site

USD 110,000 - 140,000

Full time

36 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary Annual Incentive
Comprehensive Medical Coverage
401K Plan
Certification & Training Reimbursement

Job summary

Tata Consultancy Services in Marlborough, MA seeks a senior data engineer with 10+ years of data engineering experience and 3+ years focused on Databricks/Spark. You will design scalable data pipelines, implement Delta Lake and Lakeflow, and govern data with Unity Catalog.

Proficiency in Python, PySpark, and SQL is essential, plus cloud experience and CI/CD practices. You will collaborate across teams to build Gold-layer datasets, ensure data quality, and optimize performance and costs in a

Qualifications

  • 10+ years hands-on data engineering experience with focus on Databricks/Spark.
  • Deep Databricks Lakehouse Platform expertise including Delta Lake.
  • Structured streaming, Lakeflow Declarative Pipelines, Databricks SQL, and cluster/serverless tuning.
  • Unity Catalog governance with metric views, domains, glossary for governed KPIs.
  • Experience with AI/BI Genie, Genie Spaces and Mosaic AI for AI/agent workloads.
  • Expert Python/PySpark, advanced SQL skills.
  • Data warehousing concepts (Kimball) and ETL/ELT design patterns.
  • Cloud provider experience (Azure/AWS/GCP) and S3-related services.
  • Software engineering mindset: Git, code reviews, testing, CI/CD.

Responsibilities

  • Data Pipeline Development: design, code, and deploy scalable batch/streaming pipelines with PySpark and Lakeflow.
  • Ingest data from POS, e-commerce, loyalty, and marketing clouds.
  • Data Modeling & Transformation in Medallion architecture (Bronze/Silver/Gold).
  • Define governed KPIs on Gold layer using Unity Catalog metric views.
  • Implement data quality and cleansing for Customer 360 data.
  • Tune Databricks jobs and Spark clusters for performance and cost-efficiency.
  • Manage IaC/CI-CD with Terraform and GitHub Actions; collaborate with DevOps.
  • Apply data governance features: lineage, fine-grained access, masking; ensure security.
  • collaborate with analysts to deliver consumption-ready datasets.
  • Govern Unity Catalog semantic modeling and Genie-enabled analytics delivery.

Skills

Python
PySpark
SQL
Cloud experience
Data Modeling
CI/CD practices
Communication

Tools

Databricks Lakehouse Platform
Delta Lake
Lakeflow Declarative Pipelines
Unity Catalog
Genie Spaces
Databricks One
Terraform

Job description


  • Experience: 10+ years of hands-on data engineering experience, with at least 3 years focused on the Databricks/Spark

  • Ecosystem

  • Databricks Expertise: Deep, hands-on expertise with the Databricks Lakehouse Platform, including Delta Lake,

  • Structured Streaming, Lakeflow Declarative Pipelines (formerly Delta Live Tables), Databricks SQL, and cluster/serverless configuration and optimization.

  • Business Semantics & Governance: Hands-on experience with Unity Catalog Business Semantics, including Metric Views (measures, dimensions, materialization), Domains, and Pages/Glossary to define governed, reusable KPIs once and serve them consistently across SQL, BI, and AI agents.

  • Agentic & Conversational Analytics: Working knowledge/exposure of AI/BI Genie (Genie Spaces, Genie Ontology, trusted assets), Genie Code for agentic pipeline/SQL development, Databricks One for business-user consumption, and Mosaic AI for building and serving AI/ML models and agents.

  • Programming Mastery: Expert-level proficiency in Python and PySpark. Advanced SQL skills are essential.

  • Data Warehousing Concepts: Strong understanding of data modeling principles, including dimensional modeling

  • (Kimball), data warehousing concepts, and ETL/ELT design patterns.

  • Cloud Proficiency: Proven experience working with a major cloud provider (Azure, AWS, or GCP), particularly with

  • data storage S3 and related services.

  • Software Engineering Mindset: Experience with software engineering best practices, including version control (Git),

  • code reviews, testing, and CI/CD.

  • Certification: Databricks


Job Description

Must Have Technical/Functional Skills


  • Experience: 10+ years of hands-on data engineering experience, with at least 3 years focused on the Databricks/Spark

  • Ecosystem

  • Databricks Expertise: Deep, hands-on expertise with the Databricks Lakehouse Platform, including Delta Lake,

  • Structured Streaming, Lakeflow Declarative Pipelines (formerly Delta Live Tables), Databricks SQL, and cluster/serverless configuration and optimization.

  • Business Semantics & Governance: Hands-on experience with Unity Catalog Business Semantics, including Metric Views (measures, dimensions, materialization), Domains, and Pages/Glossary to define governed, reusable KPIs once and serve them consistently across SQL, BI, and AI agents.

  • Agentic & Conversational Analytics: Working knowledge/exposure of AI/BI Genie (Genie Spaces, Genie Ontology, trusted assets), Genie Code for agentic pipeline/SQL development, Databricks One for business-user consumption, and Mosaic AI for building and serving AI/ML models and agents.

  • Programming Mastery: Expert-level proficiency in Python and PySpark. Advanced SQL skills are essential.

  • Data Warehousing Concepts: Strong understanding of data modeling principles, including dimensional modeling

  • (Kimball), data warehousing concepts, and ETL/ELT design patterns.

  • Cloud Proficiency: Proven experience working with a major cloud provider (Azure, AWS, or GCP), particularly with

  • data storage S3 and related services.

  • Software Engineering Mindset: Experience with software engineering best practices, including version control (Git),

  • code reviews, testing, and CI/CD.

  • Certification: Databricks


Roles & Responsibilities


  • Data Pipeline Development: Design, code, and deploy robust and scalable batch and streaming data pipelines

  • using PySpark, Spark SQL, and Lakeflow (Lakeflow Connect for ingestion and Lakeflow Declarative Pipelines) to ingest data from sources such as Point-of-Sale (POS), e-commerce platforms, loyalty systems, and marketing clouds.

  • Data Modeling & Transformation: Implement complex data transformations and business logic within the Medallion

  • architecture (Bronze, Silver, Gold layers). Build and optimize the final "Gold" dimension tables that will

  • serve as the single source of truth. Define governed business KPIs on top of the Gold layer using Unity Catalog Metric Views (business semantics) so metrics are computed once and reused consistently across BI, SQL, and Genie.

  • Data Quality: Implement data quality frameworks and cleansing routines to ensure the accuracy and trustworthiness

  • of the Customer 360 data.

  • Performance Optimization: Proactively monitor, debug, and tune Databricks jobs and Spark clusters for performance

  • and cost-efficiency. Implement best practices for partitioning, caching, and data layout in Delta Lake.

  • Infrastructure as Code (IaC) & CI/CD: Work with DevOps teams to manage Databricks environments, clusters, and

  • job deployments using tools like Terraform and AWS DevOps/GitHub Actions. Champion and implement CI/CD best

  • practices for data pipelines.

  • Data Governance & Security: Implement data governance features within Databricks Unity Catalog, including

  • data lineage tracking, fine-grained access controls (row/column-level security and ABAC policies), tags, and data masking to ensure compliance and security across BI and AI/agent workloads.

  • C ollaboration: Partner closely with Functional Consultants, Data Scientists, and Analytics Engineers to understand

  • their data requirements and deliver well-structured, consumption-ready datasets.

  • Semantic Modeling & Business Semantics: Build and govern Unity Catalog Business Semantics, authoring Metric Views (measures, dimensions, joins, synonyms, materialization), Domains, Pages, and Glossary terms, and certifying trusted assets so every dashboard, SQL query, notebook, and AI agent works from the same governed definitions.

  • Genie & Agentic Analytics Enablement: Configure and curate AI/BI Genie Spaces and the Genie Ontology (instructions, trusted assets, example queries, synonyms) to deliver accurate natural-language, conversational analytics for business users, and enable consumption through AI/BI Dashboards and Databricks One.


TCS Employee Benefits Summary


  • Discretionary Annual Incentive.

  • Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.

  • Family Support: Maternal & Parental Leaves.

  • Insurance Options: Aut& Home Insurance, Identity Theft Protection.

  • Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.

  • Time Off: Vacation, Time Off, Sick Leave & Holidays.

  • Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.


Salary Range-$110,000-$140,000 a year
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Databricks Engineer
Senior Databricks Engineer

Tata Consultancy Services • Malvern

On-site
USD 120,000 - 135,000
Discretionary incentive
Medical coverage
Parental leaves
+5
Databricks Data Engineer with a Solution Architect
Databricks Data Engineer with a Solution Architect

Tata Consultancy Services • Dublin (OH)

On-site
USD 140,000 - 150,000
Discretionary annual incentive
Comprehensive medical coverage
Family support leaves
+4
Databricks Data Engineer
Databricks Data Engineer

Berkley Alternative Markets IO • Chesterfield (MO)

On-site
USD 110,000 - 140,000
BI / Data Architect
BI / Data Architect

Tata Consultancy Services • Marlborough (MA)

On-site
USD 150,000 - 160,000
Discretionary annual incentive
Comprehensive medical coverage
Parental leave
+4
Sr. Manager, Data Engineering
Sr. Manager, Data Engineering

United States Digital Space LLC • United States

Remote
USD 168,000 - 187,000
Collaborative environment
Yearly bonus
Comprehensive benefits
+3
Data Engineer
Data Engineer

Sonobello • Bellevue (WA)

On-site
USD 125,000 - 140,000
Medical, Dental, Vision Insurance
401(k) with match
Paid time off
+2
Global GTM Enablement & Scale Architect
Global GTM Enablement & Scale Architect

Menlo Ventures • United States

On-site
USD 217,000 - 300,000
Staff Designated Support Engineer
Staff Designated Support Engineer

Databricks Inc. • San Francisco (CA)

On-site
USD 141,000 - 251,000
Go-To-Market (GTM) Digital Natives Program Leader
Go-To-Market (GTM) Digital Natives Program Leader

Menlo Ventures • San Francisco (CA)

On-site
USD 269,000 - 371,000
Comprehensive benefits
Annual performance bonus
Equity options
Delivery Solutions Architect - Healthcare & Life Sciences
Delivery Solutions Architect - Healthcare & Life Sciences

Databricks • United States

Hybrid
USD 180,000 - 248,000