Data Platform Architect

TaskUs

India

On-site

INR 4,000,000 - 8,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TaskUs is seeking an accomplished Data Platform Architect to drive platform design for large-scale lakehouse environments. You will decide compute engines, table formats like Iceberg or Delta Lake, and tiered storage strategies to support 1,000+ concurrent users.

In this role you will champion an agnostic data infrastructure, establish automated governance, and define coding standards, CI/CD patterns, and documentation requirements.

Qualifications

  • 8+ years of data engineering/architecture delivering production-grade lakehouse environments for high-concurrency orgs.
  • Modern data stack fluency with Databricks and high-performance warehouses (Redshift/Snowflake).
  • Expertise with open-table formats like Iceberg or Delta Lake and schema evolution.

Responsibilities

  • Platform Design & Research: Lead architectural decisions on compute engine selection, table formats, and tiered storage design.
  • Agnostic Infrastructure: Architect a decoupled data environment to ensure interoperability and avoid vendor lock-in.
  • Governance & Compliance: Design automated data governance including PII discovery, security, and auditability.
  • Standards & Frameworks: Define DoD for data pipelines, coding standards, CI/CD patterns, and docs.
  • ADR: Maintain a versioned repo of architectural decisions with rationales and trade-offs.
  • Performance & FinOps: Optimize platform performance and cost for sub-second queries at scale.
  • Technical Stewardship: Lead deep-dive reviews and mentor senior engineers.

Skills

Data Engineering
Data Architecture
Lakehouse
Databricks
Unity Catalog
Delta Lake
Iceberg
dbt Core
PySpark
DuckDB
Airflow
AWS
Azure/GCP
Kubernetes
Open-source Catalogs

Tools

Apache Iceberg
Delta Lake
dbt Core
PySpark
Apache Airflow
AWS
Azure/GCP
DuckDB

Job description

  • Platform Design & Research: Lead architectural decisions regarding compute engine selection, open-table format implementation, and tiered storage design.
  • Agnostic Infrastructure: Architect a decoupled data environment that ensures interoperability across multiple engines and prevents proprietary vendor lock-in.
  • Governance & Compliance: Design and oversee the implementation of automated data governance, including PII discovery, row/column-level security, and auditability.
  • Standards & Frameworks: Define the \"Definition of Done\" for data pipelines, establishing coding standards, CI/CD patterns, and technical documentation requirements.
  • Architectural Decision Records (ADR): Maintain a version-controlled repository of all consequential technical decisions, documenting the rationale, trade-offs, and long-term implications.
  • Performance & FinOps: Monitor and optimize platform performance and spend, ensuring sub-second query speeds for massive user bases while maintaining a lean cloud footprint.
  • Technical Stewardship: Conduct deep-dive code and design reviews for all data models and orchestration workflows; mentor and unblock senior engineering staff.
Technical Skills & Experience
  • :8+ Years in Data Engineering / Architecture: Proven experience delivering production-grade Lakehouse environments for high-concurrency (1,000+ user) organizations
  • .Modern Data Stack Fluency: Extensive experience with Databricks (Lakehouse/Unity Catalog) and high-performance warehouses like Amazon Redshift or Snowflake
  • .Open-Table Formats: Deep hands-on expertise with Apache Iceberg or Delta Lake, including optimization strategies for partitioning and schema evolution
  • .Transformation & Modeling: Mastery of dbt (Core) for complex SQL-based modeling and PySpark or Python for sophisticated data processing
  • .High-Efficiency Compute: Familiarity with vectorized/embedded engines like DuckDB for specialized or cost-sensitive processing tasks
  • .Orchestration Mastery: Advanced experience with Apache Airflow, specifically in designing resilient, dependency-aware DAGs in resource-constrained environments
  • .Cloud Ecosystems: Expert-level knowledge of AWS (S3, EC2, IAM) or equivalent services in Azure/GCP, with a focus on storage-compute separation
  • .Experience with open-source catalog implementations like Apache Polaris
  • .Knowledge of Data Ops principles and automated data quality testing frameworks
  • .Experience translating technical debt and architectural roadmaps for C-level executives
  • .Background in managing fixed-resource infrastructure (e.g., EC2/VM-based processing) vs. elastic serverless models
  • .Kubernetes (K8s). Deep understanding of Pods, Deployments, Services, ConfigMaps, and Secrets management
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Engineer
Sr. Data Engineer

R Systems • Dadri

On-site
INR 1,200,000 - 1,800,000
Data Engineering Architect
Data Engineering Architect

ADP • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Data Engineer
Data Engineer

Bounteous • Chennai District

On-site
INR 2,000,000 - 3,600,000
Data Architect
Data Architect

ReNew • Gurugram District

On-site
INR 4,000,000 - 6,000,000
Senior Data Engineer - Databricks, PySpark & Lakehouse
Senior Data Engineer - Databricks, PySpark & Lakehouse

Tata Consultancy Services • Bengaluru, Pune District, Chennai District

On-site
INR 1,500,000 - 3,000,000
Resident Solution Architect
Resident Solution Architect

Celebal Technologies • Dadri

On-site
INR 3,000,000 - 6,000,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE Clear Europe Limited • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Manager - Data Engineering
Manager - Data Engineering

HD Supply • Chennai District

On-site
INR 1,500,000 - 2,200,000
DATA ARCHITECT - Databricks
DATA ARCHITECT - Databricks

Happiest Minds Technologies • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE • Hyderabad

On-site
INR 1,800,000 - 3,000,000