Senior Data Engineer

Taskus India

United States

Remote

USD 140,000 - 190,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Taskus India is seeking a Senior Data Engineer to architect and operate resilient data pipelines in a Lakehouse environment. You will translate architects' blueprints into production-grade code, delivering clean, high-fidelity data for analytics and decision making.

You will own pipeline engineering, use dbt and PySpark, implement data SLAs, and collaborate with BI and data science teams to produce feature-ready datasets, while optimizing SQL and Spark plans to reduce cost and latency.

Qualifications

  • 6+ years of data engineering experience required.
  • Proficiency in Python, SQL, and PySpark.
  • Experience with modern lakehouse stack (Databricks or Snowflake).
  • Strong skills in dbt, CI/CD, and data modeling.

Responsibilities

  • Build and maintain high-throughput ETL/ELT pipelines to the Lakehouse.
  • Drive data quality with data SLAs and observability.
  • Collaborate with BI and data science teams to produce feature-ready datasets.
  • Optimize SQL and Spark plans to reduce cost and latency.

Skills

Data Engineering
Python
SQL
PySpark
Databricks
Snowflake
dbt
Terraform
Airflow
Kubernetes

Tools

DuckDB
Kubernetes
Airflow
Databricks
Snowflake
Kafka
dbt
Terraform

Job description

The Mission:

As a Senior Data Engineer, you are the engine room of our data strategy. You don't just "move data"; you build resilient, self-healing systems that transform raw, into high-fidelity data products. You will take the Architects blueprints and turn them into production-grade code, ensuring there is clean and reliable data.

What the Senior Data Engineer will own:
  • Pipeline Engineering: Build and maintain high-throughput ETL/ELT pipelines that ingest data from various sources into our Lakehouse.
  • Code Quality & Tooling: Drive the adoption of dbt for transformation and PySpark for heavy lifting. You will be responsible for writing modular, reusable code that follows strict CI/CD practices.
  • Observability & Reliability: Implement "Data SLAs." You will build the monitoring and alerting systems that notify the team of data drift or pipeline failures before the business notices.
  • Data Productization: Work closely with the BI team and Data Scientists to prepare "feature-ready" datasets.
  • Performance Tuning: Deep-dive into SQL and Spark query plans to optimize slow-running jobs, reducing cost and latency.
  • Local Development Advocacy: Implementing a workflow where engineers can develop and test complex SQL transformations locally using DuckDB, drastically reducing "waiting-for-cluster" time and lowering development costs.
  • Cost-Efficient Micro-Pipelines: Building specialized pipelines for specific client reporting or data-quality checks that run on single-node containers (e.g., Lambda or small ECS tasks) using DuckDB, avoiding the minimum billing cycles of larger warehouses.
  • Embedded Analytics: Exploring ways to use DuckDB as an embedded engine for internal tools or "edge" processing within the BPOs local site offices where bandwidth to the cloud might be a constraint.
Technical Skills & Experience:
  • 6+ Years in Data Engineering: You’ve lived through the "on-call" life and know how to build systems that don't break at 3 AM.
  • The Power Trio: Expert-level proficiency in Python, SQL, and PySpark.
  • Modern Lakehouse Stack: Hands-on experience with Databricks (Delta Lake) or Snowflake. You understand the nuances of the Medallion Architecture (Bronze/Silver/Gold).
  • Transformation & Modeling: Advanced experience with dbt (Data Build Tool). You treat data models like software, including version control, testing, and documentation.
  • Orchestration: Experience with Apache Airflow or Prefect, specifically building complex, idempotent DAGs.
  • Streaming Experience: Familiarity with Kafka, Kinesis, or Spark Streaming is a huge plus—our BPO operations move in real-time.
  • Infrastructure as Code (IaC): Comfort with Terraform or CloudFormation to manage your own data infrastructure.
  • In-Process Analytics (DuckDB): Proven experience using DuckDB for high-speed local development, unit testing data transformations, or as a query engine for "small-to-medium" datasets (up to 100GB) without the overhead of a distributed cluster.
  • Hybrid Execution Patterns: Ability to identify when to use heavyweight compute (PySpark/Databricks) versus lightweight compute (DuckDB/Python) to minimize cloud costs and reduce job latency.
  • Parquet/Iceberg Interaction: Experience using DuckDB to directly query data stored in S3/Azure Blob (via Parquet or Iceberg files) for rapid ad-hoc analysis or local dashboarding.
  • Orchestration: Kubernetes (K8s). Deep understanding of Pods, Deployments, Services, ConfigMaps, and Secrets management.Role & responsibilities
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer – AWS Lakehouse (Mandarin Required)
Data Engineer – AWS Lakehouse (Mandarin Required)

Bitus Labs • Irvine (CA)

On-site
USD 120,000 - 180,000
AWS Lakehouse Data Engineer
AWS Lakehouse Data Engineer

engineeringjobs.net, Inc. • Atlanta (GA)

Remote
USD 140,000 - 200,000
Senior Data Engineer
Senior Data Engineer

Madison-Davis, LLC • Chicago (IL)

On-site
USD 130,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Peyton Resource Group • Houston (TX)

On-site
USD 120,000 - 150,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Senior Data Engineer
Senior Data Engineer

Dynata, LLC (Connecticut) • United States

Remote
USD 130,000 - 150,000
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Databricks Data Engineer
Databricks Data Engineer

Compunnel, Inc. • Spring (TX)

On-site
USD 110,000 - 140,000
Data Engineer
Data Engineer

InfoVision, Inc. • United States

On-site
USD 110,000 - 150,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000