Data Analyst

Keka Technologies Private Limited

United States

On-site

USD 140,000 - 190,000

Full time

39 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Keka Technologies Private Limited is seeking a senior data engineer to build, optimize, and operate production Spark pipelines that generate attribution feeds from large-scale data. You will design data processing across modern data lake and warehouse tech, orchestrated with AWS workflows.

The role involves building partner integrations, maintaining REST APIs (PHP/Symfony), diagnosing failures, backfills, and delivering automated tests.

Qualifications

  • 7+ years of professional software or data engineering experience.
  • Advanced SQL, relational data modeling, and large-scale data processing experience.

Responsibilities

  • Build, optimize, and operate production Apache Spark pipelines for attributed conversion feeds.
  • Design and evolve data processing across data lake/warehouse with AWS-based orchestration.
  • Develop and operate REST APIs, including PHP-based APIs, with secure, backward-compatible changes.
  • Diagnose production failures, backfills, and manage staged releases and rollbacks.
  • Write unit and integration tests for pipelines, APIs, and feeds.

Skills

Java
Python
SQL
Spark

Tools

Snowflake
MySQL
Iceberg
Airflow
Symfony
PHP

Job description

  • Build, optimize, and operate production Apache Spark pipelines that generate attributed conversion and optimization feeds from large-scale impression and conversion data.
  • Design and evolve data processing and models across modern data lake and warehouse technologies, orchestrated with a production workflow scheduler on AWS.
  • Build and operate partner optimization feeds and integrations with external partners, including schemas, identity fields and unique IDs, file delivery, reconciliation, and SLAs.
  • Build, change, and operate production REST APIs, including PHP APIs, delivering secure, backward-compatible changes end to end (implementation, validation, authorization, data access, documentation, automated testing, and production diagnostics).
  • Diagnose production failures, reconcile partner-facing data, execute backfills, and manage staged releases and rollbacks.
  • Write and maintain unit and integration tests across pipelines, APIs, and feeds.
Qualifications and Education Requirements:
  • 7+ years of professional software or data engineering experience.
  • 4+ years building, optimizing, and maintaining production Apache Spark pipelines, with strong Java and Python skills across JVM Spark and PySpark.
  • Advanced SQL, relational data modeling, and large-scale data processing experience, including production work with Snowflake, MySQL, and Iceberg.
  • 3+ years operating AWS data workloads using S3 and Airflow (or equivalent production workflow orchestration); container orchestration and data-catalog experience a plus.
  • 3+ years building, changing, and operating production REST APIs, including PHP APIs built with Symfony or a closely equivalent PHP MVC framework.
  • Demonstrated proficiency using AI development tools to deliver production-quality work efficiently — including AI-assisted coding, code review, documentation, and investigation, working with agentic code harnesses, and spec-based (spec-driven) AI development — with the judgment to know when to rely on AI output and when not to.
  • Ability to independently deliver secure, backward-compatible API changes, including implementation, validation, authorization, data access, documentation, automated testing, and production diagnostics.
  • Experience implementing external provider or client integrations involving schemas, APIs, file delivery, identity fields, reconciliation, privacy-sensitive data, and SLAs.
  • Experience writing unit and integration tests for data pipelines, APIs, and web applications.
  • Ability to diagnose production failures, reconcile data, execute backfills, manage staged releases and rollbacks, and support delivery commitments.
  • Hands-on experience building and running big data systems, with a focus on performance, reliability, and data quality.
  • Experience in ad tech or advertising measurement, such as attribution, conversions, audience and impression data, identity matching, or partner optimization feeds.
  • Ability to ramp quickly and deliver with minimal onboarding in an existing, complex codebase.
Preferred Skills:
  • Direct experience building partner optimization or advertising feeds with external platforms (DSPs, publishers, or measurement partners).
  • Experience with deterministic and probabilistic identity matching — unique IDs, device IDs, IP-based matching, and identity graphs.
  • Familiarity with edge/log delivery infrastructure and pixel/impression tracking.
  • Experience with data lake table formats and query engines at large scale.
  • Experience with privacy and compliance-driven data workflows (deletion, opt-out, suppression, data retention).
  • Comfort operating in uncharted territory — turning ambiguity into production systems without a detailed guide.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

2026-2201 Data Scientist
2026-2201 Data Scientist

Mountain Cat LLC • McLean (VA)

On-site
USD 120,000 - 180,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

Skillerszone LLC • Los Angeles (CA)

On-site
USD 90,000 - 120,000
Senior Big Data Engineer
Senior Big Data Engineer

KMM Technologies, Inc. • Rockville (MD)

Hybrid
USD 140,000 - 200,000
Senior Data Engineer
Senior Data Engineer

ConsultNet Technology Services and Solutions • United States

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

The Doyle Group • Bridgeport (CT)

On-site
USD 120,000 - 170,000
Fully remote
Senior Data Engineer
Senior Data Engineer

Boston Energy Trading and Marketing LLC • Boston (MA)

On-site
USD 120,000 - 150,000
Pyspark Developer
Pyspark Developer

Tieto • Irving (TX)

On-site
USD 140,000 - 190,000
Senior Data Engineer
Senior Data Engineer

The Phoenix Group • New York (NY)

On-site
USD 120,000 - 170,000
Senior Data Engineer – 8+ Years Experience Full Time role
Senior Data Engineer – 8+ Years Experience Full Time role

hudsonmanpower • United States

On-site
USD 140,000 - 190,000