Senior Principal Data Engineer – Real-time Data

Keka Technologies Private Limited

Hinoba-an

On-site

PHP 150,000 - 230,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Algoworks is seeking a hands-on Lead Data Engineer to provide technical leadership for high-volume data ingestion and processing, focusing on real-time CDC, Databricks, SQL Server, Debezium and Azure Event Hubs. The role owns architecture across batch, full-load, and real-time pipelines and requires strong leadership with hands-on execution.

The ideal candidate will guide engineers, design production-grade ingestion solutions for scalability, resiliency, and data integrity, and drive data

Qualifications

  • Bachelor’s or Master’s degree in computer science, IT, business or related field or equivalent practical experience.
  • 10+ years of Data Engineering / Software Engineering experience.
  • Strong hands-on with Databricks and Delta Lake; scale-out pipelines.

Responsibilities

  • Lead architecture and implementation of batch, full-load, incremental and real-time CDC pipelines.
  • Design high-volume ingestion from SQL Server using CDC and Debezium.
  • Build scalable event-driven pipelines using Azure Event Hubs and Databricks.
  • Design and optimize Databricks pipelines for large-scale data ingestion and transformation.
  • Implement robust error handling, retry, replay, checkpointing, recovery and idempotency.
  • Handle schema drift and evolution without disrupting downstream processing.
  • Establish monitoring and observability for CDC lag, pipeline health and throughput.
  • Provide technical direction, design reviews, and mentorship to engineers.

Skills

Data engineering
Mentoring
Technical leadership

Education

Bachelor’s or Master’s in CS/IT/Business

Tools

Databricks
Delta Lake
SQL Server CDC
Debezium
Azure Event Hubs
Python/PySpark
Azure Data Lake
Kafka (optional)

Job description

10+ Years

Full-Time

Algoworks is an award-winning artificial intelligence, engineering services and experience transformation firm with offices across the United States, Europe, South America and India. We bring together a global team of engineers, architects, designers, researchers and operators united by rigor, accountability and a commitment to delivering measurable results.

For over 20 years, Algoworks has partnered with Fortune 500 organizations across the Americas, Europe and Asia to define, build and run technology that drives meaningful business outcomes. Our work combines human-centered design, engineering excellence and AI-powered capabilities to solve complex challenges with clarity and precision. Innovation, particularly in the responsible application of AI, is embedded in how teams approach problem-solving and continuous improvement.

At Algoworks, growth is continuous and closely tied to impact. Teams collaborate across geographies and disciplines, strengthening outcomes through shared insight and collective expertise. The culture values transparency, open dialogue and an environment where every voice is heard and contribution is recognized.

Through collaboration, accountability and a focus on results, Algoworks operates at the intersection of technology and people, building not only advanced systems but strong global teams that elevate performance and create lasting impact.

Follow the video below to know about us! Clipchamp

Role overview

We are looking for a hands-on Lead Data Engineer to provide technical leadership for high-volume data ingestion and processing, with a strong focus on real-time CDC, Databricks, SQL Server, Debezium and Azure Event Hubs.

This role will own the technical architecture and engineering direction across Full Load / Batch and Real-Time CDC pipelines, with real-time streaming expected to become the primary long-term ingestion pattern.

The ideal candidate is a strong technical leader and architect who remains hands-on, can guide and mentor engineers and can design production-grade ingestion solutions for scalability, resiliency, performance and data integrity.

Key responsibilities:
  • Lead the architecture and technical implementation of batch, full-load, incremental and real-time CDC pipelines.
  • Design high-volume ingestion from SQL Server using CDC and Debezium.
  • Build scalable event-driven pipelines using Azure Event Hubs and Databricks.
  • Design and optimize Databricks pipelines for large-scale data ingestion and transformation.
  • Implement robust error handling, retry, replay, checkpointing, recovery and idempotency.
  • Design solutions for schema drift and schema evolution without disrupting downstream processing.
  • Design and optimize Delta Lake / Delta Tables, including partitioning, compaction, data layout and performance optimization.
  • Optimize pipeline throughput, latency, parallelism, resource utilization and processing windows.
  • Establish monitoring and observability for CDC lag, connector health, consumer lag, pipeline failures, throughput and processing latency.
  • Implement reconciliation and data-quality controls to ensure source-to-target completeness and accuracy.
  • Provide technical direction, perform design/code reviews, mentor engineers and establish engineering best practices.
  • Drive technical readiness for scaling ingestion across significantly more clients, databases, tables and data volumes.
Required skills and qualifications:
  • Bachelor’s or master's degree in computer science, Information Technology, Business, or related field (or equivalent practical experience).
  • 10+ years of Data Engineering / Software Engineering experience.
  • Strong hands-on experience with Databricks and Delta Lake.
  • Strong experience designing and operating Databricks data pipelines at scale.
  • Deep understanding of:
  • Pipeline design and orchestration
  • Error handling and recovery
  • Schema drift
  • Data partitioning and optimization
  • Performance tuning
  • Strong hands-on experience with SQL Server CDC, transaction logs, LSNs and high-volume transactional databases.
  • Experience with Debezium SQL Server Connector, including configuration, offsets, snapshots, recovery and schema changes.
  • Strong experience with Azure Event Hubs, including partitioning, consumer groups, scaling, throughput and checkpointing.
  • Deep understanding of batch, micro-batch, streaming and event-driven data architectures.
  • Strong experience with Python/PySpark, SQL, Azure Data Lake and distributed data processing.
  • Experience designing production-grade solutions for retry, replay, fault tolerance, duplicate handling, reconciliation and observability.
  • Strong performance engineering and troubleshooting skills across large-scale data pipelines.
  • Ability to provide technical leadership, architecture guidance, mentoring and hands-on engineering support.
Must have skills:
  • 10+ years of Data Engineering / Software Engineering experience.
  • 3+ years working with production-scale CDC or real-time streaming architectures.
  • Strong production experience with Databricks and Delta Lake.
  • Experience processing millions to billions of records.
  • Experience with multi-client or multi-tenant ingestion architectures.
  • Experience implementing Medallion / Bronze-Silver-Gold architectures.
Good to have skills:
  • Experience with Apache Kafka / Kafka Connect and streaming ecosystems.
  • Knowledge of Azure Data Factory, Azure Functions and Azure Monitor.
  • Experience with Infrastructure as Code (Terraform/ARM/Bicep) and CI/CD for data platforms.
  • Familiarity with Unity Catalog, Databricks Workflows and advanced Spark optimization.
Key success criteria:

The person in this role should be able to:

  • Establish a scalable architecture for batch and real-time ingestion.
  • Scale pipelines across substantially more databases, clients, tables and data volumes.
  • Improve Databricks pipeline performance and processing windows.
  • Deliver reliable high-volume CDC without sustained lag, duplication, or data loss.
  • Handle schema changes and schema drift without destabilizing ingestion.
  • Ensure pipelines are idempotent and safely recoverable/replayable following failures.
  • Optimize Delta Tables and downstream processing for performance and scalability.
  • Provide clear technical leadership and mentoring for the ingestion engineering team.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer (Databricks, AWS)
Data Engineer (Databricks, AWS)

blaseek • Quezon City

On-site
PHP 900,000 - 1,500,000
Data Engineer
Data Engineer

Goodfit • Hinoba-an

On-site
PHP 900,000 - 1,200,000
Databricks Unified Data Analytics Platform Engineer
Databricks Unified Data Analytics Platform Engineer

Accenture Inc. • Metro Manila

On-site
PHP 700,000 - 1,200,000
Data Engineer
Data Engineer

SCALABLE OS CORP. • Metro Manila

Remote
PHP 900,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

WeSupport Incorporated • Taguig

Hybrid
PHP 2,232,000 - 3,906,000
Senior Databricks Data Engineer (Data Migration Project)
Senior Databricks Data Engineer (Data Migration Project)

ERNI • Mandaluyong

Hybrid
PHP 1,800,000 - 3,200,000
Private HMO & insurance from day one
13th-month pay
Training & certifications
+1
Senior Data Engineer
Senior Data Engineer

Permhunt • Manila

On-site
PHP 1,674,000 - 2,567,000
Senior Data Engineer with Databricks
Senior Data Engineer with Databricks

Indra Philippines, Inc. • Pasig

On-site
PHP 2,000,000 - 3,200,000
Senior Data Architect
Senior Data Architect

TymblHub • Hinoba-an

On-site
PHP 1,200,000 - 1,800,000
Data Engineer
Data Engineer

HRTX • Philippines

On-site
PHP 1,000,000 - 1,500,000