Senior Data Engineer

Sigma Software

Warszawa

On-site

PLN 180,000 - 320,000

Full time

21 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Sigma Software in Warszawa, Poland, is seeking a Senior Data Engineer to design and optimize large-scale data infrastructure for a real-time AdTech platform. You will build production-grade ML-oriented data systems, implement robust ingestion pipelines, and collaborate with customer engineers to ensure data quality, reliability, and scalability across cloud-native environments.

Key focus areas include writing diagnostic SQL queries, building BigQuery data pipelines for bid/win/impression logs,

Qualifications

  • 5+ years of experience in Data Engineering.
  • At least 2 years of experience with production ML or large-scale analytics pipelines.
  • Expert-level SQL skills including window functions and incremental processing patterns.
  • Strong Python skills for production-grade pipeline development.
  • Hands-on experience with Spark or PySpark.
  • Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools.
  • Experience working with cloud data warehouses at scale, preferably BigQuery.
  • Strong understanding of data modeling and point-in-time correctness.
  • Experience working with event-driven or clickstream datasets at very large scale.
  • Experience supporting business-critical production pipelines.
  • Upper-Intermediate English level or higher.

Responsibilities

  • Write and defend diagnostic SQL queries against large-scale production datasets.
  • Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery.
  • Harmonize fields across independently designed datasets and maintain versioned field mappings.
  • Develop point-in-time-correct feature tables and aggregation pipelines.
  • Design and maintain conversion and labeling pipelines with delayed label handling.
  • Own the data serving write path, schema contracts, publishing flows, and freshness SLOs.
  • Build experimentation infrastructure including traffic splitting and reporting pipelines.
  • Perform large-scale historical backfills and safe reprocessing after mapping changes.
  • Implement data isolation and safe-aggregation controls for advertiser data protection.
  • Develop automated data quality validation frameworks.
  • Collaborate closely with Customer engineers and prepare operational documentation.
  • Contribute to architecture discussions and platform scalability improvements.

Skills

Data Engineering
ML pipelines
SQL
Python
Spark
ETL/ELT
BigQuery
Data modeling
English

Tools

Airflow
Cloud Composer
Dagster
BigQuery
PySpark

Job description

Company Description
Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. We are looking for a Senior Data Engineer who enjoys solving complex distributed data challenges and building production-grade ML-oriented data systems.

Company Description
Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. We are looking for a Senior Data Engineer who enjoys solving complex distributed data challenges and building production-grade ML-oriented data systems.
You will become part of a dedicated Sigma Software team developing predictive modeling and optimization capabilities for a live advertising ecosystem. The role combines large-scale event processing, streaming and batch pipelines, experimentation infrastructure, and high-throughput data engineering in a cloud-native environment.
We as a company offer the opportunity to work on impactful global products, collaborate with experienced engineers, and contribute to architecture decisions while growing your expertise in large-scale distributed systems and modern data platforms.
CUSTOMER
Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a large-scale ad exchange handling hundreds of millions of auction requests per day and is actively investing in predictive decisioning technologies to optimize advertising outcomes in real time.
PROJECT
The project focuses on building a predictive modeling and optimization platform on top of a live ad exchange environment. The platform performs real-time supply scoring and filtering, contextual performance estimation, look-alike audience generation, and multi-objective optimization under business constraints.
The solution processes massive-scale event and auction datasets and includes feature engineering pipelines, streaming and batch ingestion, experimentation infrastructure, point-in-time-correct training data generation, and ML-oriented data services with strict operational reliability and compliance requirements.
Job Description

  • Write and defend diagnostic SQL queries against large-scale production datasets
  • Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
  • Harmonize fields across independently designed datasets and maintain versioned field mappings
  • Develop point-in-time-correct feature tables and aggregation pipelines
  • Design and maintain conversion and labeling pipelines with delayed label handling
  • Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
  • Build experimentation infrastructure including traffic splitting and reporting pipelines
  • Perform large-scale historical backfills and safe reprocessing after mapping changes
  • Implement data isolation and safe-aggregation controls for advertiser data protection
  • Develop automated data quality validation frameworks
  • Collaborate closely with Customer engineers and prepare operational documentation
  • Contribute to architecture discussions and platform scalability improvements
Qualifications
  • 5+ years of experience in Data Engineering
  • At least 2 years of experience working with production ML or large-scale analytics pipelines
  • Expert-level SQL skills including window functions and incremental processing patterns
  • Strong Python skills for production-grade pipeline development
  • Hands-on experience with Spark or PySpark
  • Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
  • Experience working with cloud data warehouses at scale, preferably BigQuery
  • Strong understanding of data modeling and point-in-time correctness
  • Experience working with event-driven or clickstream datasets at very large scale
  • Experience supporting business-critical production pipelines
  • Upper-Intermediate English level or higher
WILL BE A PLUS
  • Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
  • Experience building streaming or near-real-time ingestion systems
  • Understanding of feature stores, train/serve skew, and label leakage prevention
  • Experience in AdTech or auction-based environments
  • Experience handling delayed or incomplete labels in ML systems
  • Experience with dbt or similar transformation frameworks
  • Experience delivering solutions into Customer-owned infrastructure
  • Knowledge of GDPR/CCPA-related privacy engineering practices
  • Experience with experimentation infrastructure and statistical validation pipelines
  • Experience working in hybrid cloud/on-prem Linux environments
  • Terraform and Kubernetes experience
  • Experience optimizing warehouse cost and performance
Additional Information PERSONAL PROFILE
  • Strong analytical and problem-solving skills
  • Ownership-oriented mindset
  • Ability to work independently in a client-facing environment
  • Strong communication and documentation skills
  • Comfortable working in a fast-paced engineering environment
  • Collaborative and proactive attitude
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Scientist
Senior Data Scientist

Sigma Software • Warszawa

On-site
PLN 260,000 - 480,000
Senior Big Data Engineer (AdTech Cybersecurity)
Senior Big Data Engineer (AdTech Cybersecurity)

Sigma Software • Warszawa

On-site
PLN 180,000 - 280,000
Senior MLOps / ML Platform Engineer
Senior MLOps / ML Platform Engineer

Sigma Software • Warszawa

On-site
PLN 180,000 - 320,000
Senior Data Engineer (Databricks Migration)
Senior Data Engineer (Databricks Migration)

Sigma Software • Kraków

On-site
PLN 240,000 - 360,000
Senior Analytics Engineer (Semantic Layer)
Senior Analytics Engineer (Semantic Layer)

Sigma Software • Warszawa, Kraków, Poznań, Wrocław

On-site
PLN 180,000 - 270,000
Senior Data Engineer (Databricks Migration)
Senior Data Engineer (Databricks Migration)

Sigma Software • Województwo małopolskie

On-site
PLN 180,000 - 240,000
Senior Data Engineer (Databricks Migration)
Senior Data Engineer (Databricks Migration)

Sigma Software • Województwo mazowieckie

Hybrid
PLN 240,000 - 360,000
Principal Data Platform Engineer (Swedish Ad Platform)
Principal Data Platform Engineer (Swedish Ad Platform)

Sigma Software • Kraków

On-site
PLN 380,000 - 580,000
Principal Data Platform Engineer (Swedish Ad Platform)
Principal Data Platform Engineer (Swedish Ad Platform)

Sigma Software • Wrocław

On-site
PLN 260,000 - 380,000
Senior/Principal Golang Developer
Senior/Principal Golang Developer

Sigma Software • Warszawa

On-site
PLN 300,000 - 420,000