We are looking for a Senior Data Engineer to join our team. The Data Engineer is the most data-intensive engineering role on the engagement. If any pipeline drops data, positions drift, or signals compute incorrectly, calibration and UAT break. Correctness and operational robustness here are a prerequisite for everything else.ResponsibilitiesExecute all database migration sets spanning the Unified Data Store to ensure schema consistency across the platformBuild and maintain the full suite of Source Adapters responsible for connecting external systems into the data platformImplement the Field Mapper, establishing per-project and per-franchise field bindings that normalize source system fields into the unified schema during ingestionImplement the Link Resolver, responsible for resolving CL-to-ticket, ticket-to-ticket, and case-to-defect/requirement links across all source families, feeding directly into the Attribution ResolverImplement the Attribution Resolver, tracing the change-to-ticket-to-area-to-test chain, alongside the Counted Signals Aggregator, which builds file-to-area and area-to-area counted association tables with path normalization and counting verificationBuild the Area Vocabulary and accompanying Translation Tables to support consistent area classification across the systemDevelop all Signal Catalogue computation jobs covering the eight core signals — area fragility, recency, recent failures, change-touch, coupling, windowed area change volume, testing alignment, validation recency, and defect impact/volume/age — ensuring provenance capture throughoutTake ownership of data dictionary authoring across all storage components, documenting incrementally as new features are deliveredRequirements3+ years of hands-on relevant experience in Python software engineeringBackground in AI Data Engineering, applying data engineering practices to support AI/ML-driven systemsPractical experience with Apache Airflow for orchestrating and scheduling data workflowsWorking knowledge of Machine Learning concepts and their application within data systemsProven experience in data pipeline development, from design through implementationHands-on experience with PostgreSQL for data storage and queryingExperience designing idempotent ingest processes, including natural keys, upsert auditing, duplicate detection, and replay safetyExperience integrating multiple heterogeneous data sources into a unified systemSkilled in data quality and telemetry practices, including fill-rate counters, volume metrics, and reconciliation reportingFamiliarity with Spec Driven Development methodologyDecent communication skills with working English fluency (B2 level or higher) to understand business requirements and translate them into agentic architecturesNice to haveExperience developing API clients for Perforce or Code HubExperience integrating with the JaaS (Jira) REST APIFamiliarity with test management APIs such as QMetry, Zephyr, or similar toolsExperience building Snowflake connectors, including key-pair authentication and warehouse extract queriesKnowledge of temporal data systems and as-of read patternsWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn