We are seeking aData Scientist Leadto help build and evolve a high-quality measurement data foundation that enables trusted analytics and decision-making at scale. This role focuses on designing and delivering resilient datasets, pipelines, and reusable metrics that support hypothesis-driven analyses and experiments across the product development lifecycle (PDLC).
You’ll be hands-on where needed, drive engineering standards, and help teams move faster by improving reliability, observability, and usability across the data lifecycle-so leaders can clearly see what's driving value, what's creating friction, and what operating-model shifts materially improve outcomes as teams become more agentic.
Job Responsibilities
- Design, build, and operate scalabledata pipelines(batch and/or streaming) with clear SLAs, monitoring, and incident response practices.
- Develop and curate trusteddata products(e.g., conformed dimensions, event models, marts) with strong documentation and clear ownership.
- Build and maintainwell-defined metrics and feature-ready datasetsthat enable measurement of AI adoption and productivity outcomes (e.g., reusable aggregates, cohorting, time-windowed measures), including change control as definitions evolve.
- Drivedata quality and governancethrough validations, reconciliations, lineage, access controls, retention, and auditability aligned to requirements.
- Develop and operate workflow orchestration (e.g.,Apache Airflow) to schedule, monitor, and manage data movement and transformations.
- Model and transform data for analytics usingSQL/dbtto support trusted reporting and repeatable measurement.
- Write production-gradePython/PySparkwith disciplined testing, performance tuning, and maintainable design.
- Partner with analytics, product, and engineering stakeholders to define requirements, success criteria, and consistent interpretation of key measures- particularly where inputs spanfinance business cases,PDLC/SDLC tools, andAI tool logs.
- Establish and enforce engineering best practices (version control, code review, testing strategy, deployment processes, runbooks) and continuously improve observability and cost/performance (freshness, completeness, timeliness, scalability, spend).
- Mentor and develop a team of2, influencing technical direction through standards, reviews, and knowledge sharing.
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
- 5+ years of hands-on experience delivering production data solutions in a fast-paced engineering environment (actively coding and owning outcomes).
- Strong software engineering fundamentals (system design, data structures, object-oriented programming, testing strategies, and end-to-end development lifecycle).
- Strong understanding of data modeling (conceptual, logical, physical), including dimensional, normalized, and event-based approaches.
- Hands-on experience withDatabricksand large-scaled distributed data processing/performance tuning (Spark/PySpark).
- StrongSQLskills and experience with modern transformation tooling (e.g.,dbt), including building maintainable, testable data codebases.
- Experience designing and operating orchestration pipelines usingAirflow(or equivalent), including backfills, retries, and operational monitoring.
- Demonstrated rigor building and maintaining trustedmetrics(definitions, edge cases, validation/testing, documentation) and keeping them reliable as upstream sources change.
- Demonstrated ability to lead delivery in complex environments with multiple stakeholders and ambiguous requirements.
Preferred Qualifications
- Experience with modern lakehouse/warehouse patterns and broader cloud data platforms (e.g., Databricks, Snowflake).
- Experience with BI/semantic layers and metrics management practices.
- Exposure to experimentation or hypothesis-driven analytics approaches (e.g., measurement design to support tests, rollouts, and pre/post evaluation); deep causal specialization not required.