About Our Client:Our client is a frontier AI company building a next-generation AI platform - a generative AI + simulation-powered search engine.
The Role:Lead end-to-end delivery of data engineering initiatives. Architect and scale the core data infrastructure that powers their business - from data lakes and enterprise data platforms to AI-enabled analytics products and agentic systems. High-impact opportunity to build foundational systems that drive decision-making across research, product, and commercial teams.
Key Responsibilities
Data Infrastructure & Architecture
- Design, build, and scale data pipelines and lakehouse architectures supporting enterprise, product, and commercial analytics at scale
- Own the data lake ecosystem, defining standards for ingestion, storage, transformation, and access across structured and unstructured data
- Evolve the data stack for scalability, performance, and developer experience, optimizing for multi-cloud compute and supercomputing environments
- Build and maintain centralized feature registry / feature store as single source of truth for feature cataloging, lineage, ownership, SLAs - ensuring training/serving consistency
Data Products & Platforms
- Develop and own core data products including enterprise data platform, intelligence layer, and AI-powered analytics tools (including AI agents) for non-technical users
- Build robust data models supporting analytics, reporting, and ML across multiple business lines
- Enable Applied AI/ML Engineers (Agents) building agents that automate workflows
Governance & Data Quality
- Champion data quality, governance, observability to ensure organization-wide trust in data
- Implement lineage and auditability for training data used in generative models
Candidate Profile
- 12-15+ years designing and building data products - enterprise data platforms, analytics platforms, personalization systems in AI-native environments
- Hands-on lakehouse ecosystems, low-latency large-scale batch and streaming pipelines for ML optimized for GPU compute
- Feature stores, Spark/Ray/Dask, Databricks/Snowflake/Delta Lake/Iceberg
- Deep SQL, Spark, Python. Governance & observability tooling
- Translates complex technical concepts for product and commercial teams
- Scrappy startup experience - as technical as possible, as commercial as possible