Select how often (in days) to receive an alert:
We are seeking a highly skilled and experienced Senior Data Engineerto join our growing data team. In this role, you will be instrumental indesigning, building, and maintaining robust and scalable data pipelinesand infrastructure that power critical data-driven initiatives across ourorganization. You will work with vast datasets, cutting-edgetechnologies, and collaborate closely with AI researchers, other dataengineers, data scientists, machine learning engineers, and productteams to deliver insights that shape the future of our products and userexperiences.
What You'll Do:
Build and Optimize Data Infrastructure:
- Develop, construct, test, and maintain large-scale data ingest architecture consisting of diverse cloud-based services (messaging,storage, Kubernetes, persistent data store, serverless functions, etc).
- Create tooling like SDK, APIs to enable user self-service.
- Contribute to the design and evolution of our core data platform,ensuring its scalability, reliability, and cost-effectiveness.
- Implement robust monitoring, alerting, and logging solutions for datapipelines and infrastructure to proactively identify and resolve issues.
Design and Implement Scalable Data Pipelines:
- Design and implement highly reliable and efficient ETL/ELT processes to ingest, transform, and load data from diverse sources (e.g., real-time events, third-party APIs, rich media datasets) into our data lake and data warehouses.
- Utilize distributed data processing frameworks like Spark or similar to handle large scale data volumes with high throughput and low latency.
Ensure Data Quality and Governance:
- Describe and annotate datasets using industry standard schemasand internal specifications
- Cultivate data catalogs and metadata management solutions to i.o. improve data discoverability and understanding across theorganization.
- Implement data validation, cleansing, and reconciliation processes toensure the accuracy and integrity of our data assets.
Collaborate and Mentor:
- Work closely with stakeholders (research, engineering, product, andpeers more broadly) to translate their data needs into robust datasolutions.
- Provide technical leadership and mentorship to junior data engineers,fostering a culture of technical excellence and continuous learning.
- Contribute to the evolution of our data architecture and engineering bestpractices.
What You’ll Bring:
- Extensive Experience: 5+ years of experience in data engineering,with a strong focus on building and maintaining large-scale datapipelines and infrastructure.
- Programming Proficiency: Expert-level proficiency in at least onemajor programming language such as Python, Scala, or Java.- Distributed Data Processing: Deep experience with distributed dataprocessing frameworks (e.g., Apache Spark, Apache Beam). Strongfoundation in event-based approaches and systems includingmessaging/topics, pub/sub, queues, etc.
- Data Warehousing/Lakes: Hands-on experience with datawarehousing solutions (e.g., Databricks, Snowflake, Redshift,BigQuery) and data lake technologies (e.g., S3, HDFS). Deepexperience with managing large scale, heterogeneous datasets onDatabricks is highly preferred.
- SQL Mastery: Advanced SQL skills for data manipulation, analysis,and optimization.
- Cloud Platforms: Strong experience with one or more major cloudproviders (AWS, GCP, Azure) and their data-related services.
- Orchestration and DevOps: Familiarity with containerization andorchestration technologies (e.g., Docker, Kubernetes). Proficient atCI/CD-based deployment.
- Database Knowledge: Solid understanding of relational and NoSQLdatabases.
- Data Modeling: Expertise in data modeling, schema design, and dataarchitecture principles.
- Problem-Solving: Excellent analytical and problem-solving skills, witha track record of tackling complex data challenges.
- Communication: Strong communication and interpersonal skills, withthe ability to collaborate effectively with cross-functional teams.
- Master’s degree in Computer Science, Data Science, or a relatedquantitative field.
Bonus Points:
- Strong theoretical understanding of distributed computing conceptssuch as concurrency, parallelism, queueing, consistency,coordination protocols, etc.-Experience with machine learning pipelines and MLOps principles.
- Contributions to open-source data projects.