Job Summary: Our client, a major entertainment and streaming company, is looking for a Senior Data Engineer to support their Product Performance team within the Data Organization on a 6‑month contract, onsite in New York, NY. This team owns the data that measures the streaming product's performance, browse, playback attribution, and search, that leadership relies on to make business decisions about the platform. The role involves designing, building, and deploying batch and streaming data pipelines using Scala and Python, primarily through Databricks and Airflow. The right candidate has strong production experience with Spark and Kafka, hands‑on Airflow orchestration experience, and strong SQL skills, along with the ability to troubleshoot pipeline failures and drive root‑cause analysis. This role is onsite in New York, NY. This position is W‑2 only. We do not work with third‑party firms or C2C arrangements for this role.
Core Responsibilities
- Design, implement, and deploy data solutions for batch ingestion and stream processing using Scala and Python.
- Establish and enforce standards for code development, testing, and deployment to ensure data quality and governance.
- Extend existing data models and build new data pipelines to introduce new metrics and dimensions.
- Maintain and update existing source code powering core datasets.
- Collaborate with Product Managers, Data Analysts, and cross‑functional engineering teams to define requirements and deliver new data features.
- Work within a modern data stack including Airflow, Spark, Databricks, Delta Lake, Snowflake, GitHub Actions, and Jenkins.
- Investigate and resolve complex data pipeline failures and quality issues, performing root‑cause analysis and documentation.
- Identify and implement technical optimizations to improve performance, scalability, and cost‑efficiency.
Required Skills and Experience (Must‑Haves)
- 5+ years of software engineering experience developing large data pipelines.
- Strong Scala and Python software engineering skills.
- Hands‑on production experience with distributed systems such as Spark and Kafka.
- Hands‑on production experience with orchestration systems such as Airflow.
- Strong SQL skills, able to create queries to analyze complex datasets.
- Strong CI/CD and SDLC fundamentals.
- Algorithmic problem‑solving expertise.
- Basic understanding of AWS or another cloud provider (S3).
- Experience with scripting languages (Bash, PowerShell).
- Local to New York, NY and able to commute onsite.
Preferred Skills and Experience (Nice‑to‑Haves)
- Master's degree in Computer Science or Information Systems.
- Experience with an MPP/cloud database (Snowflake, Redshift, BigQuery).
- Experience with Jenkins and GitHub Actions.
- Experience developing APIs with GraphQL.
- Deep understanding of AWS or other cloud providers, plus infrastructure as code.
- Familiarity with data modeling and data warehousing methodologies.
- Familiarity with Scrum and Agile methodologies.
Key Competencies and Behaviors
- Self‑starting problem solver with strong analytical and communication skills.
- Comfortable owning production issues end‑to‑end, from root‑cause analysis through resolution.
- Able to balance new pipeline development with maintaining existing critical systems.
- Willing and able to learn new tools quickly in a fast‑evolving stack.
- Provides technical guidance and mentorship on best practices.
- Detail‑oriented with a strong focus on data quality and governance.
- Reliable with onsite attendance.
Work Environment
Location: New York, NY
Schedule: Onsite
Duration: 6 months
Compensation & Benefits
- Pay Range: $75.00 – $85.00 per hour (approximate; may vary).
- Medical, Dental, & Vision Insurance Plans.
- Employee‑Owned Profit Sharing (ESOP).
- 401K offered.