Data Engineer

Socket.dev

El Segundo (CA)

On-site

USD 120,000 - 130,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Socket.dev in El Segundo, CA is seeking a Data Engineer for Basketball Data Strategy to build, maintain, and operate cloud data pipelines spanning ingestion, transformation, storage, and serving layers for web and mobile products.

You will work with senior engineers to ensure data quality, implement bronze/silver/gold datasets, run backfills, and support in-game data delivery while growing toward platform ownership.

Qualifications

  • 2 to 4 years of experience in a data engineering or a similar role.
  • Strong SQL and relational DB design; familiarity with nonrelational stores is a plus.
  • Experience building cloud data pipelines in AWS; orchestration with Airflow or Prefect.
  • Proficiency in Python and Spark (PySpark) for data transformation.
  • Exposure to columnar formats (Parquet, Iceberg) and near real‑time data processing.

Responsibilities

  • Build and maintain bronze, silver, and gold data pipelines from multiple sources.
  • Deliver the serving layer and data marts in PostgreSQL for web and mobile apps.
  • Run idempotent backfills and routine reprocessing to ensure data freshness.
  • Support live and in-game data flows with near‑real‑time processing patterns.
  • Maintain lakehouse with partitioning, compaction, and file management for performance.
  • Implement data quality controls and automated validation across pipelines.
  • Contribute to the data catalog and lineage documentation.
  • Onboard new reference and entity data sources and assist data ingestion.

Skills

SQL
Python
PySpark
Airflow
PostgreSQL
Data Pipelines

Education

Bachelor's degree in Statistics/CS/Engineering or equivalent

Tools

PostgreSQL
Airflow
Prefect
Spark
Docker

Job description

Title:Data Engineer, Basketball Data Strategy
Department:Basketball Data Strategy, Basketball Operations
Reports to:Senior Data Engineer, Data Systems
Position Summary:

The Data Engineer, Basketball Data Systems is a core contributor on the Basketball Data Systems team. Working under the direction of the team's senior engineers and data-systems leadership, they build, maintain, and operate the pipelines and serving datasets of a modern cloud data platform spanning ingestion, transformation, storage, and service layers across both batch and near‑real‑time (in-game) delivery, using statistical feeds from external vendors and internal sources. This position focuses on the hands‑on implementation and day-to-day reliability of that platform: building curated bronze, silver, and gold datasets to establish patterns and data contracts, materializing them into the serving layer for web and mobile products, running backfills, and keeping pipelines performant and well maintained. It works alongside the Data Science / Analyst teams to ensure the cleanliness, integrity, accuracy, and relevance of the data to be analyzed, and alongside the Software Development team to deliver data into performant products. It is an excellent opportunity for an engineer who wants to grow toward broader platform ownership while working in a greenfield cloud stack under experienced mentorship.

Essential Functions (Duties & Responsibilities**):
  • Build and maintain data pipelines: implement and operate automated bronze, silver, and gold pipelines that collect, clean, and transform data (structured and unstructured) from internal and third‑party sources, following the platform's established medallion patterns, schemas, and data contracts.
  • Deliver the serving layer: build and maintain the gold data marts and the load that materializes them into the serving database (PostgreSQL), along with the serving views and tables that the web and mobile products read, to the shapes agreed with the Software Development team.
  • Run backfills and reprocessing: execute idempotent historical backfills and routine reprocessing, validating outputs across seasons.
  • Support live and in-game data flows: help build, operate, and monitor the incremental, near‑real‑time pipelines that deliver in-game data, applying the platform's idempotency and quality patterns under the event‑driven design the senior engineers set.
  • Maintain and tune the lakehouse: perform table maintenance such as compaction and small‑file management, partitioning, and storage‑format upkeep to keep queries fast and costs controlled.
  • Apply data‑quality controls: roll out and maintain the team's validation, quarantine, run‑ledger, and alerting seam across tables, including automated quality checks and unit testing of transformation logic.
  • Contribute to the data catalog and lineage: help generate and maintain the machine‑readable table catalog (grain, columns, sensitivity, and dependencies) that documents the ecosystem and drives orchestration.
  • Support new‑source onboarding: when scoped, land and shape new reference and entity sources, and assist with the ingestion and processing of player‑tracking and other high‑precision movement data.
  • Partner with consumers: work alongside the Data Science / Analyst and Software Development teams as the primary users of these datasets, incorporating their feedback on usability and correctness.

**Duties & Responsibilities subject to change based on organizational needs and direction from management.

Education:
  • Bachelor's in Statistics, Computer Science, Engineering, or a related field, or equivalent academic or professional experience; experience working with database solutions in a professional, best‑practices environment.
Minimum Qualifications:
  • 2 to 4 years of experience in a data engineering or similar role.
  • Database skills: strong proficiency in SQL and relational database design (e.g. PostgreSQL, SQL Server), including a working knowledge of best practices such as normalization, indexing, and query optimization. Exposure to nonrelational stores (document, time‑series, or vector) is a plus.
  • Cloud pipelines: experience building and maintaining data pipelines in a cloud environment; AWS preferred.
  • Orchestration: experience with a workflow orchestration tool (e.g. Prefect, Airflow).
  • Strong Python and SQL skills, with experience transforming data in Spark (PySpark) or a similar framework.
  • Familiarity with columnar and lakehouse formats (e.g. Parquet, Iceberg) and working with large datasets; exposure to high‑precision location or movement data (sub‑second sampling) is a plus.
  • Exposure to streaming or near‑real‑time data processing (e.g. microbatch, structured streaming, Kafka or Kinesis) is a plus.
  • Comfort with version control and CI/CD; exposure to infrastructure‑as‑code and containers (e.g. Docker) is a plus.
  • An interest in data quality and observability, and in treating data infrastructure as production software.
Other Qualifications:
  • A strong sense of organization
  • High agency, with an eagerness to learn and grow under senior mentorship
  • Adaptability and "outside-the-box" thinking
  • Knowledge of and passion for NBA basketball
Location:El Segundo (office M-F), and other occasional off‑site events
Travel:Less than 5% of the time
Hours:Full-time. Must be available to work evenings, weekends and holidays as reasonably required

The pay range for this role is $120,000 - $130,000 annually. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience and certifications. In addition to those factors, we consider the relative pay of our current employees in similar positions when making a final offer.

We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, sexual orientation, age, disability, gender identity, marital or veteran status, or any other protected class.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

The Los Angeles Lakers • El Segundo (CA)

On-site
USD 120,000 - 130,000
Data Engineer
Data Engineer

Los Angeles Lakers • El Segundo (CA)

On-site
USD 120,000 - 130,000
Basketball Data Scientist
Basketball Data Scientist

Phoenix Suns • Phoenix (AZ)

On-site
USD 120,000 - 180,000
Principal Data Engineer
Principal Data Engineer

Jobtailor • Massachusetts

Hybrid
USD 140,000 - 190,000
Full Stack Data Engineer
Full Stack Data Engineer

Cleveland Cavaliers • Cleveland (OH)

On-site
USD 110,000 - 165,000
Healthcare
Dental
Vision
+1
Data Engineer (Wizards)
Data Engineer (Wizards)

Washington Wizards • Washington

On-site
USD 130,000 - 150,000
Health benefits
Data Engineer, Basketball Data Systems — In-Game & Cloud
Data Engineer, Basketball Data Systems — In-Game & Cloud

Socket.dev • El Segundo (CA)

On-site
USD 120,000 - 130,000
Full Stack Data Engineer
Full Stack Data Engineer

Rocket Mortgage Fieldhouse • Cleveland (OH)

On-site
USD 110,000 - 160,000
Healthcare
Dental
Vision
+1
Basketball Insights Analyst, Phoenix Suns
Basketball Insights Analyst, Phoenix Suns

Phoenix Suns • Phoenix (AZ)

On-site
USD 85,000 - 110,000
Data Operations Engineer
Data Operations Engineer

Orlando Magic • Orlando (FL)

Hybrid
USD 80,000 - 110,000
18 days of personal time off
Casual work attire on non-game days
401(k) with company matching
+1