Data Engineer

Wynd Labs

United States

On-site

USD 120,000 - 170,000

Full time

14 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Fully remote team
Competitive salary
Equity package

Job summary

Wynd Labs is seeking a Data Engineer to design, optimize, and operate large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery with a focus on scalability, reliability, and performance.

This is a hands-on role involving distributed systems, large datasets, and web scraping infrastructure. The team collaborates across EST business hours and values fast, clean delivery, and robust data workflows.

Qualifications

  • Bachelor’s degree or equivalent work experience.
  • Python (advanced)—strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs.
  • Web scraping at scale—hands-on experience with high-volume scraping (proxies, rate limiting, anti-bot evasion).
  • Distributed data pipelines—experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar).
  • Data warehousing—practical experience with columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses
  • Docker & Kubernetes—containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads
  • Linux & bare-metal ops—comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed-cloud abstractions
  • CI/CD for data workflows (GitHub Actions, ArgoCD)
  • Writing Scalable API

Responsibilities

  • Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.
  • Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.
  • Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements.
  • Monitor and troubleshoot data pipeline issues, identify data quality concerns, and help implement timely fixes to maintain data accuracy and operational continuity.
  • Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented.
  • Participate in research and development projects to improve the Company’s data products and workflows.

Skills

Python
Async programming
Web scraping at scale
Distributed data pipelines
Data warehousing
CI/CD for data workflows
Linux
Docker & Kubernetes
Scalable API

Education

Bachelor’s degree or equivalent

Tools

Celery
Kafka
RabbitMQ
Databend
ClickHouse
BigQuery
Docker
Kubernetes
Helm
GitHub Actions
ArgoCD

Job description

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs. We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role:

We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance. This is a hands‑on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads.

Please note: This role requires a work schedule that overlaps sufficiently with EST business hours to collaborate effectively with the team.
Who You Are:
  • Bachelor’s degree or equivalent work experience
  • Python (advanced) — strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs
  • Web scraping at scale — hands‑on experience with high-volume scraping (proxies, rate limiting, anti‑bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media)
  • Distributed data pipelines — experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar)
  • Data warehousing — practical experience with columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost‑aware querying on cloud warehouses
  • Docker & Kubernetes — containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads
  • Linux & bare‑metal ops — comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed‑cloud abstractions
  • CI/CD for data workflows (GitHub Actions, ArgoCD)
  • Writing Scalable API
What You’ll Be Doing:
  • Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.
  • Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.
  • Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements.
  • Monitor and troubleshoot data pipeline issues, identify data quality concerns, and help implement timely fixes to maintain data accuracy and operational continuity.
  • Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented.
  • Participate in research and development projects to improve the Company’s data products and workflows.
Why Work With Us:
  • Opportunity. We are at the forefront of developing a web‑scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.
  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we’re capable of and to do work that matters. We’re not calling for people who keep things moving. We're calling for people who make everyone around them better. We prioritize low ego and high output. This is a fully remote team.
  • Compensation. You’ll receive a competitive salary, benefits and equity package.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Bedrock Robotics • New York (NY)

On-site
USD 120,000 - 150,000
Data Engineer
Data Engineer

Southern Arkansas University • Warner Robins (GA)

Remote
Flexible hours
Weekly bonus of $500–$1000 USD
Work from anywhere
Data Engineer
Data Engineer

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Data Engineer Remote Latin America
Data Engineer Remote Latin America

Fractal River • United States

Hybrid
USD 70,000 - 120,000
Personal development plan
Access to a reference library
Unlimited access to AI tools
+3
Data Engineer (Founding Team)
Data Engineer (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Early-stage equity
Data Engineer
Data Engineer

OpenAI • San Francisco (CA)

On-site
USD 235,000 - 385,000
Relocation assistance
Data Engineer
Data Engineer

Appsierra Group • United States

On-site
USD 140,000 - 180,000
Equity
Performance bonuses
Health insurance reimbursement
+3
Big Data Engineer
Big Data Engineer

Clear Fracture • United States

On-site
USD 90,000 - 130,000
Senior Data Engineer
Senior Data Engineer

Gloo • San Francisco (CA)

Hybrid
USD 150,000 - 250,000
Flexible work schedules
Remote-friendly culture
Unlimited PTO
+2