About this Position
As a Data Platform Engineer — External Data Acquisition, you will build and operate ASTRO's platform for acquiring data from external sources, including websites, web/mobile applications, APIs, public datasets, and third-party data providers.
A significant part of this role focuses on production-grade web and application data acquisition for pricing and market intelligence use cases. You will own the acquisition layer—from understanding external systems and extracting data to delivering reliable landing data for downstream processing.
Focus
External Data Acquisition
- Web & App Scraping
- Python
- HTTP/API
- Data Pipelines
- Reliability
As our Astronaut, your missions are:
- Build and maintain reliable web/app crawlers, scrapers, and external API integrations across e-commerce, marketplace, retail, and other external sources.
- Build acquisition systems that remain reliable as external websites, applications, APIs, and data contracts evolve.
- Handle real-world source complexity such as sessions, tokens, pagination, dynamic content, rate limits, retries, concurrency, and partial failures.
- Build resilient and maintainable acquisition systems with appropriate testing, idempotency, deduplication, retries, safe reruns, data-quality validation, monitoring, and alerting.
- Operate acquisition workloads using appropriate workflow orchestration, compute, data storage/warehousing, and cloud infrastructure.
- Continuously improve reliability, scalability, maintainability, and total cost of ownership, including evaluating build-vs-buy approaches for external data.
- Explore new technologies and AI-assisted development or automation to improve how external-data pipelines are built, operated, and maintained.
- Collaborate with Pricing, Data Warehouse Engineering, Data Analysts, and other stakeholders to turn external-data needs into sustainable platform capabilities.
Minimum Qualifications
In order to launch successfully, you need to have:
- Strong Python/software engineering fundamentals and experience building maintainable production systems.
- Hands-on experience building or operating web scraping, crawling, external API, or similar data acquisition systems.
- Strong understanding of HTTP, APIs, JSON, headers, cookies, sessions, authentication, pagination, and request/response behavior.
- Hands-on experience with browser/network debugging and the ability to investigate how websites or applications communicate with their backend services.
- Experience with both HTTP-based extraction and browser automation, using tools such as Requests/HTTPX, Selenium, Playwright, Scrapy, or equivalent.
- Understanding of production reliability patterns such as retry, timeout, rate limiting, concurrency, idempotency, and failure recovery.
- Solid understanding of SQL, data storage/warehousing, schema evolution, workflow orchestration, containerized workloads, and cloud infrastructure fundamentals.
- Strong troubleshooting skills, especially when working with frequently changing external systems outside your direct control.
Good to Have
- Experience acquiring data from mobile applications, marketplaces, e-commerce, grocery, food-delivery, or other large consumer platforms.
- Experience with asynchronous or distributed acquisition systems and high-volume crawling.
- Familiarity with proxy infrastructure, data observability, CI/CD, and infrastructure-as-code.
- Experience using AI-assisted development or automation to improve data acquisition workflows.
- Experience designing reusable scraping frameworks or shared acquisition platforms.
- Experience with geospatial or other non-traditional external datasets.