You'll join an established Data & AI team working on our Centralized Data Infrastructure platform. This is a Microsoft Fabric Lakehouse that serves as the single source of truth for all analytics, operations, and AI products across the company.
This is not a maintenance role. You'll be building and modeling data pipelines that directly enable AI products, every source you integrate and every model you design unlocks new capabilities for the business. The data you shape will power an AI Candidate Matching Engine, Conversational BI, an Onboarding Agent, and a Recruiter Copilot.
We need someone who can pick up complex data challenges, own their deliverables end-to-end, and ship independently within a collaborative team.
What You'll Do
- Build and maintain data pipelines across a three-zone lakehouse architecture (Bronze → Silver → Gold) on Microsoft Fabric
- Integrate new data sources — VMS platforms (FieldGlass, Beeline, Magnit), ATS (JobDiva), payroll systems, HR data, communication platforms, and job boards
- Design star schema models that transform raw operational data into analytics-ready and AI-ready datasets. You'll reshape a 71+ table ATS schema into clean dimensional models.
- Implement data quality rules — deduplication, PII classification and masking, cross-source entity resolution, validation frameworks
- Support API integrations — RESTful API consumption, webhook handling, and fallback scraping/bot pipelines where APIs aren't available
- Collaborate with the AI Products squad to ensure data pipelines produce the features and formats AI models need
What We're Looking For
Must have (3-7 years of experience)
- Strong experience with Microsoft Fabric (or equivalent: Azure Synapse, Databricks, Snowflake. Fabric experience preferred)
- Proficiency in Python and SQL for data transformation and pipeline orchestration
- Hands-on experience with ETL/ELT pipelines at scale — batch and incremental loads
- Familiarity with API integrations — REST APIs, OAuth, webhook-based data ingestion
- Understanding of data quality frameworks — validation rules, anomaly detection, reconciliation
- Good communication in English (daily collaboration with a distributed team)
Nice to have
- Experience with Microsoft Fabric specifically (lakehouses, notebooks, data pipelines, warehouses)
- Experience with ATS or staffing industry data (JobDiva, Bullhorn, or similar)
- Familiarity with data governance — RLS, PII masking, audit trails
- Experience with VMS platforms or their APIs (FieldGlass, Beeline, Magnit)
- Knowledge of medallion architecture (Bronze/Silver/Gold pattern)
- Familiarity with Git and CI/CD for data pipelines
Why This Role
- High impact Your work directly enables reporting, automation and AI products.
- Own your projects: You won't be waiting for someone to tell you what to do. You'll take ownership of specific data domains and deliver them end-to-end — from ingestion to business/AI-ready models.
- AI-first environment: The data you build is consumed by LLM-powered products, not just dashboards. You'll work alongside AI engineers pushing the boundaries of what's possible in staffing.
- Deepen your expertise: As the platform grows in complexity and the team scales toward a Data Platform Squad (3-4 engineers), there's room to become a deep technical specialist in the areas you care about most.