Locations: Wakefield, Massachusetts; Denver, Colorado; New York, New York; Atlanta, Georgia; Cary, North Carolina; Dallas, Texas; Pasadena, California; Rancho Cordova, California; West Palm Beach, Florida. Salary: $90,000.00 – $150,000.00 annually.
Responsibilities
- Design, build, and maintain data pipelines that ingest, transform, and deliver enterprise data to AI solutions and business applications.
- Develop and manage integrations across enterprise data sources using APIs, Graph connectors, event-driven architectures, and batch/streaming patterns.
- Build and maintain data stores and indexing infrastructure that support retrieval‑augmented generation (RAG) and other AI consumption patterns.
- Implement data quality, validation, and lineage controls to ensure accuracy and trustworthiness of data feeding AI workflows.
- Support and optimize data models underpinning Power BI dashboards and AI‑enabled analytics.
- Collaborate with AI Engineers to define data contracts and ensure pipeline outputs meet solution requirements for schema, latency, and freshness.
- Instrument data pipelines for monitoring, alerting, cost control, and performance optimization.
- Implement data governance and security controls including access management, encryption, and compliance with organizational data policies.
- Identify data‑related risks and support mitigation strategies in collaboration with architecture and platform teams.
Minimum Qualifications
- 4+ years of data engineering experience, with demonstrated ability to build and operate production data pipelines.
- Proficiency in Python and SQL; experience with PySpark or Spark is strongly preferred.
- Experience with Azure data services, including Azure Data Factory, Azure Synapse Analytics, Azure Data Lake Storage, and Azure SQL.
- Familiarity with Azure AI Search, Cosmos DB, or similar services used to support AI retrieval and storage patterns.
- Experience building and managing ETL/ELT pipelines with structured, semi‑structured, and unstructured data sources.
- Knowledge of data modeling, schema design, and indexing strategies for both analytical and AI workloads.
- Familiarity with infrastructure‑as‑code and CI/CD practices for data pipeline deployment (e.g., Terraform, Azure DevOps).
- Knowledge of data governance principles, including data cataloging, lineage, access control, and privacy requirements.
- Knowledge of security best practices for data solutions, including encryption at rest and in transit, role‑based access control, and private networking.
- Experience with Databricks, including Delta Lake and Unity Catalog, is a plus.
- Prior experience in professional services, engineering, or construction environments is a plus.
- Azure or Databricks certifications (e.g., Azure Data Engineer Associate, Azure Solutions Architect Expert, Databricks Data Engineer Professional) are a plus.
Benefits
- Market‑Competitive Compensation, including eligibility for an annual performance bonus.
- Comprehensive benefits program, including medical, dental, vision, life, and disability coverage.
- Well‑being program and paid parental leave.
- Hybrid work schedules and cell phone stipends.
- GEI University with continuing education assistance and tuition reimbursement.
- Professional development programs and opportunities for advancement.
- Support and financial rewards for publication awards, professional dues, and licenses.
- Paid holidays and generous paid time‑off program.
- Rewards and recognition initiatives.
- GEI‑funded profit sharing and 401(k).
- Opportunity to become an owner and shareholder.
GEI is an AA/equal opportunity employer, including disabled and veterans.