Job Duties and Responsibilities:
- Design, build, and optimize scalable ETL/ELT pipelines that ingest, transform, and integrate structured and unstructured enterprise data for AI/ML model training and inference
- Ensure data is accurate, current, reliable, and secure through cleansing, normalization, validation, automation, continuous updates, monitoring, and error handling
- Troubleshoot complex integration and performance issues, remove bottlenecks, and apply appropriate techniques such as caching, distributed processing, and specialized ETL patterns
- Own key AI data infrastructure components and ensure they are scalable, maintainable, and supported by effective monitoring, logging, documentation, data lineage, and quality metrics
- Apply enterprise architecture, AI governance, privacy, security, coding, and CI/CD standards to data solutions and communicate risks or limitations to leadership
- Partner with Software engineers, data architects, domain experts, product teams, and platform teams to align data solutions with model requirements, business goals, and enterprise standards
- Independently deliver assigned work, influence cross-team decisions, promote data engineering best practices, and evaluate technologies that improve AI data delivery
- The requirements herein describe the general nature and level of work performed by the employee but are not a complete list of responsibilities, duties, and skills required. Other duties may be assigned as needed
Requirements
Education and Experience
- Undergraduate degree in Computer Science, Data Engineering, Information Systems, or related field
- 8+ years data engineering, software development, or related experience, including work on data pipelines for AI/ML systems
- Proficiency with data engineering languages, platforms, and orchestration tools such as Python, SQL, PySpark, Snowflake, Databricks, and Apache Airflow
- Strong knowledge of relational and NoSQL databases, data warehouses or data lakes, integration platforms such as Azure Data Factory, and streaming technologies such as Kafka
- Experience with at least one major cloud data ecosystem, preferably Azure, and familiarity with infrastructure-as-code and automated deployment practices
- Proven ability to design and implement ETL/ELT pipelines, manage databases or data lakes, and integrate large-scale data systems
- Certifications in cloud data engineering or in data privacy/security and experience with MLOps or AI/ML lifecycle management preferred
Knowledge, Skills and Abilities
- Must be willing to travel 15-25% (1-2 times per quarter)
- Deep knowledge of data engineering and AI/ML pipeline practices, including the ability to design and optimize scalable, reliable, and efficient data solutions
- Strong understanding of data quality, governance, privacy, security, and regulatory requirements
- Able to independently analyze and troubleshoot complex data and integration issues and apply creative, data-driven solutions when standard approaches are insufficient
- Able to align technical solutions with business priorities, including speed, cost, risk, reliability, and desired outcomes, and exercise sound judgment in situations with limited precedent
- Able to communicate complex technical concepts clearly, collaborate across teams, and build stakeholder trust through transparency and reliable delivery
- Able to lead technical initiatives, document approaches clearly, influence peers, and promote adoption of standards and best practices
- Must be able to read, write, and speak English
An Equal Opportunity Employer including Disabled/Veterans
Pay Range
$99,000-$139,000/Annually
An Equal Opportunity Employer including Disabled/Veterans
We endeavor to make this site accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us: careers@dfamilk.com or 877-215-8701.
EEO is the Law