RADcube is hiring a Senior Data & AI Engineer in Carmel, Indiana (onsite). The role focuses on data modeling, semantic layers, and metadata foundations so AI systems can answer questions accurately, including text-to-SQL and RAG-style workflows, alongside client discovery and stakeholder alignment.
Responsibilities
- Build and maintain data models including dimensional, relational, and lakehouse designs that align with team standards.
- Investigate and document unfamiliar or legacy schemas by producing ER diagrams, data dictionaries, join paths, and lineage.
- Develop and optimize SQL, transformations, and data pipelines on cloud data platforms.
- Convert raw tables into business-facing semantic models covering metrics, dimensions, hierarchies, and relationships.
- Write and enrich schema metadata and descriptions to support improved performance for LLM text-to-SQL and generative BI accuracy.
- Partner with AI engineers on RAG pipelines, agent tools, and prompt design when the system relies on structured data.
- Test and evaluate AI-generated queries for correctness, including contributions to test sets and guardrails.
- Participate in client discovery to understand business processes, KPIs, and reporting requirements.
- Translate business questions into data requirements and validate metric definitions with stakeholders.
- Communicate data findings clearly to both technical and non-technical audiences.
- Apply data quality checks, naming standards, and documentation practices.
- Follow relevant governance and compliance requirements (GxP, HIPAA) where applicable.
- Review peers’ work and support junior engineers when needed.
Requirements
- 6+ years of experience in data engineering, analytics engineering, or BI development.
- Strong SQL skills and a solid understanding of relational and dimensional modeling.
- Proven ability to learn and navigate large enterprise schemas (for example, SAP, Salesforce, MES, or similar).
- Hands‑on experience with AWS (Redshift, Glue, Athena, S3) and/or Azure (Synapse, Fabric, Data Factory), plus Databricks or Snowflake.
- Proficiency in Python for data work.
- Practical exposure to LLMs on structured data, including text-to-SQL, semantic layers, or AI‑assisted analytics.
- Good business sense and comfort discussing KPIs and processes with stakeholders.
Technologies
SQL, Python, AWS (Redshift, Glue, Athena, S3), Azure (Synapse, Fabric, Data Factory), Databricks, Snowflake, SAP, Salesforce, MES, LLMs, text‑to‑SQL, RAG, LangChain, LangGraph, Bedrock Agents, MCP, Unity Catalog, Collibra, AWS DataZone, dbt, Cube, dbt Semantic Layer, LookML, vector databases, knowledge graphs, agentic frameworks.
What You Bring (Nice‑to‑Have)
- Experience with pharma, life sciences, manufacturing and quality, or healthcare data.
- Experience with dbt or semantic layer tooling such as Cube, dbt Semantic Layer, or LookML.
- Familiarity with vector databases, knowledge graphs, or agentic frameworks (LangChain/LangGraph, Bedrock Agents, MCP).
- Experience with data catalog tools such as Unity Catalog, Collibra, or AWS DataZone.
- AWS, Azure, or Databricks certifications.
What Success Looks Like (First 6 Months)
- Semantic models and metadata delivered for at least one accelerator or client use case.
- Measurable improvement in AI-generated query accuracy on datasets you own.
- Schema documentation that enables other team members to adopt and use it.
- Stakeholder confidence that you can understand both their data and their business context.