Role Overview
We are seeking a versatile and forward‑thinking Data & AI Analyst to join our team. This role bridges the gap between traditional data analytics and modern Generative AI implementation. The analyst will parse highly unstructured documents, leveraging advanced language models to drive automation and insights while handling large‑scale datasets. Strong engineering fundamentals and problem‑solving skills are valued over platform‑specific certifications.
Key Responsibilities
- Unstructured Document Parsing: Extract clean, structured insights from a wide variety of complex files, including PDFs, multi‑tab Excel workbooks, financial models, images, and scanned forms.
- Tool Selection & Strategy: Evaluate business problems and select the optimal technology mix, choosing standard data libraries, cloud OCR tools, or large language models.
- Business Application Integration: Bridge standalone AI models with core business applications, ensuring extracted data and insights flow smoothly into internal workflows and systems.
- Automated Report Writing: Design and implement systems that automate the creation of data‑driven reports, summaries, and executive briefs, combining structured data analytics with Generative AI text generation.
- Data Management & Processing: Query, transform, and manage large datasets within cloud data warehouses, using optimized data storage formats such as Parquet for high‑performance pipelines.
- AI & RAG Architecture: Design, build, and maintain Retrieval‑Augmented Generation (RAG) systems to ground LLMs in internal company data, utilizing tools like Azure Document Intelligence for upstream extraction.
- Collaboration & Code Hygiene: Maintain clean, documented, and reusable codebases; participate in team workflows using standard Git version control practices.
Required Skills & Qualifications
Artificial Intelligence & Document Parsing
- Advanced Document Extraction: Hands‑on experience extracting data from messy, semi‑structured, or completely unstructured files, including programmatic parsing or chunking of complex Excel/spreadsheet data and PDFs.
- Extraction Tools: Familiarity with Azure Document Intelligence or equivalent advanced layout extraction tools.
- RAG Frameworks: Practical experience building Retrieval‑Augmented Generation systems, working with vector databases and frameworks such as LangChain or LlamaIndex.
- Model Selection: Familiarity with commercial APIs (OpenAI, Anthropic, Gemini) and lightweight/open‑source models (Llama, Mistral, Phi).
Integration & Software Engineering
- Systems Integration: Experience building and consuming RESTful APIs to connect AI models, data systems, and downstream business applications.
- Version Control: Proficiency with Git and standard platform workflows (GitHub, GitLab, or Azure DevOps).
Data & Analytics Infrastructure
- Big Data Handling: Comfortable working with large‑scale datasets and optimized data formats (JSON, CSV, Parquet).
- Data Pipelines: Strong proficiency manipulating data via Python and writing clean, efficient SQL.
Preferred / Nice‑to‑Have Qualifications
- Experience within the Azure cloud ecosystem and running queries in Snowflake.
- Experience building systems that automatically synthesize data into natural language reports or automated client‑facing summaries.
- Familiarity with prompt engineering techniques and LLM evaluation frameworks.