Fulcrum Digital is a global AI-first enterprise transformation company with over 25 years of experience. We partner with enterprises across financial services, insurance, healthcare, retail, manufacturing, higher education, and logistics to move from AI experimentation to scalable business outcomes. Fulcrum Digital works with over 100 global clients, including Fortune 500 enterprises, combining deep industry expertise with capabilities in enterprise AI, digital engineering, cloud modernisation, platform integration, and generative AI.
The Role
We're looking for a Senior Data Engineer to help build and scale the data foundations behind our next generation of AI-powered services. This role blends hands-on data engineering with a strong focus on data quality, working across Databricks, PySpark, Hadoop, and modern Lakehouse architectures to ensure our enterprise data platforms are accurate, reliable, and performant. You'll also play a key role in shaping architecture for AI-driven capabilities, including LLM orchestration, RAG pipelines, and agentic workflows, alongside core non-functional requirements like scalability, security, and reliability. Working closely with cross-functional teams of engineers, architects, and analysts, you'll lead technical design, mentor other engineers, and help ship intelligent, data-driven products end to end.
What You'll Do
- Perform end-to-end validation of data pipelines across ingestion, transformation, and consumption layers
- Execute source-to-target reconciliation and data quality checks, identifying and resolving anomalies
- Define and implement data quality frameworks, metrics, and controls
- Develop and validate data pipelines using PySpark and Databricks across Hadoop, Hive, and Lakehouse environments
- Support ETL/ELT workflows and optimise data processing jobs for performance and scalability
- Write advanced SQL queries for data profiling, reconciliation, and root cause analysis, including complex joins, window functions, CTEs, and aggregations
- Validate and monitor data across Databricks Lakehouse architecture on cloud platforms such as Azure, AWS, or GCP
- Analyze production issues, conduct root cause analysis, and implement proactive monitoring and automated quality checks
- Lead team-level design for AI-powered services, including LLM and agent architecture, RAG pipelines, tool-calling integrations, and data flows
- Design, build, and productionise AI-powered features, including LLM integrations, RAG/retrieval systems, and agentic workflows using frameworks such as LangChain and LlamaIndex
- Implement agentic workflow patterns, including planning/reasoning loops, memory management, grounding, evaluation frameworks, safety mechanisms, and human-in-the-loop design
- Serve as a hands‑on technical contributor across the full engineering lifecycle, from design and coding to testing, debugging, and delivery
- Identify and resolve performance bottlenecks across services, APIs, data pipelines, and AI/LLM workloads
- Collaborate with cross‑functional teams to translate business requirements into scalable, AI-driven technical solutions
- Mentor junior and mid‑level engineers through technical guidance, pairing, code reviews, and knowledge sharing
- Improve engineering quality through code reviews, design reviews, automated testing, secure coding practices, and observability
Requirements
- 6+ years of experience in Data Engineering, Data Quality Engineering, or Data Testing
- Hands‑on experience with Databricks and PySpark
- Strong experience with Hadoop ecosystem components such as Hive, HDFS, Spark, and related Apache technologies
- Advanced SQL expertise for large-scale data validation and analysis
- Experience working with Data Warehouses, Data Lakes, and Lakehouse architectures
- Understanding of Star Schema, Snowflake Schema, and dimensional modeling
- Experience with cloud platforms (Azure, AWS, or GCP)
- Proficiency with Java, Spring Boot, RESTful APIs, Databricks, and Spark/PySpark for data pipelines
- Hands‑on experience building AI‑powered products using LLMs, RAG/retrieval systems, orchestration frameworks (e.g., LangChain, LlamaIndex), and tool‑calling/function‑calling patterns
- Strong grasp of agentic workflow concepts, including planning/reasoning loops, memory management, grounding, evaluation frameworks, safety mechanisms, and human‑in‑the‑loop design
- 6+ years of software engineering experience building scalable, distributed systems and cloud‑native applications in an agile environment
- Proven ability to lead team‑level design, evaluate technical trade‑offs, and translate business needs into secure, maintainable solutions
- Strong collaboration, communication, and problem‑solving skills, comfortable working through ambiguity
- Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience
Nice to Have
- Automated data testing frameworks
- Data observability and monitoring tools
- Experience with Delta Lake, Unity Catalog, or similar technologies
- Knowledge of Airflow, Kafka, or other Apache ecosystem tools
- Exposure to TypeScript/React and SQL
- Experience with cloud platforms (AWS/Azure), containerisation (Docker), and Kubernetes
- Experience with strategic technology partnerships or ecosystem work