An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Latent View Analytics Limited in the Philippines seeks a Senior Data Engineer focused on building AI-first infrastructure atop the Databricks stack. You will architect Lakehouse with Bronze/Silver/Gold layers and lead data pipelines for ML and LLM workloads, while enabling internal teams with self-service tooling.
You will optimize Unity Catalog governance, deploy Delta Live Tables, manage serverless SQL compute, and ensure scalable, low-latency data delivery for customer-facing AI features.
We need a senior data engineering resource who is super deep into building the infrastructure layer, has expertise in the Databricks stack and thinks AI first in terms building out the infrastructure - We are looking for this person as a Senior leader for the customer enablement stack who can support and help us get to the next level while building a truly AI centric stack for us.
Architecting the Lakehouse: Lead the design and implementation of a robust Medallion Architecture (Bronze/Silver/Gold) specifically optimized for downstream machine learning and LLM consumption.
Vector Database Integration: Architect the seamless integration of Databricks Vector Search and managed vector databases to support RAG-based applications.
Model-Ready Pipelines: Build "feature-first" data pipelines where data is versioned, lineage-tracked, and ready for training without manual preprocessing.
Unity Catalog Governance: Implement enterprise-wide data governance, security, and discovery using Unity Catalog to ensure AI models access data ethically and securely.
Delta Live Tables (DLT): Deploy and manage complex, streaming data pipelines using DLT to reduce operational overhead and increase data reliability.
Compute Optimization: Manage and optimize Serverless SQL warehouses and automated cluster scaling to balance high performance with cost-efficiency.
Internal Productization: Treat the data stack as a product, building self-service tooling that allows internal "customers" (DS/ML teams) to spin up environments and access clean data instantly.
Performance Engineering: Debug and resolve deep-seated architectural bottlenecks in Spark jobs to ensure sub-second latency for customer-facing AI features.
Technical Evangelism: Act as the bridge between core engineering and customer-facing teams, translating complex infrastructure capabilities into business value.
CI/CD for Data: Establish rigorous CI/CD practices for infrastructure-as-code (Terraform/Pulumi) and data pipeline deployments.
Monitoring & Observability: Implement advanced monitoring for data quality (Great Expectations/Monte Carlo) and model drift, ensuring the AI stack is "self-healing."
Agentic Framework Support: Design the backend infra to support LangChain or LlamaIndex workflows, ensuring the data retrieval layer is fast enough for agentic reasoning.