Senior AI Data Engineer

Cooley LLP

New York (NY)

On-site

USD 220,000 - 250,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cooley LLP is seeking a Senior AI Data Engineer to join the Innovation team. You will build production RAG and embedding pipelines, integrate LLM features, and ensure observability for AI-powered legal applications.

Work with the Northstar Lakehouse and Databricks stack, partnering with data engineers to maintain governance, data quality, and scalable AI infrastructure for the firm’s analytics needs.

Qualifications

  • 5+ years of data/AI engineering with production responsibility for RAG pipelines and AI infrastructure.
  • Strong Python proficiency and production-grade code practices.
  • Experience with LLM integration, prompt engineering, and observability.

Responsibilities

  • Design, build, and maintain production RAG pipelines and embedding infrastructure.
  • Develop and monitor LLM integration, prompts, and feature observability.
  • Collaborate with Data Engineering & Platform to ensure governed data for AI workloads.

Skills

RAG pipelines
LLM integration
Python
Databricks
Vector search

Tools

Pinecone
Databricks Vector Search
pgvector
Chroma

Job description

Senior AI Data Engineer

Cooley AI is seeking a Senior AI Data Engineer to join the Innovation team. About Cooley AI: Cooley AI is the wholly-owned subsidiary of Cooley LLP. The frontier development arm of the leading technology law firm in the world. Our mission is to bring the best-in-class practices from the technology field into the legal domain to help lead our firm and the legal industry into the AI era. All employees of Cooley AI are employees of Cooley and are seconded to Cooley AI.

Position summary: As a leading technology law firm, Cooley is determined to become a leader in the digital practice of law. The Senior AI Data Engineer is a senior technical individual contributor within the Data Products function, responsible for building and maintaining the AI infrastructure layer that powers the firm's AI-powered data applications on the Northstar Lakehouse. Working as a peer to the Senior Data Application Engineer, this role is responsible for the intelligence layer beneath the application: the RAG pipelines, embedding infrastructure, vector search implementation, LLM integration patterns, and AI feature observability that make AI-powered legal intelligence products reliable, performant, and trustworthy in production. The Senior AI Data Engineer brings production AI engineering depth to a greenfield AI application build, operating at the intersection of data engineering, applied machine learning, and enterprise application development. This is a high-impact individual contributor role and will build production RAG and LLM-powered systems with the understanding of what it takes to make AI features work reliably at scale.

Position responsibilities
  • RAG Pipeline & Retrieval Infrastructure: Design, build, and maintain production RAG pipelines that ingest, process, chunk, embed, and index firm data assets for retrieval by AI-powered legal intelligence applications. Own the full RAG pipeline lifecycle from document ingestion through retrieval quality monitoring. Own the vector search implementation including index design, embedding model selection and maintenance, retrieval strategy configuration, and ongoing retrieval quality assessment. Ensure the vector search layer returns accurate, contextually relevant results. Build and maintain embedding pipelines that transform structured and unstructured firm data assets into vector representations consumable by retrieval systems and LLM-powered features. Apply chunking strategies, metadata enrichment, and embedding quality validation appropriate to legal sector document complexity. Monitor and maintain retrieval quality in production, establishing metrics for retrieval relevance, embedding freshness, index coverage, and pipeline execution health that surface degradation before it affects application users. Collaborate with the Data Engineering & Platform function to ensure RAG pipeline data sources are correctly integrated with the Northstar Lakehouse Silver and Gold layer data assets, consuming governed, certified data rather than bypassing platform governance controls.
  • LLM Integration & AI Feature Engineering: Build and maintain the LLM integration layer for Data Applications squad products, implementing prompt engineering patterns, output validation frameworks, fallback handling, latency management, and cost-aware API usage patterns that make LLM-powered features reliable in production. Implement AI output handling patterns that ensure application features degrade gracefully when model responses are unexpected, latency is high, or retrieval quality issues affect the context provided to the LLM. Design for failure from the outset rather than retrofitting resilience. Build and maintain AI feature observability tooling that provides the Data Applications squad with visibility into LLM response quality, retrieval relevance, token consumption trends, latency patterns, and user interaction signals that inform ongoing AI feature improvement. Evaluate and integrate Databricks-native AI capabilities including FMAPI, Mosaic AI, Databricks Model Serving, and Databricks Vector Search as they mature, making architectural recommendations on adoption timing and integration patterns alongside the Platform Architect. Apply responsible AI practices to all AI feature implementations including LLM output validation, bias awareness in retrieval systems, appropriate user-facing communication of AI output reliability, and compliance with the firm's AI governance framework. Stay current on the rapidly evolving AI engineering landscape including advances in RAG architecture, embedding models, LLM APIs, and AI application patterns, bringing relevant innovations to the squad's technical roadmap.
  • Data Pipeline & Platform Integration: Build and maintain data pipelines that prepare, transform, and deliver data assets to AI consumption layers, ensuring data flowing into RAG pipelines, embedding infrastructure, and LLM context windows is accurate, current, and appropriately governed. Work within Unity Catalog access controls and data classification policies, understanding the governance framework and ensuring AI pipeline implementations respect data classification, access control, and audit logging requirements rather than working around them. Collaborate with the Data Engineering & Platform function on the data contracts between the Northstar Lakehouse Silver and Gold layers and the AI infrastructure layer, ensuring AI pipelines consume data through governed interfaces and surface data freshness and quality signals to the application layer. Implement pipeline-level data quality checks within AI ingestion pipelines, validating that data entering RAG pipelines and embedding infrastructure meets the completeness, accuracy, and freshness standards required for reliable AI feature performance. Apply production engineering discipline to all AI pipeline implementations including version control, CI/CD integration, testing frameworks, documentation standards, and deployment through the established platform CI/CD framework.
  • Data Applications Squad Collaboration: Work as a close technical partner to the Senior Data Application Engineer, owning the AI and data infrastructure layer that the React application on Databricks Appkit depends on. Coordinate proactively on the interface between the application layer and the intelligence layer so neither engineering track becomes a blocker for the other. Participate fully in Data Applications squad Agile ceremonies including sprint planning, daily standups, sprint reviews, and retrospectives. Write clear, well-estimated tickets that accurately reflect the complexity of AI infrastructure work and distinguish spike work from build work. Communicate proactively about AI infrastructure dependencies, retrieval quality issues, LLM API changes, and embedding pipeline health that may affect application behavior or sprint commitments. Contribute to the Data Applications squad engineering culture by modeling production-grade AI engineering practices, sharing knowledge of RAG architecture and LLM integration patterns, and peer reviewing code with teaching intent. Engage with the Director, Data Products on the AI feature roadmap, providing honest technical assessments of AI feature feasibility, implementation complexity, and the infrastructure investment required to make AI-powered features production-grade.
  • All other duties as assigned or required.
Skills and experience

Required:

  • After orientation at Cooley LLP, exhibit proficiency in the Microsoft Office suite, iManage, and other firm applications
  • Ability to work extended and/or weekend hours, as required
  • Ability to travel, as required
  • 5+ years of data engineering or AI engineering experience with hands-on production responsibility for RAG pipelines, LLM integration, or AI-powered application data infrastructure in a cloud data platform environment
  • Demonstrated production RAG pipeline experience including document ingestion, chunking strategies, embedding generation, vector index design, retrieval quality monitoring, and ongoing pipeline maintenance at scale
  • Strong LLM API integration experience in production applications including prompt engineering, output validation, fallback handling, latency management, and cost-aware usage patterns using OpenAI, Anthropic, Google, or comparable LLM APIs
  • Hands-on vector search implementation experience including Pinecone, Databricks Vector Search, pgvector, Chroma, or comparable vector database platforms in a production application context
  • Strong Python proficiency including modular package design, test-driven development, and production-grade code standards for AI pipeline and infrastructure development
  • Databricks experience including Delta Lake, Unity Catalog, Databricks SQL, and medallion architecture conventions
  • Experience with AI feature observability including LLM response quality monitoring, retrieval relevance metrics, token consumption tracking, and latency alerting for production AI systems
  • Production data engineering experience including pipeline development, data quality validation, CI/CD integration, and version control discipline in a cloud data platform environment
  • Comfort using GitHub for version control, pull requests, and CI/CD pipeline integration in a collaborative engineering context
  • Demonstrated Agile delivery fluency including sprint ceremonies, ticket writing, estimation, and definition of done discipline

Preferred:

  • Experience with Databricks Vector Search, FMAPI, Mosaic AI, or Databricks Model Serving in a production AI application context
  • Familiarity with LangChain, LangGraph, LlamaIndex, or comparable AI orchestration frameworks for RAG pipeline development and LLM workflow integration
  • Experience with embedding model evaluation, selection, and fine-tuning including open-source embedding models and the trade-offs between model quality, latency, and cost in production retrieval systems
  • Legal sector, professional services, or enterprise SaaS experience with exposure to complex document processing, knowledge retrieval, or AI-powered workflow automation in a regulated or high-stakes data environment
  • Familiarity with React application architecture and how AI features are surfaced to users through front-end application components, sufficient to collaborate effectively with the Senior Data Application Engineer on the application and intelligence layer interface
  • Experience with Databricks Appkit or Databricks Apps for AI-powered application hosting
  • Exposure to agentic AI patterns including tool-calling architectures, agent memory and state management, and multi-step LLM workflow orchestration using LangGraph, OpenAI Assistants, Claude with tools, or comparable frameworks
  • Experience with PySpark or distributed data processing for large-scale document ingestion and embedding pipeline development at volumes that exceed single-node processing capacity
  • Comfort using AI-assisted development tools such as Claude Code, GitHub Copilot, Cursor, or comparable to accelerate AI pipeline development, documentation, and code review workflows
Competencies
  • Strong problem-solving instincts with the ability to design AI systems that are reliable and observable in production rather than just functional in development
  • Excellent engineering discipline with high standards for AI pipeline reliability, observability, and resilience that make AI features trustworthy in production rather than impressive in demos
  • Strong collaborative instincts with the ability to work as a close technical partner to the Senior Data Application Engineer and engage effectively across data engineering, governance, and platform functions
  • Excellent attention to detail with the discipline to monitor AI output quality, retrieval relevance, and pipeline health proactively rather than reactively
  • Growth-oriented mindset with genuine enthusiasm for the rapidly evolving AI engineering landscape and a track record of bringing new techniques into production thoughtfully rather than chasing novelty
  • Excellent written and verbal communication skills with the ability to explain AI infrastructure decisions, retrieval quality trade-offs, and LLM integration constraints clearly to both technical peers and non-technical product stakeholders
  • Strong delivery orientation with consistent sprint commitment follow-through, proactive blocker communication, and definition of done discipline in an Agile squad environment
  • Strong judgment Unwavering ability to handle and maintain confidentiality regarding firm information, projects, and client data

Cooley LLP offers a competitive compensation and excellent benefits program. EOE. The expected annual pay range for this position with a full-time schedule is $220,000 - $250,000. Please note that final offer amount will be dependent on geographic location, applicable experience and skillset of the candidate. Senior level candidates may be considered for this position and would be eligible for a higher salary range based on experience. We offer a full range of elective benefits including medical, health savings account (with applicable medical plan), dental, vision, health and/or dependent care flexible spending accounts, pre-tax commuter benefits, life insurance, AD&D, long-term care coverage, backup care for children and/or adults and other parental support benefits. In addition to elective benefit options, benefited employees receive firm-paid life insurance, AD&D, LTD, short term medical benefits as well as 21 days of Paid Time Off (“PTO”) and 10 paid holidays each year. We provide generous parental leave and fertility benefits. New employees will attend a detailed benefit orientation to learn more about our many benefits and resources.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Application Engineer
Senior Data Application Engineer

Cooley AI • Seattle (WA)

On-site
USD 220,000 - 250,000
Medical benefits
Dental benefits
Vision benefits
+1
Senior AI/ML Engineer
Senior AI/ML Engineer

Cooley AI • Washington

On-site
USD 200,000 - 300,000
Medical coverage
Health Savings Account
Dental
+7
Senior Data Architect
Senior Data Architect

Cooley LLP • New York (NY)

On-site
USD 200,000 - 250,000
Senior Software Engineer
Senior Software Engineer

Cooley AI • San Francisco (CA)

On-site
USD 170,000 - 250,000
Medical benefits
PTO 21 days
Paid holidays
+1
Product Support Engineer
Product Support Engineer

Cooley AI • Seattle (WA)

On-site
USD 120,000 - 175,000
Medical coverage
Health Savings Account
Dental
+7
Senior User Experience/User Interface (“UX/UI”) Designer
Senior User Experience/User Interface (“UX/UI”) Designer

Cooley LLP • New York (NY)

On-site
USD 135,000 - 200,000
Practice Innovation Specialist
Practice Innovation Specialist

Cooley AI • Los Angeles (CA)

On-site
USD 100,000 - 150,000
Medical benefits
Health Savings Account
Dental and Vision
+4
Technology Cyber Security Architect
Technology Cyber Security Architect

Cooley LLP • Washington

On-site
USD 120,000 - 175,000
Medical and dental benefits
21 days of Paid Time Off (PTO)
Generous parental leave and fertility benefits
Data & AI Engineer
Data & AI Engineer

1P284 THE CARLYLE GROUP EMPLOYEE CO., LLC • Ova (KY)

On-site
USD 160,000 - 180,000
Retirement benefits
Health insurance
Life and disability insurance
+4
Competitive Intelligence Analyst
Competitive Intelligence Analyst

Cooley LLP • San Francisco (CA)

On-site
USD 85,000 - 120,000
Medical insurance
Paid Time Off (21 days)
Parental leave