Skip to content Back Senior Data Engineer
Anywhere| United States Direct Hire Our client, a leading global law firm, is seeking a Senior Data Engineer to design, build, and operate enterprise data platforms supporting reporting, analytics, and AI/ML initiatives. This hands‑on role, high-level IC roles will focus on data lakes and warehouses, ETL/ELT pipelines, SQL and data performance optimization, database operations, and preparing high-quality, model-ready datasets for AI/ML workflows. Ideal candidates will bring extensive data engineering experience, strong SQL and scripting/programming capabilities, cloud data platform expertise, and experience supporting data for AI/ML use cases. The role is primarily remote, but candidates must live within commuting distance of one of the firm's U.S. offices.
Primary applications and platforms include:
- Document Management: iManage (cloud), SPM, Litera CAM
- Finance: Aderant Expert Sierra, Chrome River, Time Entry
- HR: PeopleSoft, Workday
- Enterprise data lake, data warehouse, and analytics platforms
Responsibilities include: Data Platform, Data Lake & Pipeline Engineering
- Design, build, and maintain scalable data lakes, warehouses, and lakehouse environments (on-premises and/or cloud) to consolidate data from diverse enterprise sources.
- Develop and orchestrate reliable, automated ETL/ELT pipelines to ingest, transform, and deliver structured and unstructured data.
- Implement layered data architectures (e.g., raw / curated / consumption or bronze / silver / gold layers) that support reuse across reporting, analytics, and AI workloads.
- Monitor and maintain pipelines proactively to ensure high availability, timeliness, and data freshness.
- Apply data quality, validation, and error-handling practices to ensure accuracy, completeness, and consistency.
- Establish and maintain data lineage, cataloging, and metadata to support governance and traceability.
Data for AI / Machine Learning
- Collaborate with data scientists and ML practitioners to curate, prepare, and serve high-quality datasets for model training, fine-tuning, and inference.
- Build and maintain pipelines that transform raw enterprise data into clean, model-ready datasets.
- Support feature engineering, feature stores, and reusable data products for AI/ML use cases.
- Enable AI-oriented data patterns such as embedding pipelines and retrieval-augmented workflows, and support integration with vector stores where appropriate.
- Partner with engineering teams to operationalize data workflows that keep models supplied with reliable, well-governed data.
Database Administration & Operational Support
- Administer, monitor, and maintain relational database environments (on-premises and/or cloud).
- Perform and automate routine operations, including:
- Backups and restores (full, differential, and log).
- Integrity checks and consistency validation.
- Index maintenance and statistics updates.
- Monitor and troubleshoot performance issues, including CPU, memory, and I/O bottlenecks, as well as blocking, deadlocks, and long-running queries.
- Implement performance tuning strategies such as query optimization, execution plan analysis, and index design and review.
- Manage database availability and resilience, including high availability, clustering, and disaster recovery planning and validation.
- Coordinate patching, upgrades, and service releases.
- Ensure security and compliance through access controls, permissions, encryption, auditing, and vulnerability mitigation.
- Support scheduled jobs, ETL processes, and automated data workflows.
Query & Data Performance Optimization
- Write efficient queries and transformations for reporting and analytical workloads.
- Reduce dataset size and improve refresh and processing performance.
- Understand and optimize the impact of joins, filters, and aggregations across large datasets.
Analytics & Reporting
- Translate business questions into queries, metrics, and visualizations.
- Develop, maintain, and optimize dashboards and reports using leading BI tools (e.g., Tableau, Power BI, or comparable platforms).
- Design semantic models and data sources for reporting, including fact/dimension modeling (star and snowflake schemas), data shaping, and transformation.
- Optimize report performance through query tuning and data model optimization (aggregations and relationships).
Collaboration & Leadership
- Participate actively in group and cross-functional meetings.
- Deliver clear, coherent report-outs to senior management.
- Work with interdepartmental groups to innovate and improve the firm’s data capabilities.
- Mentor and train data engineers on the data platform and its business applications.
Qualifications
- Strong analytical and problem-solving skills, with a track record of owning systems end-to-end.
- Ability to work independently and manage competing priorities in operational environments.
- Effective communication with both technical and business teams.
- Proven ability to collaborate across departments to identify and drive improvements.
Experience
- Extensive experience building and supporting enterprise data platforms, pipelines, and application databases.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent experience).
- Experience managing business-critical data systems in a global environment.
- Proficiency with SQL and at least one programming/scripting language commonly used in data engineering (e.g., Python).
- Hands‑on experience with data warehousing, data lakes, or lakehouse architectures.
- Experience building ETL/ELT pipelines and working with data orchestration tools.
- Experience with cloud data platforms (e.g., AWS, Azure, Google Cloud, Snowflake, Databricks, or comparable).
- Experience preparing and serving data for AI/ML model training, fine‑tuning, or inference.
- Familiarity with distributed data processing frameworks (e.g., Apache Spark).
- Familiarity with pipeline orchestration tools (e.g., Apache Airflow or similar).
- Familiarity with ML and AI concepts, including feature stores, vector databases, and embedding/RAG pipelines.
- Familiarity with automation and scripting (e.g., Python, PowerShell, or Bash).
- Exposure to DevOps, MLOps, or CI/CD practices for data or BI deployments.
Expected salary for thisrole is $145,000 - $165,000, commensurate with experience, training, skills, qualifications, and other market factors.
#LI-BJ1 #LI-REMOTE
Job ID:7606
word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word word
mmMwWLliI0fiflO&1
mmMwWLliI0fiflO&1
mmMwWLliI0fiflO&1
mmMwWLliI0fiflO&1
mmMwWLliI0fiflO&1
mmMwWLliI0fiflO&1
mmMwWLliI0fiflO&1