Senior Data Engineer

TruLegal (formerly TRU Staffing)

United States

Remoto

USD 145.000 - 165.000

Tempo pieno

17 ore fa
Candidati tra i primi
Generatore di candidature

Non inviare un curriculum generico — genera un curriculum e una lettera di presentazione personalizzati per questo specifico impiego.

Supera i filtri ATS

Descrizione del lavoro

A senior data engineer role at a leading global law firm is seeking a hands-on expert to design, build, and operate enterprise data platforms supporting reporting, analytics, and AI/ML initiatives. The position emphasizes data lakes, warehouses, ETL/ELT pipelines, SQL optimization, and preparing high-quality model-ready datasets.

The role is primarily remote with a requirement to reside within commuting distance of a U.S.

Competenze

  • Bachelor’s degree in Computer Science, IT, Engineering, or related field (or equivalent experience).
  • Extensive experience building and supporting enterprise data platforms, pipelines, and databases.
  • Strong SQL proficiency and scripting (Python) skills.
  • Experience with cloud data platforms (AWS, Azure, GCP) and data lake/warehouse architectures.
  • Knowledge of data governance, lineage, and metadata practices.
  • Familiarity with ML data workflows and feature stores.

Mansioni

  • Design, build, and maintain scalable data lakes, warehouses, and lakehouse environments.
  • Develop and orchestrate ETL/ELT pipelines to ingest and transform data.
  • Collaborate with data scientists to prepare model-ready datasets.
  • Optimize query performance and data models for analytics and BI.
  • Maintain data quality, lineage, and governance across platforms.
  • Mentor and guide junior data engineers.

Conoscenze

SQL
Python
Data modeling
ETL/ELT design
Cloud data platforms
Data pipelines

Formazione

Bachelor’s degree or equivalent experience

Strumenti

Airflow
Spark
Tableau/Power BI

Descrizione del lavoro

Our client, a leading global law firm, is seeking a Senior Data Engineer to design, build, and operate enterprise data platforms supporting reporting, analytics, and AI/ML initiatives. This hands‑on role, high‑level IC roles will focus on data lakes and warehouses, ETL/ELT pipelines, SQL and data performance optimization, database operations, and preparing high‑quality, model‑ready datasets for AI/ML workflows. Ideal candidates will bring extensive data engineering experience, strong SQL and scripting/programming capabilities, cloud data platform expertise, and experience supporting data for AI/ML use cases. The role is primarily remote, but candidates must live within commuting distance of one of the firm's U.S. offices.

Primary applications and platforms include:
  • Document Management: iManage (cloud), SPM, Litera CAM
  • Finance: Aderant Expert Sierra, Chrome River, Time Entry
  • HR: PeopleSoft, Workday
  • Enterprise data lake, data warehouse, and analytics platforms
Responsibilities include:
Data Platform, Data Lake & Pipeline Engineering
  • Design, build, and maintain scalable data lakes, warehouses, and lakehouse environments (on‑premises and/or cloud) to consolidate data from diverse enterprise sources.
  • Develop and orchestrate reliable, automated ETL/ELT pipelines to ingest, transform, and deliver structured and unstructured data.
  • Implement layered data architectures (e.g., raw / curated / consumption or bronze / silver / gold layers) that support reuse across reporting, analytics, and AI workloads.
  • Monitor and maintain pipelines proactively to ensure high availability, timeliness, and data freshness.
  • Apply data quality, validation, and error‑handling practices to ensure accuracy, completeness, and consistency.
  • Establish and maintain data lineage, cataloging, and metadata to support governance and traceability.
Data for AI / Machine Learning
  • Collaborate with data scientists and ML practitioners to curate, prepare, and serve high‑quality datasets for model training, fine‑tuning, and inference.
  • Build and maintain pipelines that transform raw enterprise data into clean, model‑ready datasets.
  • Support feature engineering, feature stores, and reusable data products for AI/ML use cases.
  • Enable AI‑oriented data patterns such as embedding pipelines and retrieval‑augmented workflows, and support integration with vector stores where appropriate.
  • Partner with engineering teams to operationalize data workflows that keep models supplied with reliable, well‑governed data.
Database Administration & Operational Support
  • Administer, monitor, and maintain relational database environments (on‑premises and/or cloud).
  • Perform and automate routine operations, including:
  • Backups and restores (full, differential, and log).
  • Integrity checks and consistency validation.
  • Index maintenance and statistics updates.
  • Monitor and troubleshoot performance issues, including CPU, memory, and I/O bottlenecks, as well as blocking, deadlocks, and long‑running queries.
  • Implement performance tuning strategies such as query optimization, execution plan analysis, and index design and review.
  • Manage database availability and resilience, including high availability, clustering, and disaster recovery planning and validation.
  • Coordinate patching, upgrades, and service releases.
  • Ensure security and compliance through access controls, permissions, encryption, auditing, and vulnerability mitigation.
  • Support scheduled jobs, ETL processes, and automated data workflows.
Query & Data Performance Optimization
  • Write efficient queries and transformations for reporting and analytical workloads.
  • Reduce dataset size and improve refresh and processing performance.
  • Understand and optimize the impact of joins, filters, and aggregations across large datasets.
Analytics & Reporting
  • Translate business questions into queries, metrics, and visualizations.
  • Develop, maintain, and optimize dashboards and reports using leading BI tools (e.g., Tableau, Power BI, or comparable platforms).
  • Design semantic models and data sources for reporting, including fact/dimension modeling (star and snowflake schemas), data shaping, and transformation.
  • Optimize report performance through query tuning and data model optimization (aggregations and relationships).
Collaboration & Leadership
  • Participate actively in group and cross‑functional meetings.
  • Deliver clear, coherent report‑outs to senior management.
  • Work with interdepartmental groups to innovate and improve the firm’s data capabilities.
  • Mentor and train data engineers on the data platform and its business applications.
Qualifications
  • Strong analytical and problem‑solving skills, with a track record of owning systems end‑to‑end.
  • Ability to work independently and manage competing priorities in operational environments.
  • Effective communication with both technical and business teams.
  • Proven ability to collaborate across departments to identify and drive improvements.
Experience
  • Extensive experience building and supporting enterprise data platforms, pipelines, and application databases.
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent experience).
  • Experience managing business‑critical data systems in a global environment.
  • Proficiency with SQL and at least one programming/scripting language commonly used in data engineering (e.g., Python).
  • Hands‑on experience with data warehousing, data lakes, or lakehouse architectures.
  • Experience building ETL/ELT pipelines and working with data orchestration tools.
  • Experience with cloud data platforms (e.g., AWS, Azure, Google Cloud, Snowflake, Databricks, or comparable).
  • Experience preparing and serving data for AI/ML model training, fine‑tuning, or inference.
  • Familiarity with distributed data processing frameworks (e.g., Apache Spark).
  • Familiarity with pipeline orchestration tools (e.g., Apache Airflow or similar).
  • Familiarity with ML and AI concepts, including feature stores, vector databases, and embedding/RAG pipelines.
  • Familiarity with automation and scripting (e.g., Python, PowerShell, or Bash).
  • Exposure to DevOps, MLOps, or CI/CD practices for data or BI deployments.

Expected salary for this role is $145,000 - $165,000, commensurate with experience, training, skills, qualifications, and other market factors.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior Data Engineer - Data Lake / AI
Senior Data Engineer - Data Lake / AI

Onward Legal Innovation • Washington

Remoto
USD 170.000 - 200.000
Senior Data Engineer
Senior Data Engineer

Tru Staffing Inc • Northern (KY)

In loco
USD 145.000 - 165.000
Senior Data Engineer
Senior Data Engineer

The Phoenix Group • Stati Uniti

In loco
USD 120.000 - 180.000
Performance bonuses
Health benefits
Flexible work
Senior Data Engineer
Senior Data Engineer

Peyton Resource Group • Houston (TX)

In loco
USD 120.000 - 150.000
Senior Data Engineer
Senior Data Engineer

SDL Tech Search • Boston (MA)

In loco
USD 140.000 - 190.000
Remote Senior Data Platform Engineer for AI/ML
Remote Senior Data Platform Engineer for AI/ML

TruLegal (formerly TRU Staffing) • Stati Uniti

Remoto
USD 145.000 - 165.000
Senior Data Engineer - Full Time Only - Remote
Senior Data Engineer - Full Time Only - Remote

GD Resources LLC • Stati Uniti

Remoto
USD 126.000 - 154.000
Remote work
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

In loco
USD 120.000 - 150.000
Senior Data Engineer
Senior Data Engineer

Mindlance • Reston (VA)

In loco
USD 120.000 - 180.000
Senior Data Engineer
Senior Data Engineer

Landing Point • Wakefield (MA)

Ibrido
USD 110.000 - 130.000