Data Quality & ML Pipelines Engineer - Onsite NYC

Selby Jennings

New York (NY)

On-site

USD 150,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Selby Jennings is seeking an onsite Data Analyst in New York to develop and maintain batch and real-time data pipelines using Python and SQL. You will ingest, cleanse, validate, and normalize diverse datasets while building data quality controls and anomaly detection.

You will work with large financial datasets, onboard new data providers, and collaborate with engineering and business teams to improve data accessibility. Experience with ML/NLP and LLM-driven workflows is a plus.

Qualifications

  • Bachelor's, Master's, or PhD in Computer Science, Engineering, Mathematics, Statistics, Economics, or a related quantitative discipline.
  • Strong development skills in Python and SQL.
  • Experience with data engineering, analytics, quantitative research, financial data, or large-scale datasets.
  • Knowledge of ETL development, data modeling, and data pipeline design.
  • Exposure to workflow orchestration tools such as Airflow or Dagster.
  • Experience with AWS or GCP.
  • Familiarity with modern data warehouses including Snowflake, BigQuery, or Databricks.
  • Hands-on experience with machine learning, NLP, or LLM-based solutions.
  • Strong understanding of data quality, validation, reconciliation, and monitoring practices.
  • Experience working with messy, incomplete, or high-volume datasets.
  • Familiarity with Git, Linux, testing, and software engineering best practices.
  • Excellent communication skills to explain technical concepts to technical and non-technical audiences.
  • Strong analytical thinking, attention to detail, and problem-solving ability.

Responsibilities

  • Develop and maintain batch and real-time data pipelines using Python and SQL.
  • Ingest, cleanse, validate, and normalize structured and unstructured datasets from multiple sources.
  • Build and improve data quality controls, validation frameworks, reconciliation processes, and anomaly detection.
  • Work with large financial and alternative datasets to maintain accuracy, consistency, and reliability.
  • Support new data provider onboarding by reviewing specs, mapping fields, and integrating APIs.
  • Partner with engineering and business teams to define data requirements and improve accessibility.
  • Apply machine learning and AI to automate data processing, classification, extraction, and enrichment.
  • Evaluate and help implement LLM-based solutions for data mapping, documentation, and QA.
  • Monitor production workflows and investigate data issues, outliers, and operational anomalies.
  • Create documentation and data dictionaries to support scalability and process improvements.
  • Collaborate on cloud-based data platform initiatives and modern data architecture projects.
  • Own datasets and processes, identifying opportunities for efficiency and automation.

Skills

Python
SQL
Data pipelines
Data quality
Machine learning
NLP
Data analysis

Education

Bachelor's/Master's/PhD in CS/Engineering/Math/Stats/Economics

Tools

Airflow
Dagster
AWS
GCP
Snowflake
BigQuery
Databricks
Git
Linux

Job description

Selby Jennings is seeking an onsite Data Analyst in New York to develop and maintain batch and real-time data pipelines using Python and SQL. You will ingest, cleanse, validate, and normalize diverse datasets while building data quality controls and anomaly detection.

You will work with large financial datasets, onboard new data providers, and collaborate with engineering and business teams to improve data accessibility. Experience with ML/NLP and LLM-driven workflows is a plus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer - AI Pipelines & Financial Data
Senior Data Engineer - AI Pipelines & Financial Data

Carlyle Group • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 200,000
Health insurance
Retirement benefits
Paid time off
+1
ML Data Engineer: Scalable Pipelines & Data Quality
ML Data Engineer: Scalable Pipelines & Data Quality

Sesame • San Francisco (CA)

On-site
USD 120,000 - 160,000
401(k) max employer match: 3.5% of compensation
100% employer-paid health, vision, and dental benefits
Unlimited PTO and sick time
+3
Senior Data Engineer: Scalable Data Pipelines & Analytics
Senior Data Engineer: Scalable Data Pipelines & Analytics

UJA Federation of New York • New York (NY)

On-site
USD 110,000 - 130,000
Data Engineer: AI-Ready Pipelines for Finance
Data Engineer: AI-Ready Pipelines for Finance

Mondrian Alpha • New York (NY)

On-site
USD 120,000 - 190,000
Senior Data Engineer — AI Data Pipelines
Senior Data Engineer — AI Data Pipelines

Bloomberg L.P. • New York (NY), Northern (KY)

Hybrid
USD 110,000 - 190,000
401(k) match
Excellent benefits and compensation
Data Engineer - End-to-End ML Pipelines (Equity, NY)
Data Engineer - End-to-End ML Pipelines (Equity, NY)

Raydar • New York (NY)

On-site
USD 200,000 - 250,000
Equity
Comprehensive benefits package
Onsite Data Engineer: Pipelines, APIs & Analytics
Onsite Data Engineer: Pipelines, APIs & Analytics

Largeton Group • Woodbury (MN)

Hybrid
USD 83,000 - 124,000
Lead Data Platform Engineer — Cloud, AI & Scale
Lead Data Platform Engineer — Cloud, AI & Scale

Selby Jennings • New York (NY)

On-site
USD 175,000 - 250,000
Onsite Data Engineer - Finance Pipelines (6-Month Contract)
Onsite Data Engineer - Finance Pipelines (6-Month Contract)

Weekday AI (YC W21) • San Francisco (CA)

On-site
USD 82,656 - 117,096
Senior Data Pipeline Architect for AI-Driven Platforms
Senior Data Pipeline Architect for AI-Driven Platforms

BNY • Pittsburgh

On-site
USD 120,000 - 180,000