Principal Data Scientist

Jobtailor

Dearborn (MO)

On-site

USD 150,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Jobtailor seeks a lead data scientist to design and deploy end-to-end ML models for complex platform data challenges. You will own identity resolution, data quality and scalable data pipelines while partnering with cross-functional teams to deliver impactful analytics.

The role emphasizes mentoring, ML monitoring tools, and advancing data warehousing and cloud-native data services with Python, SQL, and BigQuery. On-site in Dearborn, MO with strong collaboration.

Qualifications

  • Ph.D. in Computer Science, Statistics or Mathematics.
  • 5+ years in data science or related field.
  • 5+ years Python development with Pandas/NumPy/Scikit-learn.
  • 5+ years ML frameworks and libraries (e.g., PySpark, BigQuery ML, Pandas, Scipy-Weave, etc.).
  • 3+ years SQL experience for complex datasets.

Responsibilities

  • Design, build and deploy ML algorithms for platform data challenges.
  • Develop identity resolution algorithms and production deployment.
  • Identify and resolve data quality issues across data warehouses/lakes.
  • Build ML-driven monitoring tools for anomaly detection and data quality checks.
  • Mentor data engineers and collaborate with global stakeholders.

Skills

Machine Learning
Python Development
Data Quality
SQL Query Optimization
Mentoring
Project Leadership
Communication

Education

Ph.D. in Computer Science
Ph.D. in Statistics
Ph.D. in Mathematics

Tools

Python
Pandas
NumPy
Scikit-learn
Pyspark
BigQuery
SQL
Docker
Graph ML libraries
NLTK

Job description

  • Design, build, and deploy end-to-end machine learning algorithms for complex platform data challenges
  • Develop identity resolution algorithms, refine them for accuracy and performance at scale, and support production deployment
  • Identify, analyze, and resolve complex data quality issues across large-scale data warehouses and data lakes
  • Design and implement ML-driven monitoring tools for anomaly detection, data cleansing, standardization, and validation
  • Develop internal tools, libraries, and automation scripts for data lineage, metadata management, and automated data quality checks
  • Build and manage Docker images and automate algorithms for orchestration and scheduling
  • Apply statistical analysis and predictive modeling to evaluate large-scale data pipelines, ETL/ELT processes, and data storage solutions
  • Implement data-driven improvements to optimize data flow, achieve sub-second latency, and improve resource utilization
  • Perform root cause analysis on data discrepancies, performance bottlenecks, and system failures
  • Drive strategic development of data science capabilities within the Product group and lead R&D efforts
  • Leverage Generative AI and LLMs to identify problems and develop solutions in data platform and data warehouse environments
  • Partner with data engineers and architects to design, optimize, and maintain scalable data structures and cloud-native data services
  • Manage data ingestion from various sources and formats into BigQuery and other data platforms
  • Implement transformation processes for downstream analytical consumption
  • Act as a key liaison and project leader, mentor data engineers, and collaborate with global stakeholders
  • Translate complex business requirements into technical specifications for data solutions
  • Serve as a technical conduit between central privacy, product, and engineering teams for end-to-end system design
  • Create documentation for data structures, data quality rules, and analytical findings
  • Share expertise, mentor junior team members, and foster best practices across the organization
Requirements
  • Ph.D. in Computer Science, Statistics, Mathematics or a related field
  • 5 years of experience in the job offered or a related occupation
  • 5 years of experience in data science or advanced data engineering
  • 5 years of experience with Python development, including Pandas, NumPy, and Scikit-learn, for data manipulation, analysis, and scripting
  • 5 years of experience using machine learning frameworks and libraries, including at least 5 of: Pyspark, BigQuery ML, Pandas, Scipy-Weave, Multiprocessing, Graph ML libraries, Neo4j, NLTK, or Matplotlib
  • 5 years of experience designing, implementing, and deploying machine learning models for complex data problems, including anomaly detection or predictive maintenance, with algorithm fine-tuning for performance and accuracy at scale
  • 5 years of experience researching, evaluating, and integrating new ML capabilities
  • 3 years of experience using SQL for querying, manipulating, and optimizing complex datasets, including query tuning
  • Must be legally authorized to work in the United States
  • Must live within a reasonable commuting distance from the Dearborn, Michigan worksite
Core Competencies

Demonstrates expertise in designing and deploying machine learning algorithms, with a strong focus on data quality, anomaly detection, and performance optimization. Proficient in Python and various machine learning frameworks, with a proven ability to mentor teams and collaborate across functions.

Highest-signal resume keywords
  • Machine Learning Model Deployment
  • Python Development
  • Data Quality Management
  • Statistical Analysis
  • SQL Query Optimization
Hard Skills
  • Machine Learning Algorithms
  • Data Manipulation
  • Predictive Modeling
  • Anomaly Detection
  • Data Engineering
  • ETL Processes
  • Data Cleansing
  • Statistical Analysis
  • Algorithm Fine-Tuning
  • Data Lineage
Soft Skills
  • Mentoring
  • Collaboration
  • Project Leadership
  • Communication
Certifications & Qualifications
  • Ph.D. in Computer Science
  • Ph.D. in Statistics
  • Ph.D. in Mathematics
Industry Keywords
  • Data Science
  • Data Engineering
  • Data Warehousing
  • Cloud-Native Data Services
  • Data Quality Rules
Tools & Technologies
  • Python
  • Pandas
  • NumPy
  • Scikit-learn
  • Pyspark
  • BigQuery
  • SQL
  • Docker
  • Graph ML libraries
  • NLTK
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Machine Learning Engineer, Python, AWS, SQL, GenAI
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI

Jobtailor • New York (NY)

On-site
USD 140,000 - 210,000
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Jobtailor • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Principal Data Scientist
Principal Data Scientist

Jobtailor • Acton (MA)

On-site
USD 180,000 - 260,000
Lead Data Engineer
Lead Data Engineer

Jobtailor • Massachusetts

Hybrid
USD 150,000 - 210,000
Principal Engineer, GenAI Architecture
Principal Engineer, GenAI Architecture

Jobtailor • North Carolina

On-site
USD 120,000 - 160,000
Data Scientist
Data Scientist

Jobtailor • Arlington (VA)

On-site
USD 110,000 - 160,000
Senior Manager, Data Science
Senior Manager, Data Science

Jobtailor • Charlotte (NC)

On-site
USD 140,000 - 190,000
Senior Software Engineer – AI Quality and Productivity
Senior Software Engineer – AI Quality and Productivity

Jobtailor • Missouri

On-site
USD 125,000 - 180,000
AI/ML Data Scientist, GPSSC
AI/ML Data Scientist, GPSSC

Jobtailor • Missouri

On-site
USD 120,000 - 170,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 170,000