Senior Data Engineer – AI Infrastructure Integration, High Performance Compute

Jobtailor

New Jersey

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is looking for an experienced AI/ML Platform Engineer to design, develop, test, validate, and deploy AI/ML capabilities that support infrastructure reliability, capacity forecasting, observability, and automation.

You will apply NLP, statistical modeling, embeddings, anomaly detection, and optimization across data pipelines, APIs, dashboards, and production systems, while guiding the full model lifecycle with cross-team collaboration.

Qualifications

  • 15+ years of experience delivering data science, software engineering, analytics, automation, platform engineering, risk analytics, cloud engineering, SRE, or infrastructure technology solutions.
  • 7+ years of hands-on experience applying AI/ML, NLP, statistical modeling, predictive analytics, optimization, or quantitative methods to enterprise problems.
  • Strong Python programming skills.
  • Practical experience with pandas, NumPy, scikit-learn, TensorFlow, PyTorch, spaCy, Hugging Face Transformers, Gensim, or equivalent tools.
  • Experience with the end-to-end model lifecycle.
  • Experience developing NLP, text analytics, classification, embeddings, recommendation, key driver analysis, network analysis, anomaly detection, or predictive modeling solutions.
  • Experience creating model documentation, validation evidence, implementation procedures, monitoring plans, governance artifacts, or peer review materials.
  • Working knowledge of APIs, data pipelines, relational databases, SQL, dashboards, visualization tools, automation frameworks, version control, CI/CD, observability, and production support practices.
  • Ability to analyze complex structured and unstructured data and quantify business or operational impact through metrics and reporting.
  • Demonstrated experience in Agile delivery environments using Jira, Kanban boards, Confluence, and related platforms.
  • Excellent written and verbal communication skills.
  • Ability to operate across multiple initiatives in a large, matrixed, geographically distributed technology organization.
  • BA or BS in a related quantitative or technical field is listed under Desired Qualifications; advanced Master’s degree preferred.

Responsibilities

  • Design, develop, test, validate, and deploy AI/ML-enabled capabilities for infrastructure reliability, capacity forecasting, observability, operational automation, and enterprise decision-making.
  • Apply NLP, statistical modeling, supervised and unsupervised learning, embeddings, classification, anomaly detection, forecasting, and optimization techniques.
  • Build reusable models, data pipelines, APIs, feature workflows, prompt libraries, automation components, dashboards, and integration patterns.
  • Support the full model lifecycle from use-case intake and data preparation through deployment, monitoring, performance review, and remediation planning.
  • Assess model design, assumptions, limitations, performance, controls, explainability, and implementation risks.
  • Partner with infrastructure, data science, model risk, cyber/risk, architecture, operations, and product teams.
  • Define requirements, success metrics, delivery plans, governance artifacts, and operational handoff criteria.
  • Develop production-grade code, documentation, model artifacts, validation evidence, test automation, and implementation procedures.
  • Advance MLOps, CI/CD, version control, model serving, workflow orchestration, monitoring, and hybrid cloud deployment practices.
  • Communicate technical findings, model outcomes, operational impact, risks, and tradeoffs to technical and executive audiences.

Skills

Python Programming
NLP Techniques
MLOps Practices
Agile Delivery
Data Modeling

Education

BA/BS in a quantitative field

Tools

TensorFlow
scikit-learn
Pandas
NumPy
SpaCy
Hugging Face Transformers
Gensim
Jira
Confluence
CI/CD

Job description

  • Design, develop, test, validate, and deploy AI/ML-enabled capabilities for infrastructure reliability, capacity forecasting, observability, operational automation, and enterprise decision-making
  • Apply NLP, statistical modeling, supervised and unsupervised learning, embeddings, classification, anomaly detection, forecasting, and optimization techniques
  • Build reusable models, data pipelines, APIs, feature workflows, prompt libraries, automation components, dashboards, and integration patterns
  • Support the full model lifecycle from use-case intake and data preparation through deployment, monitoring, performance review, and remediation planning
  • Assess model design, assumptions, limitations, performance, controls, explainability, and implementation risks
  • Partner with infrastructure, data science, model risk, cyber/risk, architecture, operations, and product teams
  • Define requirements, success metrics, delivery plans, governance artifacts, and operational handoff criteria
  • Develop production-grade code, documentation, model artifacts, validation evidence, test automation, and implementation procedures
  • Advance MLOps, CI/CD, version control, model serving, workflow orchestration, monitoring, and hybrid cloud deployment practices
  • Communicate technical findings, model outcomes, operational impact, risks, and tradeoffs to technical and executive audiences
Requirements
  • 15+ years of experience delivering data science, software engineering, analytics, automation, platform engineering, risk analytics, cloud engineering, SRE, or infrastructure technology solutions
  • 7+ years of hands-on experience applying AI/ML, NLP, statistical modeling, predictive analytics, optimization, or quantitative methods to enterprise problems
  • Strong Python programming skills
  • Practical experience with pandas, NumPy, scikit-learn, TensorFlow, PyTorch, spaCy, Hugging Face Transformers, Gensim, or equivalent tools
  • Experience with the end-to-end model lifecycle
  • Experience developing NLP, text analytics, classification, embeddings, recommendation, key driver analysis, network analysis, anomaly detection, or predictive modeling solutions
  • Experience creating model documentation, validation evidence, implementation procedures, monitoring plans, governance artifacts, or peer review materials
  • Working knowledge of APIs, data pipelines, relational databases, SQL, dashboards, visualization tools, automation frameworks, version control, CI/CD, observability, and production support practices
  • Ability to analyze complex structured and unstructured data and quantify business or operational impact through metrics and reporting
  • Demonstrated experience in Agile delivery environments using Jira, Kanban boards, Confluence, and related platforms
  • Excellent written and verbal communication skills
  • Ability to operate across multiple initiatives in a large, matrixed, geographically distributed technology organization
  • BA or BS in a related quantitative or technical field is listed under Desired Qualifications; advanced Master’s degree preferred
Core Competencies

Demonstrates expertise in AI/ML model development, including NLP, statistical modeling, and optimization techniques, while effectively managing the full model lifecycle and collaborating across diverse teams. Proficient in Python and familiar with tools such as TensorFlow and scikit-learn to deliver impactful data-driven solutions.

Highest-signal resume keywords
  • AI/ML Model Development
  • Python Programming
  • NLP Techniques
  • MLOps Practices
  • Agile Delivery
Hard Skills
  • Statistical Modeling
  • Predictive Analytics
  • Optimization Techniques
  • Data Pipeline Development
  • Model Lifecycle Management
  • Anomaly Detection
  • Classification Techniques
  • Feature Engineering
  • Quantitative Methods
  • Production-Grade Code Development
Soft Skills
  • Excellent Communication Skills
  • Collaboration Across Teams
  • Analytical Thinking
  • Problem-Solving
Industry Keywords
  • Infrastructure Reliability
  • Operational Automation
  • Capacity Forecasting
  • Observability
  • Model Risk
  • Data Science
  • Cloud Engineering
  • SRE
  • Analytics
  • Automation
Tools & Technologies
  • TensorFlow
  • Scikit-learn
  • Pandas
  • NumPy
  • SpaCy
  • Hugging Face Transformers
  • Gensim
  • Jira
  • Confluence
  • CI/CD
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • New York (NY)

On-site
USD 180,000 - 260,000
Senior AI Data Scientist – Solutions Developer
Senior AI Data Scientist – Solutions Developer

Jobtailor • Arizona

On-site
USD 120,000 - 180,000
Senior Lead Machine Learning Engineer
Senior Lead Machine Learning Engineer

Jobtailor • Illinois

On-site
USD 150,000 - 190,000
Staff Data Scientist
Staff Data Scientist

Jobtailor • New York (NY)

On-site
USD 180,000 - 240,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • United States

On-site
USD 180,000 - 240,000
Staff Machine Learning Engineer, Document & Vision Intelligence
Staff Machine Learning Engineer, Document & Vision Intelligence

Jobtailor • California (MO)

On-site
USD 150,000 - 230,000
Software Engineer, ML Platform
Software Engineer, ML Platform

Jobtailor • California (MO)

On-site
USD 140,000 - 200,000
Senior AI Scientist
Senior AI Scientist

Jobtailor • Malvern

On-site
USD 120,000 - 190,000
Data Scientist
Data Scientist

Jobtailor • Town of Texas (WI)

On-site
USD 90,000 - 150,000
Distinguished Data Scientist
Distinguished Data Scientist

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000