AI/ML Lead Data Engineer - Automation/Image Processing

JPMorganChase

Tampa (FL)

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive health care coverage
On-site health and wellness centers
Retirement savings plan
Tuition reimbursement
Mental health support
Financial coaching

Job summary

JPMorganChase is seeking a skilled Data Engineer in Tampa, Florida, to design and maintain scalable data pipelines and architecture on AWS. The candidate will work closely with machine learning engineers, developing solutions for the processing of scanned documents and ensuring data quality and governance throughout the lifecycle. Required qualifications include 5+ years of experience, proficiency in Java, Python, and AWS services, as well as strong leadership capabilities. A competitive salary and comprehensive benefits package are offered.

Qualifications

  • 5+ years applied experience in Data Engineering.
  • Deep hands-on experience with AWS cloud services.
  • Advanced knowledge of Oracle databases.
  • Experience with OCR and computer vision pipelines.
  • Excellent leadership and communication skills.

Responsibilities

  • Design and maintain scalable data pipelines.
  • Architect data solutions on AWS cloud services.
  • Develop robust image preprocessing and OCR integration pipelines.
  • Collaborate closely with data scientists and ML engineers.
  • Lead and mentor a team of data engineers.

Skills

Java
Groovy
Python
Image file handling
AWS cloud services
AWS EKS
Oracle databases
OCR technologies
CI/CD pipelines
Data governance
Leadership

Education

Formal training or certification on Data Engineering

Tools

Terraform
CloudFormation
Docker
Jenkins

Job description

Join us as we embark on a journey of collaboration and innovation, where your unique skills and talents will be valued and celebrated. Together we will create a brighter future and make a meaningful difference.

Job Responsibilities
  • Design, build, and maintain scalable, high-performance data pipelines and infrastructure to support ingestion, processing, and storage of large volumes of scanned document images across enterprise-wide workflows
  • Architect end-to-end data solutions on AWS cloud services to enable seamless flow of scanned images from source systems through OCR processing, model inference, and downstream data extraction and categorization pipelines
  • Develop robust image preprocessing and OCR integration pipelines that handle TIF/PNG format conversion, normalization, resolution enhancement, noise reduction, and batching to prepare scanned documents for downstream computer vision and OCR models
  • Build and optimize data pipelines that integrate OCR engine outputs, extracting structured text and metadata from scanned images and routing them into databases and analytics platforms for further processing
  • Design and manage data storage architectures and containerized deployments, using Oracle databases and AWS-native stores (S3, EFS) to efficiently catalog, index, and retrieve extracted text, classification labels, and metadata from processed document images
  • Drive the adoption of containerized deployment strategies using AWS EKS (Elastic Kubernetes Service) to deploy and scale image processing microservices, OCR engines, and data pipeline components with high availability and fault tolerance
  • Collaborate closely with data scientists and ML engineers to ensure training datasets for different models, and other computer vision models are properly curated, versioned, labeled, and accessible through well-structured data pipelines
  • Evaluate and integrate emerging data technologies and tools to continuously improve pipeline throughput, reduce processing latency for high-volume document scanning workloads, and optimize cost efficiency
  • Establish and enforce data quality, lineage, governance, and security frameworks to ensure traceability and integrity of extracted data from scanned documents throughout the entire processing lifecycle
  • Partner with security and compliance teams to ensure that scanned document data, extracted PII/PHI, and sensitive content are handled in accordance with regulatory requirements, encryption standards, and access controls
  • Lead and mentor a team of data engineers, establishing coding standards, peer review processes, CI/CD workflows, and best practices for building production-grade image and document processing pipelines
Required Qualifications
  • Formal training or certification on Data Engineering concepts and 5+ years applied experience
  • Strong proficiency in Java, Groovy, and Python for building data pipelines, image preprocessing workflows, automation scripts, and backend data services
  • Hands‑on experience with image file handling, particularly TIF/PNG format processing, multi-page document splitting, format conversion, and integration with OCR and computer vision pipelines
  • Deep hands‑on experience with AWS cloud services including S3 (for image storage), Lambda, Step Functions, and CloudWatch for building and monitoring scalable data workflows
  • Expertise in AWS EKS (Elastic Kubernetes Service) for deploying and managing containerized image processing, OCR, and data pipeline services using Docker and Kubernetes
  • Advanced knowledge of Oracle databases including PL/SQL, performance tuning, partitioning strategies, and data modeling for storing and querying large volumes of extracted document data and classification results
  • Familiarity with OCR technologies and the ability to build data pipelines that consume and structure OCR output for downstream analytics and model training
  • Understanding of data requirements for training deep learning models including dataset preparation, annotation management, and feature store integration
  • Experience with CI/CD pipelines (Jenkins) and infrastructure‑as‑code tools (Terraform, CloudFormation) for automated deployment and environment management
  • Strong understanding of data governance, data quality frameworks, metadata management, and data cataloging, particularly in the context of document‑centric and image‑heavy data ecosystems
  • Excellent leadership, communication, and stakeholder management skills with the ability to drive technical decisions across cross‑functional teams
Preferred Qualifications
  • Domain expertise in the healthcare industry
Benefits

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission‑based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on‑site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

Equal Opportunity Employment

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Lead Data Engineer - Automation/Image Processing
AI/ML Lead Data Engineer - Automation/Image Processing

Fairygodboss • Tampa (FL)

On-site
USD 120,000 - 160,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
+1
AI/ML Lead Data Engineer - Automation/Image Processing
AI/ML Lead Data Engineer - Automation/Image Processing

Next Frontier Capital • Tampa (FL)

On-site
USD 140,000 - 190,000
Health insurance
On-site health and wellness centers
Retirement savings plan
AI/ML Lead Data Engineer - Automation/Image Processing
AI/ML Lead Data Engineer - Automation/Image Processing

JPMorgan Chase & Co. • Tampa (FL)

On-site
USD 100,000 - 130,000
Applied AI/ML Lead
Applied AI/ML Lead

Next Frontier Capital • Tampa (FL)

On-site
USD 180,000 - 280,000
Applied AI/ML Lead
Applied AI/ML Lead

JPMorganChase • Tampa (FL)

On-site
USD 130,000 - 150,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
Data Scientist Lead
Data Scientist Lead

Next Frontier Capital • Tampa (FL)

On-site
USD 140,000 - 200,000
Competitive compensation
Health benefits
On-site wellness centers
+1
Data Scientist Lead
Data Scientist Lead

JPMorganChase • Tampa (FL)

On-site
USD 140,000 - 190,000
Machine Learning Engineer - Document Digitization (LLMs)-Vice President
Machine Learning Engineer - Document Digitization (LLMs)-Vice President

TwinThread • Jersey City (NJ)

On-site
USD 120,000 - 160,000
Comprehensive health care coverage
Tuition reimbursement
Mental health support
+1
Machine Learning Engineer - Document Digitization (LLMs)-Vice President
Machine Learning Engineer - Document Digitization (LLMs)-Vice President

Fairygodboss • Jersey City (NJ)

On-site
USD 164,350 - 260,000
Comprehensive health care coverage
On-site health and wellness centers
Tuition reimbursement
+1
Applied AI/ML Lead - Payments
Applied AI/ML Lead - Payments

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 260,000
Health care coverage
Retirement savings plan
On-site wellness centers
+2