Software Engineer III - Machine Learning Platform

J.P. Morgan

New York (NY)

On-site

USD 140,000 - 200,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorganChase is seeking a Software Engineer III to join the AI/ML data platform team in New York. You will build and operate scalable ML training systems on AWS and other clouds, optimize GPU workloads, and enable repeatable, governed training across environments.

You will leverage Kubernetes, CI/CD, and enterprise AI tooling to improve code quality, delivery speed, and observability, while partnering with data engineering and platform teams to define interfaces and guardrails.

Qualifications

  • Formal training or certification in software engineering concepts and 3+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (e.g., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (e.g., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2).
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment with demonstrated ability to evaluate outputs for correctness, performance, and security.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed), improving throughput and reproducibility.
  • Build and operate training infrastructure on Kubernetes (EKS and other platforms).
  • Enable Gen AI/LLM training workflows with evaluation harnesses, governance, and scalable GPU execution patterns.
  • Implement observability for training systems: metrics, logs, dashboards, alerting.
  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls).
  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation.
  • Leverage enterprise AI coding assist tools to improve code quality and productivity, with peer review and secure coding.
  • Apply SDLC toolchain knowledge to improve automation.

Skills

Python
ML deployment
Performance optimization
Observability
Code reviews
Testing
Security & governance

Tools

AWS
Kubernetes
Docker
EKS
PyTorch
TensorFlow
CI/CD for ML
Git

Job description

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level.

As a Software Engineer III at JPMorganChase within the AI/ML data platform team you serve as a seasoned member of an agile team to build and operate scalable, reliable ML training systems and pipelines on AWS and other cloud platforms. You will productionize training workloads (often GPU-based), improve performance and cost efficiency, and enable repeatable, well-governed training across environments

Job Responsibilities
  • Design, build, and maintain end-to-end ML training platform.

  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.

  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.

  • Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.

  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.

  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls)

  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow

  • Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness.

  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Required qualifications, capabilities, and skills
  • Formal training or certification on software engineering concepts and 3+ years applied experience

  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.

  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).

  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).

  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).

  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.

  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).

  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).

  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)

  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.

  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.

Preferred qualifications, capabilities, and skills
  • Experience running training workloads across multiple cloud platforms and managing portability, performance, and governance across environments.

  • Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management.

  • Experience optimizing training input pipelines (sharding, prefetching, caching, format choices such as Parquet/WebDataset) and working with large datasets.

  • Familiarity with distributed compute frameworks (Spark, Ray, Dask) for feature/dataset generation.

  • Familiarity with workflow orchestration tools (Airflow-like systems, Argo Workflows-like patterns) and model registry concepts.

  • Experience optimizing training cost/performance (right-sizing, scheduling policies, interruptible capacity strategies where applicable budge guardrails, quota planning).

  • Strong observability practice for training systems: metrics/logs/traces, GPU telemetry, dashboards, and alert tuning.

FEDERAL DEPOSIT INSURANCE ACT: This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries.

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans

Our professionals in our Corporate Functions cover a diverse range from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success. Design and deliver market-leading technology products in a secure and scalable way as a seasoned member of an agile team

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer III - Machine Learning Platform
Software Engineer III - Machine Learning Platform

Socket.dev • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Software Engineer III - Machine Learning Platform
Software Engineer III - Machine Learning Platform

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Software Engineer III - Machine Learning Platform
Software Engineer III - Machine Learning Platform

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Health care coverage
On-site health centers
Retirement savings plan
+2
Software Engineer III - Machine Learning Platform
Software Engineer III - Machine Learning Platform

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Software Engineer III - AI/ML Platform Engineer
Software Engineer III - AI/ML Platform Engineer

Worky • Jersey City (NJ)

On-site
USD 110,000 - 160,000
Senior Lead Software Engineer- AI Platform engineer
Senior Lead Software Engineer- AI Platform engineer

Next Frontier Capital • United States

On-site
USD 120,000 - 160,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
Senior Lead Software Engineer- AI/ML Platform
Senior Lead Software Engineer- AI/ML Platform

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000