Senior Lead Software Engineer- AI/ML Platform

JPMorgan Chase & Co.

Wilmington (DE)

On-site

USD 180,000 - 250,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase & Co. is seeking a Senior Lead Software Engineer to design, build, and operate foundational cloud infrastructure enabling data scientists to develop, train, and deploy AI/ML solutions across the firm.

You will lead platform reliability, scalability, and automation, collaborating with cross-functional teams to solve complex infra challenges and accelerate production-grade AI capabilities with measurable business impact.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Experience delivering secure, production-quality code in Python or Java.
  • Strong foundations in distributed systems, microservices, and platform design.
  • Proven ability to architect and operate cloud-native infrastructure on AWS and other clouds.
  • Expertise with infrastructure-as-code tooling, specifically Terraform.
  • Hands-on with Docker and Kubernetes, including AWS EKS.
  • Experience building production AI/ML platforms (training, deployment, serving).
  • Strong DevOps practices: CI/CD, release automation, observability.
  • Experience with SQL/NoSQL and Linux scripting.
  • Experience leading adoption of AI-assisted development tools with security considerations.
  • Understanding of responsible AI use and governance.

Responsibilities

  • Builds and maintains reusable AI/ML platform infrastructure and shared services to support development, deployment, and operations at scale.
  • Architects, deploys, and operates secure cloud and container-based environments for training and inference, including GPU-intensive workloads.
  • Design and implement platform tooling, automation, and infrastructure-as-code solutions to streamline model deployment, environment provisioning, release management, and operational support.
  • Develops and maintains production-grade services, APIs, SDK integrations, and workflows that support model training, serving, evaluation pipelines, and AI application lifecycle management.
  • Partners with data science, ML engineering, and application teams to translate model and compute requirements into platform standards and deployment patterns.
  • Optimizes platform reliability, scalability, latency, and cost through orchestration, scheduling, and hardware acceleration.
  • Establishes operational best practices including monitoring, logging, observability, access controls, incident response, and production troubleshooting.
  • Supports enterprise LLM operationalization, including fine-tuning workflows, inference optimization, and evaluation; contribute to documentation and engineering standards.
  • D drives adoption and governance of approved AI-assisted engineering practices across teams to improve code quality, delivery speed, and operational outcomes, while establishing measurable validation standards and promoting reuse.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including approved AI-assisted development and automation capabilities, to improve the value realized by automation at scale.

Skills

Python/Java
Distributed systems
Platform architecture
DevOps
AI/ML platforms

Education

5+ years software engineering experience

Tools

Terraform
Docker
Kubernetes
AWS EKS
CI/CD tooling

Job description

Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products.

As a Senior Lead Software Engineer at JPMorgan Chase within Corporate - AIML Data Platforms team , you will design, build, and operate the foundational cloud infrastructure that enables data scientists and machine learning engineers to develop, train, and deploy intelligent solutions across the firm. In this role you will serve as a technical leader, driving platform reliability, scalability, and automation while collaborating with cross-functional teams to solve complex infrastructure challenges. Your work will directly accelerate the firm’s AI/ML capabilities - enabling faster experimentation and production-grade deployments that create measurable business impact.

Job Responsibilities
  • Builds and maintains reusable AI/ML platform infrastructure and shared services to support development, deployment, and operations at scale.
  • Architects, deploys, and operates secure cloud and container-based environments for training and inference, including GPU-intensive workloads.
  • Design and implement platform tooling, automation, and infrastructure-as-code solutions to streamline model deployment, environment provisioning, release management, and operational support.
  • Develops and maintains production-grade services, APIs, SDK integrations, and workflows that support model training, serving, evaluation pipelines, and AI application lifecycle management.
  • Partners with data science, ML engineering, and application teams to translate model and compute requirements into platform standards and deployment patterns.
  • Optimizes platform reliability, scalability, latency, and cost through orchestration, scheduling, and hardware acceleration.
  • Establishes operational best practices including monitoring, logging, observability, access controls, incident response, and production troubleshooting.
  • Supports enterprise LLM operationalization, including fine-tuning workflows, inference optimization, and evaluation; contribute to documentation and engineering standards.
  • Drives adoption and governance of approved AI-assisted engineering practices across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within the SDLC/TLM toolchain.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including approved AI-assisted development and automation capabilities, to improve the value realized by automation at scale.
Required Qualifications, Capabilities, and Skills
  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Experience delivering secure, production-quality code in Python or Java.
  • Strong foundations in distributed systems, microservices, and platform architecture/design principles.
  • Proven ability to architect and operate cloud-native infrastructure on AWS (compute, networking, storage, security) and other major clouds.
  • Demonstrated expertise with infrastructure-as-code tooling, specifically Terraform, in large-scale cloud environments.
  • Hands-on experience with Docker and Kubernetes, including AWS EKS operations.
  • Experience building or supporting production AI/ML platforms (training, deployment, and model serving/inference), including GPU infrastructure/tooling.
  • Strong DevOps/platform engineering practices: CI/CD, release automation, automated testing, and observability (monitoring/logging/tracing).
  • Experience with SQL/NoSQL databases and data integration; strong Linux, scripting, and networking fundamentals.
  • Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching senior engineers/leads on compliant usage patterns and controls.
Preferred Qualifications, Capabilities, and Skills
  • Proficiency in Go or Python for automation, tooling development, or platform service implementation.
  • Experience with MLOps frameworks and tools such as Kubeflow, MLflow, or similar AI/ML lifecycle management platforms.
  • Working knowledge of ML frameworks (PyTorch, TensorFlow, Hugging Face, scikit-learn) for model integration and operationalization.
  • Exposure to multi-cloud or hybrid cloud architectures and platform portability strategies.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Lead Software Engineer - AI/ML
Sr Lead Software Engineer - AI/ML

JPMorgan Chase & Co. • New York (NY)

On-site
USD 130,000 - 160,000
Senior Lead Software Engineer, Cloud Platforms
Senior Lead Software Engineer, Cloud Platforms

慨正橡扯 • Jersey City (NJ)

On-site
USD 130,000 - 170,000
Sr Lead Software Engineer - AI/ML
Sr Lead Software Engineer - AI/ML

Fairygodboss • New York (NY)

On-site
USD 180,000 - 250,000
Lead Software Engineer, Cloud Platforms
Lead Software Engineer, Cloud Platforms

慨正橡扯 • Jersey City (NJ)

On-site
USD 120,000 - 150,000
Senior Lead Software Engineer- AI/ML Platform
Senior Lead Software Engineer- AI/ML Platform

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior Lead Software Engineer- AI Platform engineer
Senior Lead Software Engineer- AI Platform engineer

JPMorgan Chase & Co. • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior Lead Software Engineer-AI Foundation Services
Senior Lead Software Engineer-AI Foundation Services

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 180,000 - 240,000
Sr Lead Software Engineer - AWS - Lead AI/ML Platform Engineer
Sr Lead Software Engineer - AWS - Lead AI/ML Platform Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 200,000
Lead Software Engineer - AI Platform Reliability
Lead Software Engineer - AI Platform Reliability

JPMorgan Chase & Co. • Seattle (WA)

Hybrid
USD 180,000 - 260,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000