Stand out for this role — generate a tailored resume and cover letter in about a minute.
JPMorgan Chase & Co. seeks an experienced engineer to lead design and delivery of scalable ML-powered services. You will architect resilient microservices, integrate with enterprise systems, and drive AI-assisted practices across teams.
You will mentor developers, collaborate with architects and DBAs, productionalize ML models, and ensure secure, auditable delivery in a fast-paced environment.
This role spans architecture, hands-on delivery, and ML/AI enablement in production. Architect and implement resilient, highly scalable, fault-tolerant, low-latency services and drive target-state architecture.
Design and deploy services that integrate with enterprise systems; ensure functional, performance, scalability, security, governance, and auditability requirements are met.
Lead and mentor the development team in a high-pressured delivery environment; manage multiple deliverables across business groups and strengthen stakeholder relationships.
Collaborate with LOB users, SMEs, architects, DBAs, and system administrators to design solutions, manage enhancements, and resolve issues.
Build and mature capabilities that execute ML pipelines for fraud detection and risk assessment; support modeling teams in implementation and tooling.
Productionalize models built by data scientists, including validation readiness and quality controls prior to live usage.
Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
Design and own reusable ML platform components (e.g., feature-store patterns, delivery pipelines) and establish monitoring/alerting for performance, scalability, availability, and reliability.
Build agentic AI services to automate and enhance engineering and model-ops workflows (tool-using agents, orchestration, state management, and audit-ready traceability).
Define and implement guardrails and evaluation approaches for agentic AI in production (quality, safety, latency, and cost).
Formal training or certification on software engineering concepts and 5+ years applied experience; Hands-on practical experience delivering system design, application development, testing, and operational stability
The successful candidate demonstrates deep distributed-systems engineering expertise in Python/Java plus strong platform, delivery, and production-operability discipline; Recent hands-on software development experience in large-scale distributed systems, primarily Python and modern microservices.
Strong Python experience for AI/ML engineering, automation, model operationalization, and agentic AI services, including tool integration, monitoring, telemetry, and governance; Strong experience with REST APIs and service-oriented / microservices architecture.
Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices.
Experience developing in Linux environments.
Strong Kubernetes orchestration experience (building, deploying, and operating production services).
Messaging expertise with Kafka, MQ, or similar platforms.
Experience with backend infrastructure patterns (e.g., load balancing, autoscaling; Experience with NoSQL databases such as Cassandra; Experience with log analytics / observability tools (e.g., ELK, Splunk).
Strong SDLC knowledge and agile ways of working, including CI/CD, application resiliency, security, testing, and operational stability; Strong communication skills and proven ability to influence across senior technology and business stakeholders.
AI/ML platform exposure (MLOps, feature engineering, model hosting/operationalization; AWS and/or hybrid on-prem + cloud); Agentic AI experience: building and operating LLM-driven agents with tool integration, monitoring/telemetry, and governance/audit considerations.
AWS Certification(s) - and/or Working Knowledge
AI Certifications(s) - and/or Working Knowledge