SRE/Production Support Lead

Morgan Stanley

Alpharetta (GA)

On-site

USD 125,000 - 175,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Morgan Stanley is seeking an experienced SRE/Production Support Lead (VP level) to oversee the production environment, ensure reliability of deployed software, and drive performance optimization. The role spans incident management, automation, and cross-team collaboration across wealth management platforms.

The ideal candidate has extensive enterprise experience, strong scripting and cloud skills, and a proactive, coaching mindset to mentor teammates while maintaining 24/7 service availability.

Qualifications

  • Ten or more years of experience developing and/or supporting enterprise applications.
  • Five or more years’ experience leading a small to medium team.
  • Hands-on professional who understands both code and infrastructure.
  • Experience with Agile and DevOps/SRE concepts.
  • Working knowledge of cloud technologies (AWS/GCP/Azure).
  • Strong Unix shell scripting skills.
  • Proficiency in at least one major language (Python/Java/C#).
  • Experience in the Financial Services domain is a plus.
  • Excellent communication and collaboration across teams and stakeholders.
  • Experience with on-call incident management.

Responsibilities

  • Proactively detect, troubleshoot, and resolve issues affecting production applications with cross-team coordination.
  • Maintain clear, concise, and timely communication during outages.
  • Ensure production environment stability and reliability of platforms.
  • Develop and revise policies to guide production deployment and standards.
  • Ensure Change Management guidelines are followed for production systems.
  • Handle data access requests and production data activities.
  • Collaborate with development teams to meet production standards.
  • Build and maintain a knowledge base to reduce dependency on external resources.
  • Perform DevOps/SRE roles in application support, platform stability, and resiliency.
  • Produce documentation, standards, training materials, and conduct knowledge sharing.

Skills

SRE Leadership
DevOps
Unix Shell
Cloud (AWS/GCP/Azure)
Incident Management
Automation
On-call
Communication
Programming
AI Tools

Education

BS/MS in Computer Science or Computer Engineering

Tools

Grafana
Prometheus
Splunk
Kibana

Job description

In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is an SRE/Production Support Lead position at Vice President level, which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime.

Morgan Stanley is an industry leader in financial services, known for mobilizing capital to help governments, corporations, institutions, and individuals around the world achieve their financial goals.

Interested in joining a team that’s eager to create, innovate and make an impact on the world? Read on.

The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user and systems support, answering hotline calls, monitoring systems alerts, and taking corrective action. Technical understanding is important as well as the ability to speak to users and understand their problems. In addition to direct user support tasks, the team performs infrastructure related tasks including process configuration, hardware capacity planning, event management, release work, and support tool development to ensure any repetitive tasks are packaged to remove any element of risk.

This role will be responsible for overall stability of the Wealth Management Investment Management application platforms, participation in key optimization initiatives, and collaboration with multiple technical teams within Morgan Stanley. Partner with WM business units, various levels of management and staff to collect, analyze and make recommendations on optimizing the platform. As a team member with expertise in deep analytical triage, you will provide subject matter expertise in debugging, issue analysis and troubleshooting, working with business and technical colleagues to provide reviews and recommendations to avoid any future application issues.

What You’ll Do In The Role

  • Proactively detecting, troubleshooting, and resolving all issues affecting production applications. This involves coordination with and escalation to development and external teams where necessary, this team owns all issues escalated to us until it is resolved or a workaround is provided for end-user to continue functioning.
  • Responsible for maintaining clear, concise, and timely communication with affected parties during the investigation and resolution of any individual or system-wide outage.
  • Responsible for the stability of the Production environment.
  • Develop and continually revise (in partnerships with other teams where necessary) suitable policies and procedures to ensure appropriate application development standards are available to guide development for systems deployed to Production.
  • As the gatekeepers of the Production environment, responsible for ensuring the Change Implementation Management guidelines/policies are adhered to for all systems deployed to Production.
  • Responsible for servicing all requests for data or other activities that require access to Production systems.
  • Work with development teams at the appropriate stages in application development to ensure any new systems or projects meet the Production standard.
  • Responsible for maintaining and growing a body of knowledge that is accessible to all team members. Ensure information regarding any support related activities or issues is available and easily accessible. The goal is to improve self-reliance and reduce dependency on the availability of development or external team resources for the initial troubleshooting and resolution of problems.
  • Perform DevOps/ SRE role in Application support, Platform Stability and Resiliency.
  • Produce guidance documentation, standards and procedures, products assessments, and training material including working with the various application and infrastructure support teams ensuring that they are documenting every single troubleshooting step in Morgan Stanley knowledge base system to resolve issues in a faster time frame.
  • Serve as a fully seasoned/proficient technical resource; provide technical knowledge in outage management and proactive solutions to improve the user experience.

What You’ll Bring To The Role

  • Minimum 10 years’ experience in developing and/or supporting Enterprise Applications
  • BS/MS or equivalent, preferably in quantitative discipline (Computer Science, Computer Engineering).
  • 5+ years’ experience in leading a small to medium team of alike skillset.
  • Experienced, technically hands-on professional that understands both code and infrastructure
  • Willingness to embrace Agile and DevOps/SRE concepts.
  • Working knowledge on any of the DevOps & observability tools (Grafana, Prometheus, Splunk, Kibana)
  • Solid analytical skills, problem determination, and resolution recovery processes
  • Ability to interface and cultivate excellent working relationships with technology teams, business analysts, and vendors
  • Strong Unix Shell scripting experience required.
  • Have administrative competence in at least one major programming language or platform (for example: Perl, Powershell, Python, Java or C#, .NET)
  • Experience in AWS/GCP/Azure Cloud technologies
  • Experience in web analytics tools (preferably Adobe Experience Cloud tools) is Plus
  • Should be a fast learner of technologies in a quick paced environment.
  • Have strong organizational skills and the ability to manage multiple tasks and high-pressure situations for outage handling, management, or resolution
  • Is driven to learn about new technologies, techniques and what it takes to be an integral member of this team
  • Hands-on experience administering large-scale, high-availability systems and the tools to monitor performance and availability
  • Excellent communication and writing skills specific to technical discussions across the management layers
  • Experience with incident “on call” and ability to respond to emergencies on a 24/7 basis
  • Assisting in the investigation and troubleshooting of production issues and playing an active role in mentoring/coaching/training and development of team members
  • 3+ years’ experience in Automating tasks
  • Hands-on with AI and implementation of AI tools for operational efficiency
  • Strong ownership mentality with a focus on customer satisfaction
  • Be able to manage an outage incident, coordinating user communications, and other teams to help resolve an incident.
  • Experience working with Financial Services area will be a plus

What You Can Expect From Morgan Stanley

At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.

To learn more about our offices across the globe, please copy and paste https://www.morganstanley.com/about-us/global-offices into your browser.

Expected base pay rates for the role will be between $125,000 and $175,000 per year at the commencement of employment. However, base pay if hired will be determined on an individualized basis and is only part of the total compensation package, which, depending on the position, may also include commission earnings, incentive compensation, discretionary bonuses, other short and long-term incentive packages, and other Morgan Stanley sponsored benefit programs.

Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.

Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.

For more information, please visit: https://www.morganstanley.com/people-opportunities/eeo.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE/Production Support Lead
SRE/Production Support Lead

Socket.dev • Alpharetta (GA)

On-site
USD 125,000 - 175,000
SRE/Production Support Lead
SRE/Production Support Lead

PowerToFly • Alpharetta (GA)

On-site
USD 125,000 - 175,000
SRE/Production Support Lead
SRE/Production Support Lead

Morgan-Stanley • Alpharetta (GA)

On-site
USD 125,000 - 175,000
Comprehensive benefits
Application Support
Application Support

PowerToFly • New York (NY)

On-site
USD 120,000 - 165,000
Comprehensive employee benefits
Opportunity for career advancement
Lead Site Reliability Engineer, Vice President
Lead Site Reliability Engineer, Vice President

Socket.dev • New York (NY)

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer, Vice President
Lead Site Reliability Engineer, Vice President

Morgan Stanley • New York (NY)

On-site
USD 150,000 - 190,000
Executive Director – Site Reliability Engineering – WM Technology
Executive Director – Site Reliability Engineering – WM Technology

PowerToFly • New York (NY)

On-site
USD 195,000 - 215,000
Comprehensive employee benefits
Opportunity for career advancement
Technical Financial Program Manager - Vice President
Technical Financial Program Manager - Vice President

Morgan Stanley • New York (NY)

On-site
USD 155,000 - 215,000
Technical Financial Program Manager - Vice President
Technical Financial Program Manager - Vice President

Socket.dev • New York (NY)

On-site
USD 155,000 - 215,000
Senior Full Stack Software Engineer - Vice President
Senior Full Stack Software Engineer - Vice President

Morgan Stanley • New York (NY)

On-site
USD 150,000 - 210,000