Site Reliability Engineer

PowerToFly

Alpharetta (GA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A global financial services firm is seeking a Site Reliability Engineer for their Alpharetta, GA office. This Director-level role is responsible for ensuring the operational reliability of deployed software, minimizing downtime, and automating deployments. The ideal candidate will have over 5 years of experience, strong skills in scripting (Python, Shell), and expertise in database technologies. Excellent communication and leadership skills are essential. This position demands high-level problem-solving in a fast-paced environment.

Qualifications

  • Minimum of 5 years of experience in a production environment with a solid software development background.
  • Strong experience in handling production issues in a pressured environment.
  • Hands-on experience in designing, developing, and implementing technical solutions.

Responsibilities

  • Maintain live applications and ensure system health.
  • Engage in the lifecycle of services from inception to deployment.
  • Troubleshoot infrastructure issues and maintain a knowledge base.

Skills

Scripting languages (Shell, Python, Perl)
Database skills with DB2, Sybase, Oracle
Cloud-driven development
Performance tuning
Agile Methodology
Excellent communication skills

Education

BS degree in Computer Science or Engineering

Tools

Jenkins
Autosys
Monitoring tools (Splunk, IP Soft)

Job description

In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities.

This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime.

Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.

Services Technology is a division within Wealth Management Technology that enhances technology solutions to improve business processes and service delivery. The division leverages a range of tools and technologies to automate processes, increase efficiency, and improve the effectiveness of business services. The Services Technology organization delivers platforms that support core client and advisor experiences. Our teams build and maintain solutions for the Contact Center, digital business automation, workflow orchestration, CRM & Salesforce, and client reporting.

Job Summary

We are looking for a Site Reliability Engineer with a minimum of 5 years of industry experience, preferably working in the financial IT community. This role focuses on delivering exceptional services to both BU and Dev partners to minimize or avoid production outages. The position will concentrate on production support within the WM Product Technology team, automating deployments and working with agile teams to build and maintain stable, reliable production systems. The ideal candidate will be passionate about automation and skilled in one of the following programming languages: Python, PERL, SHELL, Ruby, JAVA, C#, or the like. Candidates should possess a strong understanding of database concepts, job scheduling, MQ, Web services, UNIX/LINUX/Windows OS, and experience debugging applications. Leadership, excellent communication, and a commitment to continuous improvement are essential.

Responsibilities
  • Maintain live applications by measuring and monitoring availability, latency, and overall system health, continuously evaluating cost and TOIL.
  • Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation, capacity planning, and launch reviews.
  • Scale systems sustainably through automation and evolve systems by implementing changes that improve reliability and velocity.
  • Troubleshoot infrastructure issues, review log files, update documentation, and maintain a knowledge base with resolutions.
  • Work closely with the application Development team to understand the platform and create tools/utilities to aid production management.
  • Collaborate with upstream data providers and consumers, reducing escalations to development teams.
  • Develop scripts and assist with code changes along with operational tasks and activities.
  • Ensure that the support team has comprehensive knowledge of the application set, maintaining support knowledgebases and documents.
  • Use analytical skills to identify trends in the environment and drive problem resolution.
  • Lead efforts to identify improvement areas to stabilize the production plant.
  • Identify risks and work with urgency, both independently and within a team.
  • Test and tune network, hardware, and software configurations to maximize performance.
  • Interface with IT Dev managers, Infrastructure teams, and serve as a Subject Matter Expert for supported applications.
  • Understand the overall business flow of supported application systems and their interface with clients.
  • Own and manage production requests, questions, issues, and conduct Root Cause Analysis for outages/incidents.
  • Provide weekend on-call rotation and availability for offshore time lead.
  • Account for Production and non-Production environments as part of 24/7 support coverage.
Skills Required
  • 5+ years of experience in a production environment with a solid software development background and a focus on performance tuning, end-to-end troubleshooting, networking fundamentals, and attention to detail.
  • Ability to resolve production issues in a high-demanding and pressured environment.
  • Hands‑on experience designing, developing, and implementing technical solutions or deep technical support.
  • Strong experience in scripting languages (Shell, Python, Perl, etc.) and cloud-driven development.
  • Strong database skills with DB2, Sybase, or Oracle.
  • Hands‑on experience with Autosys or other batch scheduling software.
  • Experience with Continuous Integration and Continuous Deployment.
  • Experience provisioning environments on demand for both Virtual Machines and containers.
  • Knowledge and hands‑on experience with monitoring tools such as Splunk, IP Soft, Spark, or Sockeye.
  • Experience with Agile Methodology (e.g., Scrum).
  • Experience automating deployments using Jenkins and Train.
  • Ability to diagnose technical problems, debug, optimize code, and automate routine tasks.
  • Hands‑on experience in application and database troubleshooting/issue resolution in a fast-paced environment.
  • Excellent communication and creative problem-solving skills.
  • Knowledge of cloud-based deployment, security, and networking concepts in Azure and AWS.
  • Hands‑on experience leveraging generative AI tools to enhance research, automation, and productivity.
Skills Desired
  • Knowledge of algorithms, data structures, complexity analysis, and software design.
  • Interest in designing, analyzing, and troubleshooting large-scale distributed systems.
Educational Qualification
  • Minimum BS degree in Computer Science, Engineering, or a related field.
Equal Employment Opportunity

Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet). The firm ensures equal employment opportunity without discrimination on the basis of race, color, religion, creed, age, sex, gender identity, sexual orientation, national origin, citizenship, disability, marital status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer, Vice President
Lead Site Reliability Engineer, Vice President

Morgan Stanley • New York (NY)

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer, Vice President
Lead Site Reliability Engineer, Vice President

Socket.dev • New York (NY)

On-site
USD 150,000 - 190,000
SRE/Production Support Lead
SRE/Production Support Lead

PowerToFly • Alpharetta (GA)

On-site
USD 125,000 - 175,000
SRE/Production Support Lead
SRE/Production Support Lead

Socket.dev • Alpharetta (GA)

On-site
USD 125,000 - 175,000
Vice President - Production Support Manager
Vice President - Production Support Manager

Morgan Stanley • South Jordan (UT)

On-site
USD 100,000 - 130,000
Executive Director – Site Reliability Engineering – WM Technology
Executive Director – Site Reliability Engineering – WM Technology

PowerToFly • New York (NY)

On-site
USD 195,000 - 215,000
Comprehensive employee benefits
Opportunity for career advancement
SRE/Production Support Lead
SRE/Production Support Lead

Morgan Stanley • Alpharetta (GA)

On-site
USD 125,000 - 175,000
Application Support
Application Support

PowerToFly • New York (NY)

On-site
USD 120,000 - 165,000
Comprehensive employee benefits
Opportunity for career advancement
Site Reliability Engineer III- Production Management
Site Reliability Engineer III- Production Management

慨正橡扯 • New York (NY)

On-site
USD 140,000 - 210,000
Site Reliability Engineer III- Production Management
Site Reliability Engineer III- Production Management

JPMorganChase • New York (NY)

On-site
USD 140,000 - 190,000