The Manager of Infrastructure and Security Operations is responsible for the availability, performance, and security of the server, storage, and endpoint infrastructure supporting company’s core processing platforms and the customers that rely on them. Reporting to the Director of Data Center Operations and Hardware Services, this position leads the infrastructure and security operations team, owns day-to-day operational execution across hosted and customer-facing environments, and partners closely with Network & Systems Operations, Software Engineering, and Client Support.
This is a hands-on leadership role in a regulated financial services environment. The Manager is accountable for infrastructure issues being identified, diagnosed, and resolved or properly escalated with supporting evidence, and for maintaining the documentation and monitoring that make that possible.
Responsibilities
Infrastructure Operations and Service Delivery
- Manage daily operation of physical and virtual server infrastructure, storage, backup, and endpoint systems supporting company products and hosted customer environments.
- Ensure availability, integrity, and performance of infrastructure platforms in accordance with established service level agreements.
- Oversee support of customer hardware, including workstations, peripherals, check scanners, and related endpoint devices.
- Ensure operational adherence to approved architectures, standards, and change control procedures.
Monitoring, Capacity, and Performance
- Build and maintain monitoring, alerting, and logging coverage across infrastructure so that degradation is detected internally before customer impact.
- Establish and maintain performance baselines by environment, site, and device class.
- Perform capacity planning and forecast infrastructure needs ahead of demand.
- Analyze performance trends and initiate corrective action on recurring conditions.
Incident, Problem, and Root Cause Management
- Lead infrastructure incident response, coordinating resources through resolution and post-incident review.
- Conduct root cause analysis on infrastructure incidents; document findings, evidence, and remediation.
- Ensure issues referred to other teams are accompanied by documented investigation, including observed symptoms, tests performed, results, and conditions ruled out.
- Track incident trends and drive problem management to eliminate recurring causes.
Hardware Lifecycle and Asset Management
- Maintain an accurate inventory of servers, firewalls, switches, storage, and customer-site hardware, including configuration, warranty, and lifecycle status.
- Plan and execute hardware installation, upgrade, replacement, and decommissioning activities with minimal business disruption.
- Develop business cases and cost estimates for infrastructure investment and refresh.
- Manage the infrastructure hardware budget in coordination with the Director.
Security Operations
- Oversee firewall administration, access control, VPN, endpoint protection, patch management, and vulnerability remediation.
- Maintain security monitoring and lead infrastructure security incident response.
- Support internal and external audits, penetration testing, and regulatory examinations, including remediation of findings.
- Ensure infrastructure practices align with applicable financial services regulatory guidance and information security policy.
Business Continuity and Disaster Recovery
- Maintain and test disaster recovery procedures for infrastructure platforms in coordination with risk and business continuity stakeholders.
- Document test results, gaps, and remediation plans.
- Ensure backup and recovery operations meet defined recovery objectives.
Documentation and Reporting
- Maintain current network and infrastructure diagrams, runbooks, and operational procedures.
- Report on infrastructure performance, incident volume, root cause findings, and operational risk to the Director and other stakeholders.
- Ensure change and configuration records are accurate and auditable.
Cross-Functional and Customer Support
- Serve as the infrastructure point of contact for Software Engineering, Client Support, and Network & Systems Operations.
- Engage directly with customer staff and their third-party IT providers to establish facts during performance and stability investigations.
- Coordinate with hardware, software, and service vendors on support, maintenance, and issue resolution; hold vendors to contracted service levels.
Team Leadership
- Supervise assigned staff, including assigning work, setting priorities, and ensuring adequate operational coverage.
- Establish and enforce team standards for diagnostics, documentation, and escalation.
- Coach and develop staff; conduct performance reviews and support career growth.
- Participate in hiring, training, and cross-training in coordination with Human Resources.
Qualifications
- Bachelor's degree in Computer Science, Information Technology, or a related field; equivalent additional experience may be substituted for the degree requirement.
- Five or more years of experience in IT infrastructure or systems operations.
- Two or more years of direct people management experience.
- Demonstrated experience operating production infrastructure supporting external customers under service level commitments.
- Working knowledge of server hardware, virtualization, storage, backup and recovery, and endpoint infrastructure.
- Working knowledge of network fundamentals, firewalls, VPN, and their interaction with application performance.
- Experience with infrastructure monitoring and alerting platforms, including establishing baselines and thresholds.
- Demonstrated experience conducting structured root cause analysis and documenting technical findings.
- Working knowledge of IT security practices, patch and vulnerability management, and change management discipline.
- Strong written communication skills, with the ability to document technical findings for both technical and non-technical audiences.
Additional Information
Preferred Qualifications
- Experience in a FinTech, bank, core processing provider, or other regulated financial services environment.
- Experience supporting endpoint and peripheral hardware in customer locations.
- Experience with structured process improvement methodology such as Six Sigma or DMAIC.
- Certifications such as CompTIA Security+ or Server+, CCNA, ITIL Foundation, VMware, or Microsoft infrastructure credentials.
Knowledge, Skills, and Abilities
- Investigative discipline: gathers evidence and verifies claims before acting or escalating.
- Sound analytical and problem-solving judgment under production pressure.
- Effective collaboration across technical teams with adjacent or overlapping scope.
- Professional customer-facing communication with customer staff and third-party providers.
- Ability to manage competing operational priorities and multiple concurrent projects.
- Ability to maintain confidentiality of member and customer data.