Overview
GENERAL DUTIES This position encompasses professional and responsible highly technical consultative and/or administrative work unique to the design and engineering of complex datacenters, networks and cloud deployments. Under administrative direction of a university IT manager, with broad latitude for independent action or decision, is responsible for the development of comprehensive infrastructure solutions. Engineer and operate mission critical infrastructure and be responsible for the access and protection of sensitive data, systems, and networks. May lead a team or participate in a team in planning, designing, implementing, and maintaining datacenter and cloud operations that includes the management of IT facilities, network, security, telecommunications, servers, storage, and virtualization components.
Assignment Levels There are three Levels of this title. Assignment Levels 2 and 3 may oversee multiple projects and supervise employees. This specification describes typical assignments for this title; related duties may be assigned as needed. To view the complete job description, go to http://www.cuny.edu/about/administration/offices/hr/classified-civil-service/ccsjobs/ and view the Job Description for IT Infrastructure Engineer.
Contract / Campus Information
CONTRACT TITLE IT Infrastructure Engineer
FLSA Non-exempt
CAMPUS SPECIFIC INFORMATION The Office of Computing and Information Services (CIS) at the City University of New York (CUNY) supports the IT and telecommunications needs of CUNY's 26 colleges. CIS supports enterprise IT and applications, identifies and develops new technologies that advance the University's core mission, operates and maintains the University's network, the enterprise Data Center, and the CUNY Service Desk, and manages the processes that safeguard the University’s IT assets. Additionally, CIS maintains the security posture by operating the CUNY Security Operations Center (SOC), develops disaster recovery plans for business continuity, and supports the CUNYfirst ERP solution across CUNY’s 26 colleges. The Infrastructure Operations Engineer is responsible for the engineering, administration, optimization, and operational management of the University's enterprise observability and monitoring platforms, with primary responsibility for Datadog observability services and SolarWinds network monitoring solutions. The role will support the design, implementation, monitoring, and continuous improvement of observability and network monitoring capabilities across on-premises, hybrid, and cloud environments. Reporting to the Director of Data Center Operations, the Infrastructure Operations Engineer will partner with infrastructure, cloud, networking, cybersecurity, and application teams to provide end-to-end visibility into enterprise systems and services. The position will be responsible for developing monitoring strategies, engineering dashboards and alerts, analyzing operational data, supporting incident response, and improving service reliability through proactive observability practices.
Key Responsibilities
- Engineers, administers, and optimizes the University's enterprise observability and monitoring platforms, including Datadog for infrastructure, application, cloud, and log observability, and SolarWinds for enterprise network monitoring, alerting, performance management, and operational visibility.
- Administers and maintains SolarWinds monitoring solutions, including Network Performance Monitor (NPM), Network Configuration Manager (NCM), and related modules to support enterprise network operations, performance monitoring, and configuration management.
- Designs, develops, and maintains enterprise monitoring dashboards, service maps, alerting strategies, and operational reporting to support infrastructure and application teams.
- Configures and manages observability integrations across servers, cloud platforms, databases, network devices, applications, and enterprise services.
- Analyzes telemetry, metrics, logs, traces, and events to identify performance bottlenecks, operational anomalies, and service degradation.
- Supports incident response activities through observability-driven troubleshooting, root cause analysis, and operational analytics.
- Develops monitoring standards, alert tuning methodologies, and operational best practices to reduce alert fatigue and improve service reliability.
- Partners with infrastructure, cloud, networking, cybersecurity, and application teams to establish observability requirements for new systems and enterprise initiatives.
- Supports monitoring of Microsoft Azure, AWS, virtualization platforms, storage systems, network infrastructure, and enterprise applications.
- Develops automated monitoring solutions using APIs, scripting, and infrastructure automation tools.
- Produces operational health reports, capacity trends, service-level metrics, and executive dashboards.
- Supports enterprise initiatives related to AIOps, predictive analytics, event correlation, and service reliability engineering practices.
- Maintains observability platform documentation, operational procedures, monitoring standards, and knowledge articles.
- Participates in incident, problem, and change management processes and provide on-call support as required.
- Performs special projects, ad hoc assignments, and related duties consistent with the scope and responsibilities of the position, which may require travel to different work locations and participation in on-call support outside of standard business hours, including evenings and weekends, as needed.
Notes
- Until further notice, this position is eligible for a hybrid work schedule.
- An appointment to this Competitive title/position in the Classified Civil Service Title Series will be made with a Provisional status. Employees in provisional status must pass a competitive civil service examination and be appointed from a civil service list to remain in the title/position.
Minimum Qualifications
- A baccalaureate degree in computer science, engineering or a related field from an accredited college or university and five (5) years of satisfactory full-time experience in the design and engineering of complex datacenters, networks and/or cloud deployments; or
- A baccalaureate degree from an accredited college or university and six (6) years of satisfactory full-time experience as described in “1” above; or
- A high school diploma or its educational equivalent and ten (10) years of satisfactory full-time experience as described in “1” above; or
- Education and/or experience which is equivalent to the above. The following may substitute for some of the required experience: College Education (undergraduate Credits) may substitute for up to four years of the required experience; Graduate Credits may substitute for up to two years; Certifications may substitute for one year each. (Details in the original job description.)
Preferred Qualifications
- 5+ years of professional experience supporting enterprise IT infrastructure environments including servers, storage, networking, virtualization, and cloud platforms.
- 3+ years of experience supporting enterprise monitoring, alerting, and operational management tools in complex IT environments.
- Experience engineering and administering enterprise observability platforms including Datadog, Dynatrace, SolarWinds, Splunk Observability, or similar solutions, with demonstrated experience supporting infrastructure monitoring, network performance monitoring, application observability, alerting, and operational analytics.
- Experience developing Datadog dashboards, monitors, synthetic testing, APM, log pipelines, and cloud integrations; experience developing and maintaining SolarWinds dashboards, alerts, reports, network discovery, configuration management, and performance monitoring solutions in large enterprise environments.
- Strong understanding of infrastructure monitoring, application performance monitoring, observability principles, telemetry, metrics, logging, tracing, and event correlation.
- Experience supporting Microsoft Azure, AWS, hybrid cloud environments, and enterprise infrastructure platforms.
- Familiarity with AIOps, operational analytics, service reliability engineering (SRE), and automation practices.
- Experience with PowerShell, Python, APIs, Terraform, and other automation technologies.
- Ability to work productively both independently and collaboratively within cross-functional teams supporting enterprise IT operations.
- Strong communication and interpersonal skills with the ability to effectively collaborate with technical and non-technical stakeholders across the organization.
Compensation & Benefits
- Level 1 (Incumbent Minimum) - $105,827
- Benefits: CUNY offers a comprehensive benefits package to employees and eligible dependents based on job title and classification. Pension and Tax-Deferred Savings Plans are available. Part-time eligibility details apply.
How to Apply
For full consideration, submit a cover letter and resume online via CUNY's web-based job system, addressing how your experience and credentials meet the responsibilities and qualifications outlined. The direct link to the job opening from external sources is: https://hrsa.cunyfirst.cuny.edu/psc/erecruit/EMPLOYEE/HRMSCG/c/HRS_HRAM_FL.HRS_CG_SEARCH_FL.GBL?Page=HRS_APP_JBPST_FL&Action=U&FOCUS=Applicant&SiteId=1&JobOpeningId=32393=5&PostingSeq=1 Current CUNY employees must apply through CUNYfirst Employee Self Service using their login credentials. After you login, click the Careers tile on the Employee Self Service Menu page to view job openings.
Closing Date: Open until filled.
Equal Employment Opportunity
CUNY encourages people with disabilities, minorities, veterans and women to apply. EEO/AA/Vet/Disability Employer. Job ID: 32395; Location: Central Office; Job Type: Full-Time