As a NOC Engineer, you will help keep the infrastructure our clients depend on healthy, patched, secure, and performing. You will monitor systems across our client base, respond to alerts with urgency, and build automation that enables the team to operate more efficiently. This is a fully remote position. Patching and maintenance windows may occasionally fall outside of standard business hours.
Position Responsibilities
- Ensuring Maximum Uptime: Own and maintain system uptime targets by monitoring infrastructure health, responding to alerts, and resolving incidents before they impact users.
- Tools Deployment and Maintenance: Deploy, configure, and maintain monitoring and management tools to ensure visibility and control across infrastructure and endpoints.
- Infrastructure Maintenance: Maintain and optimize core infrastructure components to ensure reliability, performance, and scalability.
- Endpoint Patching: Execute and track endpoint patching schedules to maintain security compliance and system stability.
- Backup Management: Manage and monitor client backup platforms, verifying job success, remediating failures, and testing restores to ensure data is recoverable when it matters.
- Automation: Build and maintain automation scripts and workflows to reduce manual effort, improve consistency, and increase operational efficiency.
- Documentation: Keep monitoring configurations, runbooks, and maintenance procedures documented and current so the team can follow standardized processes.
- Escalation and Handoffs: Escalate complex incidents with full context and hand off open work cleanly so nothing stalls between teammates.
To excel in this position, candidates should possess the following attributes:
- Detail-oriented: Consistently does the little things right, from alert triage to patch verification, because small misses in a NOC can become big outages.
- Calm under pressure: Stays composed and methodical during alert storms and incidents, prioritizing clearly and communicating status effectively.
- Process-driven: Works comfortably within documented processes, follows runbooks, and improves them rather than working around them.
- Curious and improvement-minded: Looks for the recurring problem behind repeated alerts and explores how automation and AI-enabled tooling can reduce manual work.
- Client-centric: Empathetic and attentive to client needs, with a consistent focus on setting and meeting expectations and helping clients achieve their business objectives.
Skills, Experience and Qualifications
- Hands-on experience with RMM platforms such as Datto RMM.
- Strong understanding of Group Policy and Intune policy management for endpoint configuration and control.
- Experience administering backup platforms, including job monitoring, failure remediation, and restore testing.
- Working knowledge of monitoring tools, alert tuning, and incident triage.
- Administration of Microsoft 365 and Windows Server environments.
- Networking fundamentals, including TCP/IP, DNS, DHCP, VLANs, VPN, and firewalls.
- Scripting in PowerShell or Python to automate repetitive operational tasks.
- Clear written documentation habits for runbooks, configurations, and handoffs.
Experience
- Two or more years in a NOC, systems administration, or infrastructure support role.
- Experience in an MSP or other multi-client environment is strongly preferred.
- A track record of running patching cycles and managing monitoring or RMM tooling.
- Exposure to automation projects; interest in AI-enabled operations is a plus.
Qualifications
- Certifications such as CompTIA Network+ or Security+, Microsoft certifications, or RMM vendor certifications are advantageous.