We are helping our client find Network Operations Engineers to monitor and support the critical networks, cloud infrastructure, databases, and authentication systems that keep live autonomous vehicle operations running reliably.
In this role, you'll monitor critical infrastructure, triage alerts, respond to incidents, and escape complex issues to specialized engineering teams. You'll also partner closely with Site Reliability Engineering (SRE) to identify recurring issues and turn operational insights into long-term reliability improvements.
The ideal candidate has experience working in a NOC, TechOps, or similar mission-critical operations environment and is comfortable troubleshooting infrastructure issues in real time. You are highly responsive, technically curious, and able to communicate clearly and remain composed during high-pressure incidents.
As a Network Operations Engineer, you'll:
- Proactively monitor critical IT services, networks, and infrastructure supporting live autonomous vehicle operations in a 24/7 environment.
- Monitor the real-time health of cloud infrastructure, databases, DNS, authentication services, and other systems required for continuous operations.
- Acknowledge, investigate, and triage system alerts, serving as a primary responder for incidents affecting live services.
- Troubleshoot incidents and participate in escalation calls to support investigation, resolution, and root cause identification.
- Follow established operational procedures and escalation paths to route complex issues to the appropriate engineering teams.
- Partner with SRE and engineering teams to identify recurring issues and improve system reliability.
- Participate in incident retrospectives and recommend improvements that reduce outages and strengthen operational processes.
- Document incidents, troubleshooting activities, resolutions, and escalation details accurately.
- Prepare clear shift handoffs and service status reports to maintain continuity across 24/7 operations.
Required Qualifications:
- Experience: 5+ years in a NOC, SOC, TechOps, or similar structured operations environment, with a strong understanding of incident management lifecycles and SLAs.
- Infrastructure Monitoring: Proven experience monitoring enterprise networks, cloud infrastructure (AWS or GCP), and critical databases.
- Core Services: Strong understanding of TCP/IP, DNS, authentication services EntraID, OpenAuth, SSL/TLS and general networking.
- Tooling: Proficiency with modern monitoring and observability platforms (e.g., Grafana, Datadog, Kentik, Logic Monitor, or similar alert management systems).
- Communication: Excellent written and verbal communication skills, with the ability to convey critical technical issues clearly during high-pressure situations.
- Availability: Willing and able to work 100% on-site in Scottsdale, AZ on an assigned shift that includes at least one weekend day (Saturday or Sunday) every week, including holiday coverage as scheduled.
Preferred Qualifications:
- Experience with ticketing and incident management systems such as Incident.io, Pagerduty, Opsgenie, ServiceNow and JIRA
- Basic scripting for operational tasks (Python, Bash)
- ITIL or comparable incident/service management framework familiarity
- CompTIA Network+ and Security+ certifications
- Cisco Certified Network Associate (CCNA) or similar vendor-specific networking certifications (e.g., Juniper JNCIA)
Key Responsibilities:
- Provide proactive, 24x7 monitoring of the critical IT services, networks, and infrastructure supporting Robotaxi Operations from the Scottsdale operations center.
- Acknowledge, investigate, and triage system alerts.
- Act as the primary responder for incidents impacting live services.
- Follow established operational procedures and escalation pathways to route complex issues to the appropriate specialized engineering teams.
Service Health Tracking:
- Monitor the real‑time health of cloud environments, databases, DNS, and authentication services required for continuous operations.
Reliability Improvement:
- Participate in incident retrospective calls to provide feedback and recommendations to improve system reliability and reduce future outages.
Reporting:
- Document incidents thoroughly and produce shift handoff reports and service status summaries.
- Participate in escalation calls to aid in troubleshooting, resolution, and root‑causing of incidents.
To achieve true 24x7 coverage, we are staffing three shifts: Day, Swing, and Overnight. Candidates apply to the role generally; shift assignment is determined during the interview process based on business need and candidate availability. Please come prepared to discuss which shifts you can work. Every schedule includes at least one weekend day (Saturday or Sunday) each week. Holiday coverage is required as part of the 24x7 rotation.
- Onsite: Onsite in Scottsdale, AZ | 5 days in office
- Shifts: Day, Swing or Overnight; at least 1 weekend day