Company Introduction
Availity is one of the leading health information networks in the United States, processing more than 4 billion transactions annually and connecting more than two million healthcare providers and over two thousand technology partners to health plans nationwide. Our teams of technology, business, and customer service professionals in Bangalore, India, work together to transform healthcare delivery in the United States through innovation and collaboration. Availity’s technologists help develop cutting‑edge revenue cycle solutions that help hospitals, health systems, and physicians maximize payments and optimize their workflows.
Job Description
The NOC Cloud Infrastructure Engineer is responsible for 24×7 operational monitoring, support, and incident response for cloud infrastructure with a strong emphasis on cloud networking. This role ensures availability, performance, and reliability of cloud platforms through proactive monitoring, structured incident management, and disciplined NOC operations. The engineer works closely with Senior NOC Engineers and the Operations Center (OC) to deliver continuous coverage and rapid restoration.
Roles & Responsibilities
- Must Have Skills: 24x7 NOC operations experience (shift‑based); incident triage & escalation (P1–P4); AWS and/or Azure operational experience; VPC/VNet, subnets, routing, load balancers; DNS and connectivity troubleshooting; Linux & Windows server administration; monitoring with CloudWatch, Azure Monitor, Splunk, Grafana; ServiceNow / ITSM incident handling.
- Good to Have Skills: Basic Kubernetes / EKS operations; Terraform familiarity; automation / scripting basics; cloud database awareness; ITIL Foundation.
Core Responsibilities
- Monitor cloud infrastructure across compute, storage, and networking layers.
- Respond to alerts and incidents in accordance with NOC SOPs and escalation matrices.
- Perform initial triage for cloud and network‑related incidents (P1–P4).
- Troubleshoot connectivity, routing, DNS, and load‑balancer related issues in cloud environments.
- Escalate complex or unresolved issues to Senior NOC Engineers or specialized teams.
- Support Major Incident Management (MIM) bridges with technical diagnostics and status updates.
- Maintain accurate incident records, timelines, and communications in ITSM tools.
- Participate in structured shift handovers to ensure follow‑the‑sun continuity.
Monitoring & Observability
- Continuously monitor cloud and network health using centralized dashboards.
- Detect availability, performance, latency, and connectivity issues before customer impact.
- Validate alerts, reduce noise, and follow defined alert response procedures.
- Contribute to improvements in dashboards, alerts, and runbooks.
Monitoring & Tooling (Required)
- AWS CloudWatch (metrics, logs, alarms)
- Azure Monitor and Log Analytics
- Splunk for log analysis and alert correlation
- Grafana and/or Prometheus for metrics visualization
- ServiceNow or Jira for alert‑to‑incident tracking
Technical Requirements
- Hands‑on operational experience with AWS and/or Azure.
- Proven experience with VPC/VNet design concepts, subnets, route tables, NAT/IGW, DNS, load balancers, VPN/peering, and security groups/NSGs.
- Ability to diagnose latency, packet loss, routing, firewall, and name‑resolution issues in cloud environments.
- Linux and Windows server administration in cloud environments.
- Understanding of cloud storage services and backup/restore concepts.
- Awareness of cloud databases (RDS / Aurora / Azure SQL) and common operational symptoms.
- Basic operational knowledge of Kubernetes / EKS workload health (Preferred).
- Familiarity with Terraform and scripted operational tasks (Preferred).
Required Experience
- 5–7 years of experience in NOC, Cloud Infrastructure, Network Operations, or SRE roles.
- Hands‑on experience supporting cloud networking in 24×7 operational environments.
- Experience with incident response, escalation, and shift‑based operations.
- CCNA or equivalent networking certification (preferred).
- AWS Certified Solutions Architect or Advanced Networking Specialty (preferred).
Behavioral & Operational Skills
Strong troubleshooting and analytical skills across cloud and networking domains. Clear communication during incidents and shift handovers. Ability to operate calmly under pressure in a 24×7 NOC environment. Team‑oriented mindset with a strong sense of operational ownership.
Preferred Certifications
- AWS Certified Solutions Architect or Advanced Networking Specialty
- Microsoft Azure Administrator or Azure Network Engineer Associate
- ITIL Foundation (preferred)
- CCNA or equivalent networking certification (preferred)
Shift Model
24×7 rotational shifts with structured handovers and escalation support.
Video Camera Usage
Availity fosters a collaborative and open culture where communication and engagement are central to our success. As a remote‑first company, we are also camera‑first and provide all associates with camera/video capability to simulate the office environment. Video participation is required for all virtual meetings to ensure security, prevent unauthorized access, and maintain focused collaboration.