N Software Engineer NetApp, Durham, North Carolina, US

Artha Nexgen

Durham, Northern (NC, KY)

Hybrid

USD 148,000 - 220,000

Full time

24 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Comprehensive benefits package

Job summary

Artha Nexgen is seeking a Cloud Infrastructure / Site Reliability Engineer to bridge development and operations. You will design, deploy, monitor, and automate cloud services in a Kubernetes-based environment, supporting SaaS/IaaS offerings.

The role includes on-call rotations and working with cloud teams across RTP and other hubs, demanding a proactive, self‑starter mindset. You will lead automation, troubleshoot complex stack issues, and drive reliability improvements while collaborating with

Qualifications

  • Bachelor's degree in Computer Science or equivalent experience.
  • 5+ years in scripting and infrastructure automation.
  • Deep knowledge of containers, Kubernetes, and distributed systems.
  • Proficiency in Linux/Unix environments; CoreOS knowledge is a plus.
  • Experience with AWS, Azure, or Google Cloud.
  • Ability to lead a scrum team and manage backlogs.
  • Willingness to participate in on-call rotations and odd hours.

Responsibilities

  • Automation and Efficiency: identify automation opportunities and build deployment tooling.
  • Team Collaboration: work with engineers to ensure scalable deployments and reliability.
  • Debugging and Advanced Support: diagnose bottlenecks and provide tiered support.
  • Monitoring & Maintenance: use Prometheus, Grafana, Stackdriver and related tools to improve health.
  • Incident Response: perform RCA for production incidents and follow SRE best practices.
  • Documentation: create runbooks and capture critical system knowledge.
  • Security Management: stay updated on security and resolve complex issues.
  • Issue Tracking: use Atlassian tools to track and resolve problems.

Skills

Scripting & automation
PowerShell
Python
Go
Ruby
Kubernetes
Docker
Serverless
Distributed systems
DevOps / SRE
Linux / CoreOS
AWS
Azure
Google Cloud
Scrum / Leadership
On-call readiness

Education

Bachelor's degree in Computer Science
Master's degree
Equivalent experience

Tools

Prometheus
Stackdriver
ElasticSearch
Grafana
SolarWinds
Atlassian tools

Job description

Job Summary

As a Cloud Infrastructure / Site Reliability Engineer, you will operate at the intersection of development and operations. You will engage and enhance all aspects of the cloud services lifecycle from design through deployment, operation, and refinement. You will be responsible for maintaining these services by measuring and monitoring their availability, latency, and overall system health and building automation for efficient cloud operations management. You will play a crucial role in sustainably scaling systems through automation and driving changes that improve reliability and velocity. As part of your responsibilities, you will administer cloud-based environments that support our SaaS/IaaS offerings implemented on a microservices, container-based architecture (Kubernetes). In addition, you will oversee a portfolio of customer‑centric cloud services (SaaS/IaaS), ensuring their overall availability, performance, and security. You will work closely with NetApp and cloud service provider teams (to include Azure) from Research Triangle Park (RTP), D.C., Pittsburg and more. Due to the critical nature of the services we support, this position involves participation in a rotation‑based on‑call schedule as part of our global team. This role offers the opportunity to work in a dynamic, global environment, ensuring the smooth operation of vital cloud services. To be successful in this role, you should be a motivated self‑starter and self‑learner, possess strong problem‑solving skills, and be someone who embraces challenges.

Responsibilities
  • Automation and Efficiency: Identify tasks and areas where automation can be applied to achieve time efficiencies and risk reduction. Develop software for deployment automation, packaging, and monitoring visibility.
  • Team Collaboration and Influence: Work in tandem with other Cloud Infrastructure Engineers and developers to ensure maximum performance, reliability, and automation of our deployments and infrastructure. Consult and influence developers on new feature development and software architecture to ensure scalability.
  • Debugging, Troubleshooting, and Advanced Support: Undertake debugging and troubleshooting of service bottlenecks throughout the entire software stack. Additionally, provide advanced tier 2 and 3 support for NetApp's Cloud Data Service solutions.
  • Analysis, and Infrastructure Maintenance: Continuously monitor, analyze, and measure system health, availability, and latency using tools like Prometheus, Stackdriver, ElasticSearch, Grafana, and SolarWinds. Develop strategies to enhance system and application performance, availability, and reliability. In addition, maintain and monitor the deployment and orchestration of servers, docker containers, databases, and general backend infrastructure.
  • Incident Response and Troubleshooting: Address and perform Root Cause Analysis (RCA) of complex live production incidents and cross‑platform issues involving OS, Networking, and Database in cloud‑based SaaS/IaaS environments. Implement SRE best practices for effective resolution.
  • Document system knowledge as you acquire it, create runbooks, and ensure critical system information is readily accessible.
  • Security Management: Stay updated with security protocols and proactively identify, diagnose, and resolve complex security issues.
  • Issue Tracking and Resolution: Use Atlassian’s tool chain along with first party cloud service management tools to track and resolve issues based on their priority.
  • Directly influence the decisions and outcomes related to solution implementation: measure and monitor availability, latency, and overall system health.
Job Requirements
  • This role requires US Citizenship, due to potential need for Security Clearance.
  • 5+ years experience in scripting and infrastructure automation using tools such as PowerShell, Python, Go or Ruby.
  • Deep working knowledge of Containers, Kubernetes, Serverless computing implementation, and distributed systems design patterns.
  • Knowledge of DevOps/SRE development methodologies.
  • Proficiency in Linux/Unix and CoreOS.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Ability to lead a scrum team, influence stakeholders to effectively maintain a product backlog, manage sprints.
  • This position will have ON‑CALL rotations as well as an ask to work odd hourss.
Education

A Bachelor of Science Degree in Computer Science, a master’s degree; or equivalent experience is required

Compensation

The target salary range for this position is 147,900 - 220,000 USD. The salary offered will be determined by the candidate's location, qualifications, experience, and education and may be outside of this range. Final compensation packages are competitive and in line with industry standards, reflecting a variety of factors, and include a comprehensive benefits package. This may cover Health Insurance, Life Insurance, Retirement or Pension Plans, Paid Time Off, various Leave options, Performance‑Based Incentives, employee stock purchase plan, and/or restricted stocks (RSU’s), with all offerings subject to regional variations and governed by local laws, regulations, and company policies. Benefits may vary by country and region, and further details will be provided as part of the recruitment process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer
Senior Software Engineer

NetApp • Waltham (MA)

Hybrid
USD 196,000 - 293,000
Health Insurance
Retirement Plans
Employee stock purchase plan
Software Engineer - Cloud Engineering (C++)
Software Engineer - Cloud Engineering (C++)

NetApp • Morrisville (NC)

Hybrid
USD 147,000 - 220,000
Health Insurance
Employee stock purchase plan
Restricted stock units (RSUs)
Software Engineer – Cloud Storage (GO , Python automation)
Software Engineer – Cloud Storage (GO , Python automation)

NetApp • San Jose (CA)

On-site
USD 180,000 - 210,000
Health Insurance
Employee stock purchase plan
RSUs
Lead Software Engineer – Cloud Storage (GO , Python automation)
Lead Software Engineer – Cloud Storage (GO , Python automation)

NetApp • San Jose (CA)

On-site
USD 215,000 - 245,000
Information Systems Engineer - NOC
Information Systems Engineer - NOC

NetApp • Morrisville (NC)

Hybrid
USD 113,000 - 168,000
Health Insurance
Retirement Plan
Employee Stock Purchase Plan
Senior Software Engineer
Senior Software Engineer

NetApp, Inc. • San Jose (CA)

Hybrid
USD 196,000 - 293,000
Health Insurance
Retirement Plan
Paid Time Off
Sr. Software Engineer (Systems Engineering, C/C++)
Sr. Software Engineer (Systems Engineering, C/C++)

NetApp • Morrisville (NC)

Hybrid
USD 196,000 - 293,000
Health Insurance
Life Insurance
Retirement or Pension Plans
+5
Software Engineer - Cloud Volumes
Software Engineer - Cloud Volumes

NetApp, Inc. • San Jose (CA)

Hybrid
USD 113,000 - 168,000
Senior Software Engineer
Senior Software Engineer

Worky • California (MO)

On-site
USD 196,000 - 293,000
Information Systems Engineer - NOC
Information Systems Engineer - NOC

NetApp, Inc. • Morrisville (NC)

Hybrid
USD 113,000 - 168,000