Site Reliability Engineer

Graphnet Health Ltd.

United States

Remote

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Graphnet Health Ltd. is seeking a Site Reliability Engineer to join the TechOps Team and own reliability of the CareCentric platform. The role involves monitoring, incident response, and collaboration with Operations, Security, Development, and Project Teams.

You will leverage Azure services, implement observability, and automate repetitive tasks. On-call rotation and continuous improvement are part of the job.

Qualifications

  • Minimum of four years working within an SRE function.
  • Experience of service monitoring and alerting.
  • Azure PaaS components including AKS, Application Insights / Log Analytics, App Services, Azure SQL, Storage.
  • Networking (NSG, VLANs etc).
  • Proven experience with Azure DevOps (ADO) and/or GitHub.
  • Experience in the support of Cloudflare.
  • Understanding of Terraform or other IAC tooling.
  • Experience of routine infrastructure upgrades and platform patching.
  • Demonstrable PowerShell administration / scripting skills.
  • Experience of Service Desk Systems.

Responsibilities

  • Provision of first-class infrastructure support to customers and systems for all Managed Service Platforms.
  • Implement observability and proactively monitor our Azure Cloud footprint, utilising dashboards, alerts and runbooks.
  • Incident response, analysis, remediation and associated post-incident root cause analysis reviews.
  • Support performance / reliability testing and capacity planning activities.
  • Perform routine system upgrades and platform patching.
  • Promote automation of repetitive tasks, reducing manual intervention.
  • Eventual participation in the 24x7 On-Call Rota.
  • Excellent, demonstrable troubleshooting and problem-solving skills.
  • Able to work well as an individual and as part of a team.
  • Take part in architectural discussions with technical / non-technical team members.
  • Organisational skills, with the ability to manage personal workloads in accordance with agreed timescales, whilst working under pressure.
  • An eye for detail and a desire to adhere to best practices.
  • Strong inter-personal and communication skills.
  • Have a desire to keep up with the latest tools and techniques.
  • Verbal, written communication and documentational skills.
  • Ability to explain technical concepts to key stakeholders of all levels.

Skills

SRE experience
Monitoring & alerting
Azure
AKS
Application Insights / Log Analytics
App Services
Azure SQL
Networking
Azure DevOps / GitHub
Terraform / IAC
PowerShell
Service Desk Systems

Tools

Azure DevOps (ADO)
GitHub
Terraform
PowerShell

Job description

The TechOps Team are responsible for the installation and support for all Technical Platforms, Operating Systems, Database Management Systems and associated products delivered to external customers of Graphnet. The Site Reliability Engineer (SRE) will work alongside specialists within the Team, taking responsibility for all aspects of the Team's work. With emphasis on the Technical Services function, you will be managing and supporting the infrastructure upon which Graphnet's CareCentric product operates. The SRE will work closely with many of Graphnet's departments, including Operations, Security, Development, and Project Teams, ensuring that our comprehensive support service is maintained. The successful candidate will be provided with on-the-job training for all supported Platforms, Operating Systems, Databases and Applications, but is expected to have prior, demonstrable experience of Microsoft Azure, Networking and Windows Server Technologies.,

  • Provision of first-class infrastructure support to customers and systems for all Managed Service Platforms
  • Implement observability and proactively monitor our Azure Cloud footprint, utilising dashboards, alerts and runbooks
  • Incident response, analysis, remediation and associated post-incident root cause analysis reviews
  • Support performance / reliability testing and capacity planning activities
  • Perform routine system upgrades and platform patching
  • Promote automation of repetitive tasks, reducing manual intervention
  • Eventual participation in the 24x7 On-Call Rota
  • Excellent, demonstrable troubleshooting and problem-solving skills
  • Able to work well as an individual and as part of a team
  • Take part in architectural discussions with technical / non-technical team members
  • Organisational skills, with the ability to manage personal workloads in accordance with agreed timescales, whilst working under pressure
  • An eye for detail and a desire to adhere to best practices
  • Strong inter-personal and communication skills
  • Have a desire to keep up with the latest tools and techniques
  • Verbal, written communication and documentational skills
  • Ability to explain technical concepts to key stakeholders of all levels
Education and Skills: Essential:
  • Minimum of four years working within an SRE function
  • Experience of service monitoring and alerting
  • Microsoft Azure PaaS components including (but not limited to):
  • Azure Kubernetes Service (AKS)
  • Application Insights / Log Analytics
  • App Services
  • Azure SQL
  • Storage
  • Networking (NSG, VLANs etc)
  • Proven experience with Azure DevOps (ADO) and / or GitHub
  • Experience in the support of Cloudflare
  • An understanding of Terraform or other IAC tooling
  • Experience of routine infrastructure upgrades and platform patching
  • Demonstrable PowerShell administration / scripting skills
  • Experience of Service Desk Systems
Desirable: Any skills or experience in the areas below would be considered an advantage for any potential candidate, but are not essential:
  • Experience or exposure of Terraform IAC software tooling
  • Failover Cluster Management
  • Understanding of Grafana Analytics and Monitoring Solution
  • Nessus Vulnerability Assessment Solution
  • NinjaOne Patch Management
  • Knowledge of ITIL Foundation principles and application
  • Knowledge of ISO27001, ISO27018, ISO9001 and CE+ Certifications
  • Any vendor certifications
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Cosm Inc. • El Segundo (CA), Northern (KY)

On-site
USD 110,000 - 145,000
Site Reliability Engineer
Site Reliability Engineer

Moultrie • Birmingham (AL)

On-site
USD 110,000 - 170,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000
Azure Cloud SRE & Reliability Engineer
Azure Cloud SRE & Reliability Engineer

Graphnet Health Ltd. • United States

Remote
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

CT19 • Massachusetts

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG • Raleigh (NC)

On-site
USD 140,000 - 190,000
Healthcare
Retirement planning
Volunteer days
+1