Site Reliability Engineer

Skyhigh Security

Frisco (TX)

Hybrid

USD 110,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Retirement Plans
Medical, Dental and Vision Coverage
Paid Time Off
Paid Parental Leave
Support for Community Involvement

Job summary

A leading cloud security company based in Texas seeks a Site Reliability Engineer to monitor and maintain high-availability production environments. The role involves incident management, troubleshooting, and collaboration with engineering teams. Candidates should have a Bachelor's degree in Computer Science and 7+ years of SRE experience, as well as expertise in Linux systems and various monitoring tools like Prometheus and Grafana. This position offers a hybrid work model, enhancing work-life flexibility.

Qualifications

  • Bachelor’s degree in computer science or related area with 7+ years of SRE experience.
  • System admin experience on Linux environments.
  • Experience with Prometheus, Grafana for monitoring setups.

Responsibilities

  • Monitor and troubleshoot operational issues in a high-availability production environment.
  • Perform root cause analysis on major incidents.
  • Implement proactive monitoring and alerting solutions.

Skills

SRE experience in a large enterprise organization
System admin experience on Linux environments
Experience with end-to-end monitoring setup
Experience with Prometheus, Grafana, ELK
Experience with Cloud Technologies like AWS
Experience with containerized workloads tools
Network knowledge (TCP/IP, UDP, DNS)
Ability to script/program with languages like Python
Experience with configuration management tools
Strong communication and analytical skills

Education

Bachelor’s degree in computer science or related

Tools

Prometheus
Grafana
AWS
Kubernetes
Jenkins
GitHub

Job description

Job Title

Site Reliability Engineer

About Skyhigh Security

Skyhigh Security is a dynamic, fast‑paced, cloud company that is a leader in the security industry. Our mission is to protect the world’s data, and because of this, we live and breathe security. We value learning at our core, underpinned by openness and transparency. Since 2011, organizations have trusted us to provide them with a complete, market‑leading security platform built on a modern cloud stack. Our industry‑leading suite of products radically simplifies data security through easy‑to‑use, cloud‑based, Zero Trust solutions that are managed in a single dashboard, powered by hundreds of employees across the world. With offices in Santa Clara, Aylesbury, Paderborn, Bengaluru, Sydney, Tokyo and more, our employees are the heart and soul of our company. Skyhigh Security is more than a company; here, when you invest your career with us, we commit to investing in you. We embrace a hybrid work model, creating the flexibility and freedom you need from your work environment to reach your potential. From our employee recognition program, to our “Blast Talks” learning series, and team celebrations (we love to have fun!), we strive to be an interactive and engaging place where you can be your authentic self.

Follow us on LinkedIn & Twitter: @SkyhighSecurity.

Role Overview

The Site Reliability Engineer at Skyhigh Security will be responsible for monitoring, maintaining, and troubleshooting operational issues of a high‑availability production environment.

Job Summary

The Site Reliability Engineer at Skyhigh Security will be responsible for monitoring, maintaining and troubleshooting operational issues of a high‑availability production environment. The SRE will also act as a bridge between Operations, Engineering and Product Management teams and you will represent the customer point of view to continue driving enhancements to our products and uptime. SREs are responsible for managing and improving the operational aspects of systems, such as monitoring, alerting, incident response, and vendor interactions.

Only US Citizens are eligible.

About the role
  • Perform Incident Management and Change Management to maintain the continuous availability of all Cloud Infrastructure services.
  • Ensure all SRE and operating procedures are maintained and executed.
  • Maintain a 24x7 production environment with a high level of service availability and perform quality reviews, manage operational issues.
  • Perform root cause analysis for major incidents and drive the process by involving required stakeholders.
  • Perform problem management by analyzing metrics, alarms and dashboards to troubleshoot problem areas, report issues to assist in performance tuning and fault finding.
  • Implementation of proactive monitoring, alerting, trend analysis, and self‑healing solutions.
  • Explore and innovate new technologies, features, and tools to improve the platform and automate operational tasks using Bash, Python or any other programming language.
  • Manage and maintain Runbooks and Standard Operating procedures
  • Manage, coordinate, and document all types of maintenance activities and outages.
  • Perform patching and upgrades for vulnerability management.
  • Work closely with the teams to initiate the development of new ideas into internal tools.
  • Understand the existing architecture and work with various Engineering teams to develop and execute strategies to provide a high‑quality production service.
  • Capable of working a flexible work schedule in a 24x7 environment with rotational shifts.
About you
  • Bachelor’s degree in computer science, electrical engineering or a related area, with 7+ years of SRE experience in a large enterprise organization
  • System admin experience on Linux environments.
  • Experience with end‑to‑end monitoring setup for infra and applications
  • Experience with Prometheus, Grafana, ELK, Opensearch, Cloudwatch, PagerDuty and other monitoring tools.
  • Solid experience with Cloud Technologies such as AWS and OCI.
  • Good experience with containerized workloads tools like Kubernetes.
  • Network knowledge (TCP/IP, UDP, DNS, Load balancing) and prior network administration experience is required.
  • Experience with BGP, NAT, TCP/IP, iBGP, Proxies, Cross connects.
  • Experience with L2/L3 switching, knowledge of Juniper and Cisco routing devices.
  • Experience understanding and managing web servers (Apache, Tomcat, Nginx)
  • Ability to script/program with one or more high level languages, such as Python, Go, etc.
  • Experience with any configuration management tools like Salt or Puppet or Ansible or similar.
  • Experience with source control tools such as Github and SVN.
  • Experience with deployment tools Jenkins, Harness etc.
  • Experience with SQL and NoSQL databases like Redis, Crate, Elasticsearch.
  • Experience in performing and writing Root Cause Analysis documents.
  • Strong communication and analytical/problem‑solving skills.
  • Systematic approach and to drive problems to resolution.
  • Good to have experience/knowledge of GCP, Azure
  • Experience in Security domain will be added advantage
  • Experience with open‑source technologies like Kafka, Hadoop, HBase, Zookeeper, Oozie will be an added advantage.
Company Benefits and Perks
  • Retirement Plans
  • Medical, Dental and Vision Coverage
  • Paid Time Off
  • Paid Parental Leave
  • Support for Community Involvement

We’re serious about our commitment to a workplace where everyone can thrive and contribute to our industry‑leading products and customer support, which is why we prohibit discrimination and harassment based on race, color, religion, gender, national origin, age, disability, veteran status, marital status, pregnancy, gender expression or identity, sexual orientation or any other legally protected status.

Seniority level
  • Mid‑Senior level
Employment type
  • Full‑time
Job function
  • Engineering and Information Technology
  • Industries: Software Development
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Site Reliability Engineer
Site Reliability Engineer

VantageScore® • San Francisco (CA)

On-site
USD 150,000
Medical insurance
Dental insurance
401(k) plan
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Oaks (PA)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare coverage
401(k) matching
Tuition reimbursement
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Site Reliability Engineer - 7 Month Contract
Site Reliability Engineer - 7 Month Contract

Orion Health group • Fort Worth (TX), Town of Texas (WI)

On-site
USD 110,000 - 170,000