SRE - Remote

Glassbox

New York (NY)

On-site

USD 120,000 - 150,000

Full time

37 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Glassbox seeks an experienced Site Reliability Engineer (SRE) to join our global Cloud team. You will support AWS-based environments and Kubernetes-based systems (EKS), driving reliability, monitoring, and automation.

You will build scalable monitoring, runbooks, and proactive alerting; participate in 24x7 on-call rotations and collaborate with DevOps to move products into production.

Qualifications

  • At least 3 years of experience as an SRE/DevOps or in a similar cloud/monitoring role.
  • Hands-on experience with AWS.
  • Strong knowledge and practical experience with Kubernetes and EKS.
  • Scripting experience with Bash and Linux systems.
  • Experience with cloud monitoring, management, and alerting tools.
  • Strong troubleshooting skills in production environments.
  • Willingness to participate in a 24x7 on-call rotation.
  • Ability to work effectively in a collaborative team and fast-paced environment.
  • Assertive, quick learner, and proactive problem solver.

Responsibilities

  • Provide technical and operational support for customers according to defined SLAs.
  • Work in cloud-based environments (AWS) and operate Kubernetes-based systems (EKS).
  • Design, build, and maintain automation to improve reliability and observability.
  • Develop scalable monitoring and alerting solutions to detect issues before they impact customers.
  • Create and maintain runbooks for NOC/SOC teams.
  • Serve as Tier-2 escalation for production incidents and participate in on-call rotations.
  • Implement automation-driven improvements using scripts and configuration management tools.
  • Collaborate with Cloud DevOps to transition products from development to production.

Skills

AWS
Kubernetes
EKS
Bash scripting
Linux
Cloud monitoring
On-call rotation
Team collaboration
Fast-paced environment

Education

Bachelor's degree in Computer Science/Information Systems or related field

Tools

Prometheus
Grafana

Job description

The Company

Glassbox's mission is to empower enterprises to shape trusted, frictionless digital experiences.

The Company

Glassbox's mission is to empower enterprises to shape trusted, frictionless digital experiences. Glassbox is a leading force in shaping digital experiences. It helps organizations uncover digital issues, boost conversion rates, enhance accessibility, prevent fraud, and more. Leveraging AI-driven customer intelligence, Glassbox enables enterprises to deliver secure, proactive, and preventative digital experiences. Its solutions are trusted by highly regulated organizations, including SoFi, Cal, and many others. We are growing and have been recognized by G2 as one of 2024's Top 50 Software Companies in the world.

The Opportunity

Glassbox is looking for an SRE to join our global Cloud team.

What will you do?
  • Provide technical and operational support for customers according to defined SLAs.
  • Work in cloud-based environments (primarily AWS) and operate/support Kubernetes-based systems (EKS).
  • Design, build, and maintain advanced automation systems to enhance reliability, monitoring, and operational efficiency across production environments.
  • Develop scalable monitoring and alerting solutions to proactively detect issues before they impact customers.
  • Build and maintain runbooks for NOC/SOC teams.
  • Serve as Tier-2 escalation for production incidents, including collaboration with DevOps and participation in a 24x7 on-call rotation.
  • Implement automation-driven improvements using scripts and configuration management tools to streamline system operations.
  • Leverage modern technologies and tooling to optimize system performance, observability, and resilience.
  • Work closely with the Cloud DevOps team to transition products from development to the production environment via continuous integration and deployment processes
What will you need?
  • At least 3 years of experience as an SRE/DevOps or in a similar cloud/monitoring role.
  • Hands-on experience with AWS.
  • Strong knowledge and practical experience working with Kubernetes and EKS.
  • Scripting experience with Bash and hands-on experience working with Linux systems.
  • Experience working with cloud monitoring, management, and alerting tools.
  • Strong troubleshooting skills in production environments.
  • Willingness to participate in a 24x7 on-call rotation.
  • Ability to work effectively as part of a collaborative team, with strong interpersonal skills and a positive, team-oriented mindset.
  • Assertive, confident, a fast learner, and comfortable working in a fast-paced environment
Advantage:
  • Experience with Azure and AKS.
  • Knowledge of additional programming languages.
  • Experience with Prometheus, Grafana, or similar monitoring/observability tools.
  • Bachelor's degree in Computer Information Systems, Management Information Systems, Computer Science, or another related field experience.
  • AWS or Azure certifications
Our Commitment

At Glassbox, we value curiosity, ownership, and a constant drive to learn and improve, no matter the gender, age, nationality, religion, or background. We believe diverse perspectives make us better, and that potential matters just as much as experience, and encourage people of all shapes and sizes to apply.

If this role excites you and you're motivated to make an impact - we'd love to hear from you, even if you don't meet every listed qualification.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE: Scale Reliability with Cloud Automation
Remote SRE: Scale Reliability with Cloud Automation

Glassbox • New York (NY)

On-site
USD 120,000 - 150,000
Infrastructure/Cloud DevOps - SRE
Infrastructure/Cloud DevOps - SRE

Bayside Solutions • Cupertino (CA)

On-site
USD 150,000 - 230,000
SRE (Site Realiability Engineer)
SRE (Site Realiability Engineer)

STRATIS Cloud Tech Solutions INC • Arkansas

On-site
USD 110,000 - 150,000
Competitive salary
Growth and learning opportunities
Friendly, collaborative team
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
CloudDevs: Senior Site Reliability Engineer (SRE)
CloudDevs: Senior Site Reliability Engineer (SRE)

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
SRE/Devops Engineer
SRE/Devops Engineer

INSPYR Solutions • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours