Staff Site Reliability Engineer

United States Digital Space LLC

United States

On-site

USD 180,000 - 240,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health & wellness
Equity

Job summary

AWAKE is seeking a Staff Engineer to design reliable, scalable systems and mentor engineers within the Platform Infrastructure team. You will influence roadmaps, collaborate across AI/ML, Data, Platform, and Product teams, and drive reliability through SLIs/SLOs.

As a Staff Engineer, you will shape architecture, promote best practices, and lead initiatives that improve platform resilience and efficiency, while navigating a fast-paced production environment.

Qualifications

  • 7+ years of experience in Production Engineering, Backend Engineering, SRE, DevOps or similar role.
  • Strategic visionary: your strong technical background enables you to look beyond solving the immediate problem, planning for the future.
  • Proficient Problem-Solver: strong coding ability in at least one language (e.g., Golang, Python, Java, Typescript) with the capability to solve complex issues through code.
  • Track Record of Success: demonstrated experience delivering medium to large-scale projects that drive meaningful improvements in platform reliability and scalability.
  • Reliability Expertise: deep understanding of production reliability concepts, including SLIs, SLOs, and incident management.
  • Strong Communicator: excellent verbal and written communication skills with the ability to influence and collaborate across technical and non-technical teams.
  • Fast-Paced Experience: familiarity with working in dynamic, reliability-focused production environments (preferred).

Responsibilities

  • Design and implement reliable, observable systems to scale the platform.
  • Lead cross-team initiatives with technical leadership.
  • Collaborate with AI/ML, Data, Platform, Product teams to develop best-in-class services.
  • Establish production standards, processes, and tools for operational excellence.
  • Advocate for SLIs, SLOs, and reliability-focused metrics across the organization.
  • Mentor team members and foster technical growth.
  • Drive continuous improvement with creative ideas and challenging the status quo.

Skills

Production Eng
Backend Eng
SRE/DevOps
Strategic thinking
Problem solving
Communication
Fast-paced env

Job description

the company® is the AI marketing platform for 1:1 personalization redefining the way brands and people connect. We’re the only marketing platform that combines powerful technology with human expertise to build authentic customer relationships. By unifying SMS, RCS, email, and push notifications, our AI-powered personalization engine delivers bespoke experiences that drive performance, revenue, and loyalty through real-time behavioral insights.

Recognized as the#1provider in SMS Marketing by G2, the company partners with more than 8,000 customers across 70+ industries. Leading global brands like Crate and Barrel, Urban Outfitters, and Carter’s work with us to enable billions of interactions that power tens of billions in revenue for our customers.

With a distributed global workforce and employee hubs in New York City, San Francisco, London, and Sydney, the company’s team has been consistently recognized for its performance and culture. We’re proud to be included inDeloitte’s Fast 500(four years running!),LinkedIn’s Top Startups,Forbes’ Cloud 100 (five years running!),Inc.’s Best Workplaces, and theHuman Rights Campaign Foundation's Corporate Equality Index!

About the Role

Our Platform Infrastructure team is the backbone of everything we do at the company, providing a resilient and cost-effective platform that seamlessly handles billions of events from over 100 million customers daily. We own everything from compute, persistence, and networking to observability and deployments. Joining our team offers a high-growth career opportunity to collaborate with some of the world’s most talented engineers in a high-performance, high-impact culture.

As part of the Infrastructure and Platform organization, the Production Engineering Team is focused on delivering a fast and reliable platform that empowers the company engineers to deliver solutions quickly and safely. We build scalable systems that automate routine tasks so we can focus on other impactful efforts. Reliability, scalability, and security are our areas of expertise. We focus on release, observability, and cost optimization. Our mission is to create robust platforms and tools that allow stakeholders to concentrate on delivering exceptional products.

As a Staff Engineer, you will take a strategic role in designing and implementing solutions that enhance the reliability and scalability of our systems, while mentoring others and influencing technical roadmaps across the organization.

What You’ll Accomplish
  • Design and Deliver High-Impact Solutions: Design and implement systems that enhance reliability, observability, traceability, and incident management, ensuring the platform scales effectively
  • Lead Strategic Initiatives: Take ownership of cross-team collaborations and drive impactful projects by providing technical leadership and guidance
  • Partner Across Teams: Collaborate with engineers from AI/ML, Data, Platform, and Product teams to develop best-in-class services
  • Partner with engineers from AI/ML, Data, Platform, Product, and other groups to deliver best-in-class services
  • Establish Standards and Best Practices: Define and enforce production standards, processes, and tools to ensure operational excellence
  • Champion Reliability Goals: Advocate for and implement SLIs, SLOs, and other reliability-focused metrics across the engineering organization
  • Mentorship and Knowledge Sharing: Guide and mentor team members, fostering technical growth and helping to develop the next generation of engineering leaders
  • Innovate and Inspire: Drive continuous improvement by bringing creative ideas and challenging the status quo
Your Expertise
  • 7+ years of experience in Production Engineering, Backend Engineering, SRE, DevOps or similar role
  • Strategic visionary: Your strong technical background enables you to look beyond solving the immediate problem, planning for the future.
  • Proficient Problem-Solver: Strong coding ability in at least one language (e.g., Golang, Python, Java, Typescript) with the capability to solve complex issues through code
  • Track Record of Success: Demonstrated experience delivering medium to large-scale projects that drive meaningful improvements in platform reliability and scalability
  • Reliability Expertise: Deep understanding of production reliability concepts, including SLIs, SLOs, and incident management
  • Strong Communicator: Excellent verbal and written communication skills with the ability to influence and collaborate across technical and non-technical teams
  • Fast-Paced Experience: Familiarity with working in dynamic, reliability-focused production environments (preferred)

You'll get competitiveperks and benefits, from health & wellness to equity, to help you bring your best self to work.

For US based applicants
  • The US base salary range for this full-time position is $180,000 - $240,000 annually+ equity + benefits
  • Our salary ranges are determined by role, level and location

#LI-HB1

By applying for this position, your data will be processed as per the company's Privacy Policy.the company Company ValuesDefault to Action

  • Move swiftly and with purpose
  • Be One Unstoppable Team
  • Champion the Customer
  • Act Like an Owner

Learn more about AWAKE, the company’s collective of employee resource groups.

If you do not meet all the requirements listed here, we still encourage you to apply! No job description is perfect, and we may also have another opportunity that closely matches your skills and experience.

At the company, we know that our Company's strength lies in the diversity of our employees. the company is an Equal Opportunity Employer and we welcome applicants from all backgrounds. Our policy is to provide equal employment opportunities for all employees, applicants and covered individuals regardless of protected characteristics. We prioritize and maintain a fair, inclusive and equitable workplace free from discrimination, harassment, and retaliation. the company is also committed to providing reasonable accommodations for candidates with disabilities. If you need any assistance or reasonable accommodations, please let your recruiter know.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Segmentation
Engineering Manager, Segmentation

United States Digital Space LLC • United States

On-site
USD 210,000 - 240,000
Equity
Health benefits
Career growth opportunities
Senior Software Engineer, ML/AI Platform
Senior Software Engineer, ML/AI Platform

United States Digital Space LLC • United States

On-site
USD 180,000 - 250,000
Health benefits
Equity
Wellness programs
Staff Software Engineer, Onsite Customer Growth
Staff Software Engineer, Onsite Customer Growth

United States Digital Space LLC • New York (NY)

Hybrid
USD 205,000 - 275,000
Health & wellness
Equity
Hybrid work
Support Engineer
Support Engineer

United States Digital Space LLC • United States

On-site
USD 75,000 - 85,000
Software Engineer II, Integrations
Software Engineer II, Integrations

United States Digital Space LLC • United States

On-site
USD 135,000 - 170,000
Health & wellness
Equity
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Attentive • Wilmington (DE)

On-site
USD 180,000 - 240,000
Health & wellness
Equity
Staff Site Reliability Engineer
Staff Site Reliability Engineer

JobCubby • Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity
Health benefits
Staff Software Engineer, User Profile
Staff Software Engineer, User Profile

United States Digital Space LLC • United States

On-site
USD 180,000 - 230,000
Health benefits
Equity
Wellness programs
Senior Software Engineer, Intelligent Messaging, AI Journeys
Senior Software Engineer, Intelligent Messaging, AI Journeys

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 230,000
Equity
Health benefits
Global remote-friendly team
Software Engineer II, Segment Optimization
Software Engineer II, Segment Optimization

United States Digital Space LLC • United States

On-site
USD 135,000 - 170,000