Systems Engineer

United States Digital Space LLC

United States

Hybrid

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity participation
Health insurance
401(k) retirement plan
Flexible paid time off

Job summary

United States Digital Space LLC is seeking a highly motivated software engineer for our Production Platform Organization. You will build infrastructure to collect, store, and provide reliability data for monitoring, collaborating with Product Managers and SREs to measure service quality at enterprise scale.

You will design and implement SLI/SLO plans, help define platform reliability objectives, and work with cross-functional teams to improve software development and product reliability

Qualifications

  • Proven track record as a software engineer or similar role with a deep understanding of developing and maintaining distributed systems.
  • System design experience for secure, highly available distributed systems.
  • Programming experience in Go, Rust, or Python.
  • Deep understanding of uptime metrics using SLOs/SLIs.

Responsibilities

  • Build production reliability testing infrastructure and reporting.
  • Define uptime metrics with SLOs/SLIs and implement plans to verify systems.
  • Collaborate with engineering teams to understand system interactions at scale.
  • Provide clear feedback to product and engineering teams to improve reliability.

Skills

Go
Rust
Python
Distributed systems
System design
Reliability metrics

Tools

Clickhouse
Prometheus
GraphQL
Postgres

Job description

About Us

At the company, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. the company protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by the company all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. the company was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company.

At the company, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a \"normalized\" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in.

Available Locations: Austin, TX

About the Department

the company’s engineering teams build and maintain the systems and products that power our global platform. A global platform which is within approximately 50 milliseconds of about 95% of the Internet connected population, serving on average, over 46 million HTTP requests per second.

About the role

the company engineering delivers code to production at a tremendous pace, and depends on automated testing to do so without incidents. The SLO team builds and runs the internal platform and tooling that empowers other engineering teams to set up Service Level Indicators (SLIs) and effectively measure their Service Level Objectives (SLOs). This enables all engineering teams to effectively measure their service and feature reliability that verify the interactions between systems and products in production at huge scale.

We are looking for a highly motivated software engineer to join our Production Platform Organization. You will build the infrastructure necessary to collect, store, and make reliability data easily accessible for monitoring needs. You’ll need to communicate effectively and proactively with engineers across the company to deeply understand the behaviors of our systems and refine their reliability objectives. You will also work closely with Product Managers and Product Site Reliability Engineers on quality of service measurements for enterprise customers.

What You Will Do

  • Build the Platform: Create and maintain production reliability testing infrastructure and availability reporting.
  • Define Reliability Metrics: Measure uptime metrics like correctness, availability, and latency SLIs/SLOs. Develop, document, and execute SLI/SLO plans to verify systems continue to operate as expected.
  • Collaborate Cross-Functionally: Collaborate with engineering teams to understand how their systems function and interact with other the company systems in production at a huge scale.
  • Communicate & Improve: Provide clear and concise feedback to engineering and product teams as an excellent communicator. Help drive continued improvements in the software development and reliability measurement processes.

What You Will Need

  • Experience: Proven track record as a software engineer or similar role with a deep understanding of developing and maintaining distributed systems.
  • System Design: Experience designing, implementing, and maintaining secure and highly-available distributed systems.
  • Programming Languages: Programming experience with one of the following languages: Go, Rust, or Python.
  • Reliability Metrics: Deep understanding and hands-on experience measuring uptime metrics like correctness, availability, and latency using SLOs/SLIs.

Bonus Points

  • Experience working with Clickhouse, Prometheus, GraphQL, and Postgres.
  • Experience working with data pipelines with a focus on reliability and scale.
  • Experience working with synthetic traffic & load testing tools.
  • Experience developing reliable, extensible platforms that other engineers can trust and leverage.

Equity

This role is eligible to participate in the company’s equity plan.

Benefits

the company offers a complete package of benefits and programs to support you and your family. Our benefits programs can help you pay health care expenses, support caregiving, build capital for the future and make life a little easier and fun! The below is a description of our benefits for employees in the United States, and benefits may vary for employees based outside the U.S.

Health & Welfare Benefits

  • Medical/Rx Insurance
  • Dental Insurance
  • Vision Insurance
  • Flexible Spending Accounts
  • Commuter Spending Accounts
  • Fertility & Family Forming Benefits
  • On-demand mental health support and Employee Assistance Program
  • Global Travel Medical Insurance

Financial Benefits

  • Short and Long Term Disability Insurance
  • Life & Accident Insurance
  • 401(k) Retirement Savings Plan
  • Employee Stock Participation Plan

Time Off

  • Flexible paid time off covering vacation and sick leave
  • Leave programs, including parental, pregnancy health, medical, and bereavement leave

What Makes the company Special?

We’re not just a highly ambitious, large-scale technology company. We’re a highly ambitious, large-scale technology company with a soul. Fundamental to our mission to help build a better Internet is protecting the free and open Internet.

Project Galileo: Since 2014, we've equipped more than 2,400 journalism and civil society organizations in 111 countries with powerful tools to defend themselves against attacks that would otherwise censor their work, technology already used by the company’s enterprise customers--at no cost.

Athenian Project: In 2017, we created the Athenian Project to ensure that state and local governments have the highest level of protection and reliability for free, so that their constituents have access to election information and voter registration. Since the project, we've provided services to more than 425 local government election websites in 33 states.

1.1.1.1: We released 1.1.1.1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Systems Engineer, DevTools
Principal Systems Engineer, DevTools

United States Digital Space LLC • United States

Hybrid
USD 200,000 - 281,000
Medical/Rx Insurance
Dental Insurance
401(k) Retirement Savings Plan
+1
Senior Software Engineer, Network Performance & Reliability
Senior Software Engineer, Network Performance & Reliability

United States Digital Space LLC • United States

Hybrid
USD 150,000 - 230,000
Equity plan
Medical/Rx Insurance
401(k) Retirement Savings Plan
Senior Manager, Engineering
Senior Manager, Engineering

United States Digital Space LLC • United States

Hybrid
USD 170,000 - 230,000
Systems Engineer
Systems Engineer

Webhosting • Austin (TX)

On-site
USD 150,000 - 230,000
Equity plan
Senior Customer Engineer, Majors - Toronto, CA
Senior Customer Engineer, Majors - Toronto, CA

United States Digital Space LLC • East Portal Distributed Camping Area (CO)

On-site
USD 140,000 - 180,000
Senior Systems Engineer, Identity & Access Management
Senior Systems Engineer, Identity & Access Management

United States Digital Space LLC • United States

Hybrid
USD 140,000 - 210,000
Presales Customer Engineer (Sydney)
Presales Customer Engineer (Sydney)

United States Digital Space LLC • United States

Hybrid
USD 140,000 - 210,000
Product Security Engineer
Product Security Engineer

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 160,000
Medical/Rx Insurance
401(k) Retirement Savings Plan
Employee Stock Participation Plan
+1
Data Center Engineer
Data Center Engineer

United States Digital Space LLC • United States

Hybrid
USD 85,000 - 125,000
Software Engineer-
Software Engineer-

United States Digital Space LLC • United States

Hybrid
USD 140,000 - 190,000