Systems Engineering, Metrics and Alerting

United States Digital Space LLC

United States

Hybrid

USD 76,000 - 105,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a Software Engineer for Observability to design, deliver, and operate a scalable observability platform. You will tackle bottlenecks in the Metrics & Alerting pipeline and contribute to highly distributed systems.

Ideal candidates have Go experience, strong knowledge of TSDBs/Columnar stores, Linux scalability, and familiarity with Prometheus, Alertmanager, and Thanos.

Qualifications

  • Software engineering background with proficiency in Go.
  • Proficiency in data structures and databases like TSDBs, Columnar stores.
  • Proficiency in distributed Linux environments.
  • Proficiency in designing high-scale distributed systems.
  • Proficiency in Prometheus, Alertmanager, Thanos.
  • Experience in fast, high-growth environments.
  • Experience in a 24/7/365 service environment.
  • Exquisite written and verbal communication skills.
  • Familiarity with Internetworking, networking protocols OSI 2-7 and BGP.
  • Strong bias for action.
  • Bonus: high-bandwidth transit Internetworking and routing.
  • Bonus: passion for code simplicity and performance.

Responsibilities

  • Design, deliver, and operate software and a platform that advances Observability capabilities.
  • Solve scaling bottlenecks in critical services of the Metrics & Alerting stack.
  • Work on highly distributed and scalable systems.
  • Participate in knowledge sharing and mentoring across the team.
  • Participate in global on-call rotation for services owned by your team.
  • Research and introduce cutting-edge technologies.
  • Contribute to open-source.

Skills

Go
Distributed systems
Data structures & TSDBs/Columnar
Linux environments
Networking fundamentals (OSI 2-7, BGP)
Strong written & verbal communication

Tools

Prometheus
Alertmanager
Thanos

Job description

About Us

At the company, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. the company protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by the company all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. the company was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company.

At the company, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in.

Available Locations:
  • London
  • LisbonAbout the Department

Production Engineering is responsible for the world’s most reliable, observable, performant, and safe network ecosystem. Our customers rely on our products and systems to safely modify, troubleshoot, and release products without external impact.

Our external customers rely on us to provide seamless and predictable incident, traffic, policy management, resulting in the fastest and safest network services in the world.

We are accountable for the overall performance of internal and external facing services, guiding our product teams to optimal configurations and maximum efficiency. From the moment that a packet enters the the company ecosystem, we know exactly what its expected purpose and behaviour is and we are capable of determining and exposing anomalous behaviour.

The the company network makes it possible to solve challenges at massive scale and efficiency which would be impossible for almost any organization.

About the Team

This role is for the internal Observability Team, responsible for the observability platform and stack to make our engineering teams productive. This includes (but is not limited to) areas like metrics, alerting, error tracking, logging, tracing, and more.

In this role, you can expect to:

  • Design, deliver, and operate software and a platform that progresses the company's Observability competency
  • Solve scaling bottlenecks in critical services in our Metrics & Alerting pipeline
  • Work on highly distributed and scalable systems
  • Participate in the constant cycle of knowledge sharing and mentoring
  • Participate in the global on-call rotation for the services your team owns
  • Research and introduce cutting-edge technologies
  • Contribute to open-source

We are a small team, well-funded, growing and focused on building an extraordinary company. This is a software engineering/systems engineering role and is a superb opportunity to be part of a high performing team to help to support the company’s mission and help build a better internet.

You may be a good fit for our team if you have:

  • A Software Engineering background and proficiency in high-level programming languages (e.g., Go)
  • Proficiency in Data structures and databases like TSDBs, Columnar stores or related
  • Proficiency in distributed Linux environments
  • Proficiency in designing high-scale distributed systems
  • Proficiency in Prometheus, Alertmanager, Thanos
  • Experience working in a fast, high-growth environment
  • Experience working in a 24/7/365 service environment
  • Exquisite written and verbal communication skills
  • Familiarity with Internetworking, networking protocols Layer 2-7 of the OSI model and BGP
  • Strong bias for action

Bonus points if you have:

  • Experience with high-bandwidth transit Internetworking and routing
  • Passion for code simplicity and performance
Compensation

-

For Portugal based hires: Estimated annual salary is between €66,000 - €91,000.

Equity

-

This role is eligible to participate in the company's equity plan

What Makes the company Special?

We’re not just a highly ambitious, large-scale technology company. We’re a highly ambitious, large-scale technology company with a soul. Fundamental to our mission to help build a better Internet is protecting the free and open Internet.

Project Galileo: Since 2014, we've equipped more than 2,400 journalism and civil society organizations in 111 countries with powerful tools to defend themselves against attacks that would otherwise censor their work, technology already used by the company’s enterprise customers--at no cost. Athenian Project:
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior System Engineer
Senior System Engineer

United States Digital Space LLC • United States

Hybrid
USD 76,000 - 105,000
Senior Engineering Manager, Observability
Senior Engineering Manager, Observability

United States Digital Space LLC • United States

Hybrid
USD 220,000 - 303,000
Medical Insurance
401(k) Retirement Savings
Equity participation
+1
Systems Engineer, Network Protocols & Distributed Systems
Systems Engineer, Network Protocols & Distributed Systems

United States Digital Space LLC • United States

Hybrid
USD 140,000 - 210,000
Flexible work hours
Take-what-you-need vacation
RSUs
+2
Software Engineer, Security Rules
Software Engineer, Security Rules

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 180,000
Equity participation
Health Insurance
Software Engineer - Egress (Go/Rust)
Software Engineer - Egress (Go/Rust)

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 190,000
Equity
System Engineer - Network Systems
System Engineer - Network Systems

United States Digital Space LLC • United States

Hybrid
USD 90,000 - 120,000
Senior Engineering Manager — Network Connectivity
Senior Engineering Manager — Network Connectivity

United States Digital Space LLC • United States

Hybrid
USD 150,000 - 230,000
Equity plan
Senior Software Engineer, Network On-Ramps
Senior Software Engineer, Network On-Ramps

United States Digital Space LLC • United States

Hybrid
USD 110,000 - 160,000
Senior Engineering Manager, Observability
Senior Engineering Manager, Observability

CloudFlare • Austin (TX), Northern (KY)

Hybrid
USD 190,000 - 270,000
Equity plan
Health & wellness benefits
401(k) retirement plan
+1
Systems Engineer, Growth Engineering
Systems Engineer, Growth Engineering

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 190,000