Software Engineer, Observability

Whatnot Inc.

Kraków

Hybrid

PLN 240,000 - 360,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Whatnot is building a world-class live commerce platform with a remote-first, globally distributed team. The Infrastructure Reliability Engineering group seeks seasoned engineers to rethink observability at scale—ensuring visibility into state, performance, and user experience across our complex software stack.

You will work closely with Core Infrastructure, Platform, and Developer Tools teams to design end-to-end observability—from data collection to normalization, routing, processing, and

Qualifications

  • 7+ years of experience designing and building large-scale distributed systems.
  • Experience in Python, Elixir, or Go is preferred.
  • Strong understanding of observability principles including metrics, logs, events, and traces.
  • Experience with cloud-native environments (AWS or GCP) and Kubernetes.
  • Familiarity with OpenTelemetry and modern data collection pipelines.

Responsibilities

  • Redesign and implement observability across the stack from data collection to visualization.
  • Collaborate with Core Infrastructure, Platform, and Developer Tools teams to scale observability.
  • Leverage analytics engines and vendors to predict incidents and enable rapid resolution.
  • Use AI-assisted troubleshooting to improve problem diagnosis and mean time to recovery.

Skills

Distributed systems design
Python
Go
Observability
Cloud-native (AWS/GCP)
Kubernetes
OpenTelemetry
Communication

Tools

OpenTelemetry

Job description

Join the Future of Commerce with Whatnot!

Whatnot is the largest live shopping platform in North America and Europe to buy, sell, and discover the things you love. Whether it's trading cards, fashion, electronics, or live plants, our sellers are building real businesses across hundreds of categories. We're building live commerce at a scale that's never been done in the West, and there's no playbook to copy. The people here are shaping how an entirely new industry develops.


As a remote co-located team, we're inspired by our values and anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia. We move fast, stay close to our users, and focus on the work that drives the most impact.


We're one of the fastest growing marketplaces and were recently named the #1 Best Startup Employer in America by Forbes. Check out the latest Whatnot updates on our news and engineering blogs and join us as we enable anyone to turn their passion into a business and bring people together through commerce.


Role

The Infrastructure Reliability Engineering team is looking for seasoned Software Engineers who will be responsible for rethinking and redesigning how we do Observability at Whatnot. As our scale, traffic, and complexity continue to grow; yesterday’s tools, vendors and platforms are becoming obsolete and impractical. Your job is to ensure that we have visibility into the state, performance, reliability and user’s experiences of our software stack, and that this visibility remains 10x and 100x our current scale.


This is hands-on, software-first, systems engineering work. You will work closely with the Core Infrastructure, Platform and Developer Tools teams to redesign and implement every step of observability, starting from collecting data in the code, through normalization, routing and processing, to querying and visualisation. We’re at the scale where it’s no longer reasonable to use simple, off-the-shelf solutions. You will have to use a combination of analytics engines and vendors to make sure we are able to predict incidents, and should they occur, have a swift path to resolution. At Whatnot, we are utilizing AI agents heavily during troubleshooting - our observability stack should take leverage of that.


You will be working on creating a platform that:



  • Can identify leading indicators for issues before they become a problem


  • Understands that not all data is equal and can work both with high-volume, low-signal data and low-volume, high-signal data


  • Uses industry good practices – based on open standards (Otel, Semconv)


  • Utilizes any set of open-source software or vendors to achieve its goals


  • Interfaces with infrastructure to provide better signals for infrastructure management than CPU utilization


  • Correlates events from infrastructure, CI/CD pipeline, experimentations or load tests with logs and time-series data to quickly identify \"what has changed?\"


  • Provides good experience for a wide range of users: application engineers, incident responders, leadership


  • Aids in driving reliability and experience for our customers, staying reliable and effective itself



This is a highly visible role. The Reliability team provides foundational systems and frameworks that allow Whatnot to scale rapidly while remaining stable and trustworthy for buyers and sellers.


About You

People who do well at Whatnot tend to be comfortable figuring things out as they go, biased toward action, and genuinely curious about what they're building. They care more about outcomes than credit and stay close to the product and the people using it.



  • 7+ years of experience designing and building large-scale distributed systems. Experience in Python, Elixir, or Go is preferred, but strong engineers from other backend stacks who are eager to learn are welcome.


  • You identify as a software engineer first. You want to build systems and write code, not just configure infrastructure or respond to pages.


  • Have a strong understanding of observability principles, knowing different types of metrics, logs, events and traces; and where they are applicable


  • Strong fundamentals in designing, building, and operating shared production services and frameworks.


  • Experience with one or more of the following:



    • OpenTelemetry


    • OLAP databases


    • Creating developer-facing tools, libraries, frameworks


    • High-traffic, real-time, or event-driven systems




  • Comfortable in cloud-native environments such as AWS or GCP with Kubernetes and infrastructure as code.


  • Strong collaborator with clear written and verbal communication skills.



EOE

Whatnot is proud to be an Equal Opportunity Employer. We value diversity, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, parental status, disability status, or any other status protected by local law. We believe that our work is better and our company culture is improved when we encourage, support, and respect the different skills and experiences represented within our workforce.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Observability
Software Engineer, Observability

Whatnot • Kraków

Hybrid
PLN 520,000 - 580,000
Remote Observability Platform Engineer for Scalable Systems
Remote Observability Platform Engineer for Scalable Systems

Whatnot Inc. • Kraków

Hybrid
PLN 240,000 - 360,000
Remote Observability Engineer - Scale & Reliability
Remote Observability Engineer - Scale & Reliability

Whatnot • Kraków

Hybrid
PLN 520,000 - 580,000
Senior Backend Engineer, International Engineering
Senior Backend Engineer, International Engineering

Whatnot • Kraków

On-site
PLN 100,000 - 130,000
Staff Software Engineer, Ads
Staff Software Engineer, Ads

Whatnot • Kraków

Hybrid
PLN 903,000 - 1,069,000
Health Insurance
Work From Home
Home office allowance
+7
Lead Test Automation Engineer
Lead Test Automation Engineer

Expereo • Katowice

Hybrid
PLN 120,000 - 160,000
Hybrid remote model
Competitive compensation
Local benefits
DevOps Engineer
DevOps Engineer

Duco Technology Ltd • Wrocław

On-site
PLN 180,000 - 240,000
Competitive salary
Commission bonus
Private Medical Insurance - Enel Med
+7
Senior Software Engineer
Senior Software Engineer

Monaire • Poland

Remote
PLN 180,000 - 320,000
Competitive salary and equity
Comprehensive health insurance
Remote-first, flexible work culture
+1
Staff Core Platform Engineer
Staff Core Platform Engineer

n8n • Poland

On-site
PLN 300,000 - 520,000
Competitive pay
Equity
Remote-first environment
+4
Lead DevOps Engineer — Observability & CI/CD Champion
Lead DevOps Engineer — Observability & CI/CD Champion

Duco Technology Ltd • Wrocław

On-site
PLN 180,000 - 240,000