Senior Production Engineer

Whatnot

New York (NY)

Hybrid

USD 207,000 - 290,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flexible time off
Health insurance
WFH support
Care benefits
401k match
Parental leave
Dogfooding budget

Job summary

Whatnot is hiring a Production Engineer to embed with product, platform, and infrastructure teams. You will hunt anomalies, improve reliability, and own fixes across live video and real-time commerce systems.

The role emphasizes deep debugging, capacity planning, and strong collaboration in a fast-paced, remote-friendly environment. You will work across payments, search, and the bidding path, focusing on reducing toil and driving top-line performance while supporting peak events and critical

Qualifications

  • Bachelor's degree in Computer Science or related field or equivalent work experience.
  • 6+ years building and debugging production services at scale.
  • Software engineer identity with focus on production systems and reliability.
  • Strong systems and distributed systems fundamentals: failure modes, saturation, queueing, cascading failures.
  • Depth in observability and production debugging; experience in capacity/performance engineering, load testing, incident command, or large migrations.
  • Ability to work well embedded across teams and trust-building.
  • Fluency across languages and cloud-native environments (AWS/GCP) with Kubernetes and IaC; backend in Python, Elixir, Go.
  • Bonus: experience with high-traffic, real-time, event-driven, or live streaming systems.

Responsibilities

  • Go deep with the embedded team to improve reliability, performance, and scalability.
  • Hunt anomalies across traffic, latency, error, and cost signals and trace root causes.
  • Connect signals to business impact and determine what matters for key cohorts.
  • Work on risk-prone areas: payments, live video, search, bidding, and security.
  • Eliminate scale bottlenecks in production and address systemic issues.
  • Prepare for peak events with capacity modeling, load testing, and drills.
  • Improve observability to detect the next issue quickly.
  • Apply AI to operations for anomaly detection and toil reduction.
  • Share on-call with the embedded team and lead incident response.

Skills

Distributed systems
Production engineering
Observability
Incident response
Capacity planning
Python
Elixir
Go
Cloud-native (AWS/GCP)

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes

Job description

Join the Future of Commerce with Whatnot! Whatnot is the largest live shopping platform in North America and Europe to buy, sell, and discover the things you love. Whether it's trading cards, fashion, electronics, or live plants, our sellers are building real businesses across hundreds of categories. We're building live commerce at a scale that's never been done in the West, and there's no playbook to copy. The people here are shaping how an entirely new industry develops. As a remote co-located team, we're inspired by our values and anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia. We move fast, stay close to our users, and focus on the work that drives the most impact. We're one of the fastest growing marketplaces and were recently named the #1 Best Startup Employer in America by Forbes. Check out the latest Whatnot updates on our news and engineering blogs and join us as we enable anyone to turn their passion into a business and bring people together through commerce.

Role

That growth is the reason this role exists. Production Engineers are software engineers who embed with product, platform, and infrastructure teams to find and fix what breaks at scale. We think of them as hunters: engineers who go looking for the anomaly, the bottleneck, the slow degradation nobody has noticed yet, and then own it through to a fix, working with whoever owns the system to make sure it stays fixed. Live video and real-time commerce run on the same critical path here. Every auction is live and money moves during the show, so system behavior and business outcomes are tightly coupled in a way few platforms experience. Milliseconds are visible to buyers and sellers, and the signals that matter most are often small, concentrated in a specific cohort, and invisible in aggregate. Surfacing them early, and knowing which ones are worth acting on, is the core of this work. Production Engineers embed with a team and go deep on its systems, while staying accountable for what falls between teams. Some of the most consequential problems at our scale live in the seams, in the interaction between two services that each look healthy on their own, so we expect hunting both inside your area and across it. You'll be joining early, which means broader discovery across the platform before the embedding settles, and real influence over how Production Engineering takes shape here.

What You'll Do
  • Go deep with the team you embed with, raising the reliability, performance, and scalability of their systems alongside them and writing the code that gets it there.
  • Hunt anomalies across traffic, latency, error, and cost signals, both inside your area and in the seams between teams, and trace them to root cause across services, storage, and clients.
  • Connect operational signals to business impact: which degradations matter, for which cohorts, and which are safe to leave alone.
  • Work where the risk concentrates: payments, live video, search, the bidding path, and security.
  • Eliminate scale bottlenecks in production, then remove the class of problem rather than the instance.
  • Prepare for peak events through capacity modeling, load testing, and failure drills.
  • Improve observability where it's thin, so the next anomaly is found in minutes rather than quarters.
  • Put AI to work on operations: anomaly detection, on-call assistance, remediation, toil reduction. Mostly greenfield, and yours to prove out.
  • Share on-call with the team you embed with, and act as an escalation point for live production incidents.
  • Lead incident response for complex cross-team failures and drive the systemic fixes that follow.

We offer flexibility to work from home or from one of our global office hubs, and we value in-person time for planning, problem-solving, and connection. Team members in this role must live within commuting distance of our San Francisco or New York City hub.

You

People who do well at Whatnot tend to be comfortable figuring things out as they go, biased toward action, and genuinely curious about what they're building. They care more about outcomes than credit and stay close to the product and the people using it.

As Our Next Production Engineer, You Likely Have
  • Bachelor's degree in Computer Science, a related field, or equivalent work experience.
  • 6+ years building and debugging production services at scale. You may have come to that through software engineering, production engineering, SRE, or systems engineering; what matters is the depth, not the title.
  • A software engineer's identity first, paired with a pull toward the messy end of production. You want to write code and build systems, and you're the one who notices the graph that looks slightly wrong and can't let it go.
  • Strong systems and distributed systems fundamentals: failure modes, saturation, queueing behavior, cascading failure, and how Linux, networking, and storage behave under load.
  • Depth in observability and production debugging, plus experience in one or more of: capacity and performance engineering, load and resilience testing, incident command, or large-scale migrations under production constraints.
  • The ability to work well embedded. You can walk into another team's codebase, earn trust quickly, and leave the system better than you found it.
  • Fluency across languages and technologies, and comfort in cloud-native environments such as AWS or GCP with Kubernetes and infrastructure as code. Our backend is primarily Python and Elixir, with Go for performance-sensitive infrastructure, and we care more about depth and adaptability than a specific stack.
  • Bonus: high-traffic, real-time, event-driven, or live streaming systems.
Benefits
  • Flexible Time off Policy and Company-wide Holidays (including a spring and winter break)
  • Health Insurance options including Medical, Dental, Vision
  • Work From Home Support
    • Home office setup allowance
    • Monthly allowance for cell phone and internet
  • Care benefits
    • Monthly allowance for wellness
    • Annual allowance towards Childcare
    • Lifetime benefit for family planning, such as adoption or fertility expenses
  • Retirement; 401k offering for Traditional and Roth accounts in the US (employer match up to 4% of base salary) and Pension plans internationally
  • Monthly allowance to dogfood the app
    • All Whatnauts are expected to develop a deep understanding of our product. We're passionate about building the best user experience, and all employees are expected to use Whatnot as both a buyer and a seller as part of their job (our dogfooding budget makes this fun and easy!).
  • Parental Leave
    • 16 weeks of paid parental leave + one month gradual return to work *company leave allowances run concurrently with country leave requirements which take precedence.
EOE

Whatnot is proud to be an Equal Opportunity Employer. We value diversity, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, parental status, disability status, or any other status protected by local law. We believe that our work is better and our company culture is improved when we encourage, support, and respect the different skills and experiences represented within our workforce.

Compensation Range: $207K - $290K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Production Engineer
Senior Production Engineer

Whatnot • San Francisco (CA)

Hybrid
USD 170,000 - 230,000
Flexible Time off
Health Insurance
Work From Home support
+3
Senior Production Engineer
Senior Production Engineer

Whatnot Inc. • New York (NY), Northern (KY)

Hybrid
USD 170,000 - 210,000
Flexible time off
Health insurance
Work-from-home support
+2
Software Engineer, 2027 New Grad
Software Engineer, 2027 New Grad

Whatnot • New York (NY)

Hybrid
USD 155,000 - 160,000
Health Insurance
Work From Home Support
401k plan
+2
Data Engineer, Payments
Data Engineer, Payments

Whatnot • Seattle (WA)

Hybrid
USD 180,000 - 260,000
Health Insurance
Work From Home Support
Home office setup allowance
+3
Fullstack Engineer, App Platform
Fullstack Engineer, App Platform

Whatnot • San Francisco (CA)

Hybrid
USD 207,000 - 230,000
Generous time off
Health insurance (Medical, Dental, Vis
WFH support
+8
Software Engineer, Trust Experience
Software Engineer, Trust Experience

Whatnot • San Francisco (CA)

On-site
USD 170,000 - 230,000
Health Insurance options including Medical, Dental, Vision
Generous Holiday and Time off Policy
Home office setup allowance
+2
Software Engineer, Enterprise Support
Software Engineer, Enterprise Support

Whatnot • Seattle (WA)

Hybrid
USD 170,000 - 230,000
Health Insurance
Work From Home Support
Home office setup allowance
+6
Software Engineer, Enterprise Support
Software Engineer, Enterprise Support

Whatnot • San Francisco (CA)

Hybrid
USD 170,000 - 230,000
Flexible time off
Health insurance
WFH support
+3
Engineering Manager, Customer Experience
Engineering Manager, Customer Experience

Whatnot • New York (NY)

Hybrid
USD 180,000 - 260,000
Health Insurance
WFH Support
Home office setup
+6
Fullstack Engineer, Ads
Fullstack Engineer, Ads

Whatnot • New York (NY)

Hybrid
USD 207,000 - 230,000
Health Insurance options
Work From Home Support
Home office setup allowance
+4