Staff Software Engineer (Infrastructure )

United States Digital Space LLC

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The company in San Francisco is seeking a Staff Software Engineer to go deep on the Nova OLAP engine and the surrounding infrastructure. You will design, implement, and optimize core query execution, storage formats, and distributed compute, while mentoring engineers and guiding architectural choices with product teams.

You will own performance, cost-efficiency, and reliability at scale, drive multi-month initiatives, and help the team ship robust data infrastructure that powers thousands of

Qualifications

  • 7+ years of industry experience in backend or infrastructure engineering.
  • Hands-on experience building or extending analytical/OLAP systems—query engines, columnar storage, large-scale data processing frameworks.
  • Track record of driving significant cost optimization on cloud infrastructure at scale (compute, storage, network).
  • Strong computer science fundamentals: distributed systems (partitioning, replication, consistency, failover), data structures and algorithms, concurrency and multi-threading, performance optimization.
  • Production experience with modern cloud infrastructure — AWS (S3, DynamoDB, EC2), Kafka, Redis/ElastiCache, Kubernetes, Terraform — or strong equivalents.
  • Proficiency in Java, C++, or Python.
  • Demonstrated technical influence beyond your immediate team: leading design discussions, driving cross-team alignment, mentoring engineers.

Responsibilities

  • Build and evolve core query engine infrastructure.
  • Design and implement capabilities as Nova expands to support more data types.
  • Design for high-throughput automated query workloads at scale.
  • Own and execute projects that reduce infrastructure cost while maintaining latency and throughput.
  • Profile and optimize JVM performance and memory management.
  • Build guardrails and observability to catch expensive or pathological queries.
  • Strengthen Nova’s reliability posture and participate in on-call rotation.
  • Contribute to capacity planning and operational tooling.
  • Lead design and execution of multi-month architecture projects.
  • Mentor senior engineers on distributed systems thinking and production debugging.
  • Collaborate with Product and Data teams to translate capabilities into customer value.

Skills

distributed systems thinking
mentoring engineers
technical leadership
communication of tradeoffs
cost optimization

Tools

AWS
Kafka
Redis/ElastiCache
Kubernetes
Terraform
Java
C++
Python

Job description

the company is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, the company is the best-in-class solution for product, data, and marketing teams. Learn more at the company.com.

As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do.

the company’s Commitment to Diversity Equity & Inclusion (DEI):

the company believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive.

About the Role & Team

Every AI insight, every experiment, every cohort at the company starts with a query. Our in-house OLAP engine, Nova, processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide.

We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non‑stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically.

We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost‑efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10.

This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers.

What You’ll DoBuild and evolve core query engine infrastructure
  • Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management.
  • Design and implement new capabilities as Nova expands to support more warehouse‑imported data types, such as metrics, profiles, and dimensions.
  • Design for high‑throughput automated query workloads — as AI agents become a primary source of queries, ensure Nova’s architecture supports sustained, concurrent, and programmatic query patterns at scale.
Drive cost and performance at scale
  • Own and execute projects that materially reduce infrastructure cost — compute, storage, network, and memory — while maintaining or improving latency and throughput.
  • Profile and optimize JVM performance: GC tuning, memory management, concurrency, and data layout decisions that compound at our scale.
  • Build guardrails and observability to catch expensive or pathological queries before they impact the system.
Improve reliability and operational excellence
  • Strengthen Nova’s reliability posture: identify systemic failure modes, drive durable fixes, and raise the bar on how we detect and respond to production issues.
  • Participate in on‑call rotation to root‑cause incidents and turn one‑off fixes into architectural improvements.
  • Contribute to capacity planning, safe rollout practices, and the operational tooling that keeps Nova healthy.
Influence through technical leadership
  • Lead the design and execution of multi‑month projects that improve Nova’s architecture, performance, or capabilities.
  • Contribute to technical direction through design docs, architecture discussions, and code reviews — helping the team make principled tradeoffs.
  • Mentor senior engineers on distributed systems thinking, production debugging, and system design.
  • Collaborate with Product, Middleware, Data Pipeline, and other engineering teams to ensure Nova’s capabilities translate into customer value.
Who You Are

You are an experienced systems engineer who:

  • Gets energy from working deep inside a complex distributed system — understanding how data flows through it, where the bottlenecks are, and how to make it meaningfully better.
  • Has built or significantly extended an OLAP engine, columnar database, query processor, or large‑scale data processing system — not just operated one.
  • Thinks about cost, performance, and reliability as interconnected concerns, not separate workstreams.
  • Communicates clearly about technical tradeoffs and earns influence through the quality of your work and ideas, not through title.
  • Finds it natural to help other engineers level up — through pairing, design reviews, or just being the person who explains the “why” behind a system’s design.
QualificationsRequirements
  • 7+ years of industry experience in backend or infrastructure engineering, with depth in distributed data systems.
  • Hands‑on experience building or extending analytical/OLAP systems — query engines, columnar storage, large‑scale data processing frameworks, or equivalent.
  • Track record of driving significant cost optimization on cloud infrastructure at scale (compute, storage, network).
  • Strong computer science fundamentals: distributed systems (partitioning, replication, consistency, failover), data structures and algorithms, concurrency and multi‑threading, performance optimization.
  • Production experience with modern cloud infrastructure — AWS (S3, DynamoDB, EC2), Kafka, Redis/ElastiCache, Kubernetes, Terraform — or strong equivalents.
  • Proficiency in Java, C++, or Python.
  • Demonstrated technical influence beyond your immediate team: leading design discussions, driving cross‑team alignment, mentoring engineers.
Nice to Have
  • Experience with specific OLAP or query engine systems: Druid, ClickHouse, Presto/Trino, BigQuery, Snowflake, or similar.
  • Deep JVM expertise — GC tuning, profiling, memory optimization at production scale.
  • Experience with columnar data formats and encodings (Arrow, Parquet, ORC, or custom formats).
  • Familiarity with product analytics, experimentation platforms, or event‑driven data systems.
  • Contributions to open‑source data infrastructure projects or published work in the data systems space.

Our values:

At the company, our values guide how we show up for one another and for our customers:

  • Humility: We operate from a place of empathy and openness, seeking to understand many points of view.
  • Ownership: We take
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer (Infrastructure )
Senior Software Engineer (Infrastructure )

United States Digital Space LLC • San Francisco (CA)

On-site
USD 160,000 - 260,000
Staff Software Engineer
Staff Software Engineer

Amplitude • San Francisco (CA)

On-site
USD 140,000 - 180,000
Staff Software Engineer — Nova OLAP & AI Platform
Staff Software Engineer — Nova OLAP & AI Platform

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer; Analytics Compute Platform team
Senior Software Engineer; Analytics Compute Platform team

United States Digital Space LLC • San Francisco (CA)

On-site
USD 170,000 - 250,000
Medical coverage
Parental benefits
Mental health benefits
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Socket.dev • New York (NY)

On-site
USD 140,000 - 210,000
Competitive salary
Meaningful equity
Health, dental, and vision coverage
+1
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Nova AI Software Inc. • New York (NY)

On-site
USD 150,000 - 210,000
Competitive salary
Meaningful equity
Health, dental, vision
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Nova Intelligence • New York (NY)

On-site
USD 150,000 - 190,000
Competitive salary
Equity
Health, dental, and vision coverage
+1
Senior Software Engineer
Senior Software Engineer

Novasoftware • New York (NY)

On-site
USD 100,000 - 150,000
Competitive compensation
Meaningful equity
Full health, dental, and vision coverage
Senior Software Engineer (Infrastructure )
Senior Software Engineer (Infrastructure )

Amplitude • San Francisco (CA)

On-site
USD 183,000 - 276,000
Excellent medical, dental and vision
401(k) with employer match
Flexible time off
+4
Senior Software Engineer; Analytics Infrastructure
Senior Software Engineer; Analytics Infrastructure

United States Digital Space LLC • San Francisco (CA)

On-site
USD 167,000 - 230,000
Health insurance
Flexible time off
Stipends for wellness/learning/home
+3