Infrastructure Engineer, TL

Arena Intelligence, Inc.

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation and equity
Health benefits
Cutting-edge AI work
Transparent culture

Job summary

Arena Intelligence is seeking an experienced backend engineer to build the core infrastructure behind our online evaluation systems. You will design low-latency APIs for leaderboards, models, and arenas, and help scale the AI gateway stack across providers.

Collaborate with researchers and product leadership to ship enterprise-grade infrastructure with observability, security, and governance. You’ll contribute across the backend for Leaderboards and Evals, driving reliability and developer

Qualifications

  • 4+ years of backend engineering experience focusing on distributed systems
  • Proficiency in Go and/or Rust with API or proxy system experience
  • Experience with LLM provider APIs and challenges: streaming, token management, rate limits
  • Solid cloud skills: AWS or GCP, Kubernetes, Terraform, Postgres/Redis
  • Product-minded: focus on developer experience of APIs
  • Comfort with ambiguity in a startup environment

Responsibilities

  • Build API-based products from the ground up for leaderboards, models, and arenas
  • Solve streaming problems and ensure graceful recovery and consistent responses
  • Ship enterprise-grade infrastructure: rate limiting, authentication, metering, audit logging, SOC 2
  • Build observability with tracing, latency breakdowns, usage tracking, dashboards
  • Integrate with Arena data and benchmarks, turn ideas into products
  • Contribute across the backend stack for Leaderboards and Evals as needed

Skills

Backend engineering
Distributed systems
Go/Rust
Cloud infrastructure
Product mindset
Ambiguity tolerance

Tools

Bifrost
Kong
Envoy
Tyk
Custom gateways

Job description

About Arena Intelligence

Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.

Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.

We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.

About the Role

Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.

This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.

You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.

What You’ll Do
  • Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.

  • Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.

  • Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.

  • Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.

  • Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.

  • Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.

You’ll have
  • 4+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.

  • Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.

  • Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.

  • Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.

  • A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."

  • Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.

Nice to Have
  • Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).

  • Background in ML infrastructure, model serving, or evaluation frameworks.

  • Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.

  • Experience building billing infrastructure around systems like Stripe, Metronome and Orb

  • Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).

What we offer
  • We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.

  • Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.

  • The opportunity to work on cutting-edge AI with a small, mission-driven team

  • A culture that values transparency, trust, and community impact

Come help build the space where anyone can explore and help shape the future of AI.

Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer – TL
Infrastructure Engineer – TL

Arena • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation
Health and wellness benefits
Cutting-edge AI projects
+1
Site Reliability Engineer
Site Reliability Engineer

Arena Intelligence, Inc. • United States

On-site
USD 140,000 - 210,000
Equity
Health benefits
Cutting-edge AI
+1
Site Reliability Engineer
Site Reliability Engineer

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Equity
Health benefits
+1
Site Reliability Engineer
Site Reliability Engineer

Arena • San Francisco (CA)

On-site
USD 180,000 - 280,000
Equity
Health benefits
Cutting-edge AI work
+1
Infrastructure Engineer
Infrastructure Engineer

People Culture Talent • San Francisco (CA)

On-site
USD 200,000 - 350,000
Comprehensive health, dental, vision benefits
Opportunity to work on cutting-edge AI
Culture of transparency and community impact
+1
Backend Engineer
Backend Engineer

Arena Intelligence, Inc. • United States

On-site
USD 120,000 - 160,000
Health insurance
Equity
Cutting-edge AI work
Backend Engineer
Backend Engineer

Arena • San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Health benefits
Cutting-edge AI work
+1
Senior Software Engineer, ML Infrastructure
Senior Software Engineer, ML Infrastructure

Arena • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive salary
Equity
Health benefits
+2
Senior Infra Engineer — Real-Time AI Data & API
Senior Infra Engineer — Real-Time AI Data & API

Arena • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer, Product
Staff Software Engineer, Product

arena • San Francisco (CA)

On-site
USD 130,000 - 160,000
Competitive compensation
Health and wellness benefits
Opportunity to work on cutting-edge AI