Founding Platform & Reliability Engineer

Embedding VC

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive base salary
Equity opportunities
High autonomy environment
Visa sponsorship available

Job summary

A leading AI-native platform company is seeking a Founding Platform & Reliability Engineer to oversee the design and reliability of their entire infrastructure. This role demands a mix of hands-on implementation and strategic architecture decisions, ensuring system reliability and cost efficiency in a rapidly evolving environment. Key qualifications include over 5 years of experience in production systems and strong engineering skills in cloud-native environments. Enjoy competitive compensation and a hybrid work setup in the Bay Area.

Qualifications

  • 5+ years building and operating production systems where reliability and scaling are core.
  • Comfortable working across cloud infrastructure and distributed systems.
  • Ability to operate with ambiguity and define problems before solving them.
  • Deep knowledge of observability practices: dashboards, alerting, tracing, incident response maturity.
  • Ability to design resilient interactions with external dependencies and communicate tradeoffs.

Responsibilities

  • Define and operationalize SLOs/SLIs across critical user journeys.
  • Implement reliability patterns at external boundaries.
  • Act as a senior technical voice influencing architecture and best practices.
  • Stand up end-to-end observability: logs, metrics, traces, dashboards to answer what broke and why now.
  • Own the direction of our infrastructure architecture, including serverless vs containerized decisions and scaling.

Skills

Cloud-native experience (AWS or GCP)
Strong software engineering skills
Deep knowledge of observability practices
Ability to design resilient interactions
Can communicate tradeoffs to peers
Tradeoffs communication

Tools

GCP
Node.js
TypeScript
Python
React / Next.js

Job description

Founding Platform & Reliability Engineer
🎨 About OpenArt

OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We’re building the next generation of creative tools powered by cutting-edge AI, enabling anyone to create videos, visuals, characters, and stories with unprecedented speed and imagination. We believe the future of creativity is AI-native, and we're shaping that future.

🚀 Why Join OpenArt
  • Small team, massive surface area, senior engineers own real systems, not slices.

  • Ship at real scale, your work goes to millions of users, fast.

  • Founder-led engineering culture, both founders are technical and deeply involved in product and architecture.

  • AI-native product, you’ll design how cutting-edge AI models are exposed as real user experiences.

  • High ownership, low process, we value judgment, clarity, and speed over bureaucracy.

  • 7-10X growth in revenue for the past 2 years. Now you’ll play a critical role in helping the company scale to the next stage.

🎯 About the Role

We’re looking for a Founding Platform & Reliability Engineer who can own the design, scalability, and reliability of our entire infrastructure stack end-to-end, from high-level architecture decisions to hands-on implementation, observability, and cost optimization.

This is NOT a role for traditional operators or narrow DevOps specialists. You should be comfortable working across cloud infrastructure, distributed systems, backend services, and developer tooling, making pragmatic decisions that balance product velocity, system reliability, and cost efficiency—especially in a fast-evolving, AI-native environment.

You will work closely with the founders and product engineers to design and evolve the platform that powers OpenArt, shaping key decisions such as serverless vs. containerized architecture, multi-provider AI reliability, and scaling systems to millions of users—while acting as a force multiplier for the entire engineering team.

🛠 What You’ll Do
  • Define and operationalize SLOs/SLIs across critical user journeys (generation, editing, payments/credits, uploads, etc.), and use them to drive prioritization (including error budgets)

  • Participate in an on-call rotation and lead incident response improvements (alert quality, runbooks, escalation paths). Establish blameless postmortems and ensure action items are implemented.

  • Implement reliability patterns at external boundaries, and build mechanisms for per-vendor “health” measurement and routing/fallback policies

  • Stand up end-to-end observability: structured logs, metrics, traces, and dashboards that let engineers answer “what broke” and “why now” quickly.

  • Build deploy safety practices: automated rollbacks, canarying, feature-flag patterns, and reliable CI/CD gates.

  • Own the direction of our infrastructure architecture, including defining when serverless is the right approach versus when we should evolve toward containerized or more managed systems, and guiding the team through those transitions as we scale.

  • Build cost observability and cost-control primitives: per-request cost attribution, caching strategies, capacity planning, and budget alerts.

  • Act as a senior technical voice, influencing architecture, tooling, engineering best practices, and raising the overall engineering bar.

🧑💻 What We’re Looking For

Core Requirements

  • 5+ years building and operating production systems where reliability and scaling are core.

  • Strong software engineering skills (you can ship production code, not just configure tools).

  • Cloud-native experience (AWS or GCP), ideally with serverless/event-driven systems and at least one container path (Fargate/ECS/Cloud Run/Kubernetes).

  • Deep knowledge of observability practices: dashboards, alerting, distributed tracing, and incident response maturity.

  • Ability to design resilient interactions with external dependencies (timeouts, retries/backoff/jitter, circuit breakers, idempotency).

  • Can communicate tradeoffs to non-infra peers clearly

  • Ability to operate with ambiguity and define problems before solving them.

Nice to Have

  • Have designed an internal platform abstraction (e.g., API gateway / workflow engine / job orchestration) that enabled multiple product teams to ship faster with fewer incidents.

  • Have shipped concrete reliability outcomes: e.g., reduced MTTR, improved SLO attainment, lowered p95 latency, or reduced infra/unit costs

  • Prior startup experience or experience owning large surface-area features.

⚙ Tech Stack You’ll Work With

GCP, Cloud Run, Modal, Upstash, Sentry, Amplitude, Firebase, Redis, React / Next.js, Node.js, TypeScript, Python, etc.

đź’° Compensation
  • Competitive base salary and bonus program

  • Equity - meaningful ownership in what you build

  • High autonomy, high growth environment

🌍 Work Setup
  • Bay Area preferred (hybrid allowed)

  • Visa sponsorship available

  • We’ll consider remote

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Data Engineer (Core Data Platform)
Founding Data Engineer (Core Data Platform)

SwiftCruit • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Competitive base salary
Equity options
High autonomy
Senior Full-Stack Product Engineer (NA/Remote)
Senior Full-Stack Product Engineer (NA/Remote)

Embedding VC • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Meaningful equity ownership
High-autonomy work environment
Visa sponsorship available
Senior Product Engineer (Full-Stack)
Senior Product Engineer (Full-Stack)

OpenArt • United States

Hybrid
USD 250,000 - 350,000
Meaningful ownership in what you build
High autonomy and growth environment
Visa sponsorship available
Growth Engineer - Globalization
Growth Engineer - Globalization

Embedding VC • San Francisco (CA)

Hybrid
USD 300,000 - 400,000
Growth Engineer - Globalization
Growth Engineer - Globalization

Socket.dev • San Francisco (CA)

Hybrid
USD 300,000 - 400,000
Senior Product Engineer (North America)
Senior Product Engineer (North America)

Embedding VC • San Francisco (CA)

On-site
USD 250,000 - 350,000
Meaningful equity ownership
High-autonomy work environment
Visa sponsorship available
Growth Data Engineer
Growth Data Engineer

Embedding VC • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Bonus program
Equity ownership
Product UI/UX Designer
Product UI/UX Designer

Embedding VC • San Francisco (CA)

Hybrid
USD 110,000 - 170,000
AI Social Content Creator (full-time)
AI Social Content Creator (full-time)

Embedding VC • San Francisco (CA)

Hybrid
USD 70,000 - 90,000
Remote-friendly work environment
Hybrid options
Visa sponsorship available
Brand Marketing Manager
Brand Marketing Manager

OpenArt AI • San Francisco (CA)

On-site
USD 90,000 - 130,000
Competitive base salary and bonus
Equity
High autonomy, high growth environment