Member of Technical Staff - Platform Engineering

Creandum

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Modal is building an open cloud infrastructure layer. We seek seasoned engineers to boost reliability while scaling the platform, customers, and team.

You will shape deployment processes, manage on-call rotations, and drive operational excellence across Kubernetes, Postgres, and Redis in a rapidly growing environment. Strong production coding skills and cloud expertise are essential, with in-person work possible in NYC or Stockholm.

Qualifications

  • 5+ years of experience writing high-quality production code.
  • 2+ years of on-call experience for critical production services.
  • Strong cloud skills, and deep familiarity with at least one hyperscaler (AWS preferred).
  • Familiarity with auto scaling, fleet management, and capacity planning at scale.
  • Experience operating databases, monitoring, CI/CD, and other infrastructure at scale.
  • Experience owning and scaling Kubernetes clusters to thousands of nodes is a plus.

Responsibilities

  • Identify architectural changes to improve reliability and performance.
  • Foster a culture of reliability across the engineering organization.
  • Define and implement deployment, upgrade, and other operational processes.
  • Operate systems like Kubernetes, Postgres, Redis, etc.
  • Participate in on-call rotations and respond to production incidents.

Skills

Production-grade code
On-call experience
Cloud computing (AWS)
Kubernetes
Scalability & capacity planning

Tools

Kubernetes
Postgres
Redis

Job description

About Us:

AI needs a new infrastructure layer. We're building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

At Modal, we sell cloud services atop which our customers run their critical production systems. As a rapidly growing new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform, customer base, and our team.

This role is for people who are deep systems thinkers, love stacking nines, and thrive from making others move faster at scale. Responsibilities include:

  • Identifying architectural changes to improve reliability and performance.

  • Fostering a culture of reliability across Modal’s engineering organization.

  • Defining and implementing operational processes such as deployments, upgrades, etc.

  • Operating systems like Kubernetes, Postgres, Redis, etc.

  • Participating in on-call rotations, and responding to production incidents.

Requirements:
  • 5+ years of experience writing high-quality production code.

  • 2+ years of on-call experience for critical production services.

  • Strong cloud skills, and deep familiarity with at least one hyperscaler cloud (AWS preferred).

  • Familiarity with auto scaling, fleet management, and capacity planning at scale.

  • Experience operating databases, monitoring, CI/CD, and other infrastructure, at scale

  • Experience owning and scaling Kubernetes clusters to thousands of nodes a plus.

  • Experience with systems safety research (e.g. STAMP) and control theory a plus.

  • Ability to work in-person in our NYC or Stockholm offices.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Platform Engineering
Member of Technical Staff - Platform Engineering

Modal • New York (NY)

On-site
USD 150,000 - 190,000
Member of Technical Staff - Platform Engineering
Member of Technical Staff - Platform Engineering

EuroPython • Town of Sweden (NY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Systems
Member of Technical Staff - Systems

modal • New York (NY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - SDK
Member of Technical Staff - SDK

modal • New York (NY)

On-site
USD 140,000 - 190,000
Customer Engineer
Customer Engineer

Modal Labs • New York (NY)

On-site
Confidential
Member of Technical Staff - Machines
Member of Technical Staff - Machines

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Machines
Member of Technical Staff - Machines

Linuxconfig • Northern (KY), New York (NY)

On-site
USD 180,000 - 230,000
Customer-Facing AI Infra Engineer — Ship Fixes & Automations
Customer-Facing AI Infra Engineer — Ship Fixes & Automations

Modal Labs • New York (NY)

On-site
Confidential
Member of Technical Staff - Product (Backend)
Member of Technical Staff - Product (Backend)

Modal Labs • New York (NY)

On-site
USD 140,000 - 190,000
Member of Technical Staff - Storage
Member of Technical Staff - Storage

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote-friendly options