Member of Technical Staff, Principal Infrastructure Engineer

Edison Scientific, Inc.

New York (NY)

On-site

USD 200,000 - 350,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary and equity
Full healthcare coverage
Parental leave

Job summary

Edison Scientific, Inc. is seeking a Principal Member of Technical Staff to lead the platform infrastructure that powers autonomous scientific discovery.

You will design, scale, and operate Kubernetes-based systems, build CRDs and operators, and ensure reliable compute orchestration for AI agents and long-running workloads. Collaborating with backend, ML, and research teams, you will shape infrastructure best practices, drive scaling strategies, and implement observability and security across

Qualifications

  • 10+ years of infrastructure or platform engineering experience in production environments.
  • Deep hands-on Kubernetes expertise and experience with CRDs and operators.
  • Experience operating and scaling clusters with thousands of resources.
  • Strong knowledge of Kubernetes internals and cloud networking.
  • Proficiency in at least one systems language for operator development.
  • Hands-on with IaC tools and GitOps workflows.

Responsibilities

  • Architect, implement, and operate Kubernetes clusters for thousands of concurrent resources.
  • Design CRDs and operators to model long-running AI workloads and pipelines.
  • Drive cluster scaling strategies, node pools, autoscaling, and quotas.
  • Build and maintain infrastructure-as-code for reproducible environments.
  • Develop robust scheduling, placement, and affinity for heterogeneous workloads.
  • Establish observability, monitoring, and incident response for infra systems.
  • Own storage and networking strategy in Kubernetes, including CSI and network policies.
  • Troubleshoot complex distributed infrastructure issues and guide teams.
  • Collaborate with backend, ML, and research teams to translate workload needs into infra patterns.

Skills

Kubernetes expert
CRDs & Operators
Cluster scaling
IaC (Terraform, Pulumi)
GitOps
Cloud networking
Security (RBAC, PSP)
Programming for operators

Tools

Kubebuilder
Operator SDK
Terraform/Pulumi
Crossplane
Prometheus/Grafana/Datadog
CSI/Storage

Job description

About Us

Edison Scientific builds and deploys AI scientist agents to accelerate science and the development of new medicines. We are an ambitious team run by scientists and engineers from leading institutions across biology, physics, chemistry, and AI.

About The Role

As a Principal Member of Technical Staff you'll play a key role in designing, scaling, and operating the core platform infrastructure that powers autonomous scientific discovery. Your primary focus will be the orchestration for our agents at scale, building and managing clusters that orchestrate thousands of persistent, stateful workloads, developing custom resource definitions (CRDs) and operators, and ensuring the reliability and efficiency of our compute layer at scale.

Our mission is to build an AI scientist, and you'll own the infrastructure foundation it runs on. AI agents performing long-running scientific research demand resilient scheduling, lifecycle management, and resource orchestration far beyond typical cloud-native workloads. This role will influence platform architecture, establish infrastructure best practices, and partner closely with backend engineers, ML engineers, and researchers to deliver a production-grade environment that lets science move faster.

This position is part of the Platform Infrastructure team.

Key Responsibilities
  • Architect, implement, and operate Kubernetes clusters that support thousands of concurrent, persistent resources (agents, jobs, services) with high availability and efficient resource utilization.
  • Design and develop custom resource definitions (CRDs) and Kubernetes operators to model and manage domain-specific workloads such as AI agent lifecycles, research pipelines, and long-running compute tasks.
  • Drive the strategy for cluster scaling, node pool management, autoscaling policies, and resource quota frameworks to handle rapid workload growth.
  • Build and maintain infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible, version-controlled environment management.
  • Design and implement robust scheduling, placement, and affinity strategies to optimize cost, performance, and fault tolerance for heterogeneous workloads (CPU, GPU, memory-intensive).
  • Establish and uphold best practices around observability, monitoring, alerting, and incident response for infrastructure systems (Prometheus, Grafana, Datadog, or similar).
  • Own storage and networking strategy within Kubernetes - including persistent volume management, CSI drivers, service mesh, network policies, and ingress architecture.
  • Troubleshoot complex, cross-system infrastructure issues and guide others through effective debugging and remediation in distributed environments.
  • Collaborate closely with backend, ML, and research teams to understand workload requirements and translate them into reliable infrastructure patterns.
Required Qualifications
  • 10+ years of professional infrastructure or platform engineering experience, with deep hands-on Kubernetes expertise in production environments.
  • Experience designing and implementing custom resource definitions (CRDs) and Kubernetes operators (using frameworks such as Kubebuilder, Operator SDK, or controller-runtime).
  • Track record of operating and scaling Kubernetes clusters supporting thousands of persistent or long-lived resources (stateful workloads, persistent pods, long-running jobs).
  • Deep understanding of Kubernetes internals - API server, etcd, scheduler, controller manager, kubelet - and how they behave at scale.
  • Expertise with cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS) and associated networking, storage, and IAM primitives.
  • Proficiency in at least one systems or backend language for operator development and infrastructure tooling.
  • Hands‑on experience with infrastructure-as-code tools (Terraform, Pulumi, or Crossplane) and GitOps workflows.
  • Strong working knowledge of container networking (CNI plugins, service mesh, network policies), storage (CSI, persistent volumes, StatefulSets), and security (RBAC, Pod Security Standards, secrets management).
  • Ability to operate autonomously, make sound technical judgments, and drive projects from concept through production.
Preferred Qualifications
  • Experience with data‑intensive platforms, scientific computing, or ML/AI infrastructure.
  • Prior experience in startups or small teams with significant architectural ownership and ambiguity.
  • Experience scaling systems, teams, or platforms through periods of rapid growth.
Why join us?

We're a Fast-moving, Mission‑driven Culture Where Smart People Do Their Best Work And Actually Enjoy Doing It. We Also Offer Full‑time Employees The Following

  • Competitive salary and equity
  • Full healthcare coverage; we pay 100% of premiums for you and your dependents
  • Support for growing families, including a yearly new parent stipend and fertility coverage through Carrot
  • Mental health support through Rula, our in‑network therapist and psychiatrist network with fast availability
  • 12 weeks of paid parental leave for maternity, paternity, and adoption
  • Pet care support with a yearly employer-funded stipend for your animal companions
  • Commuter benefits so you can pay for transit and parking with pre‑tax dollars
  • 401(k) company matching
  • $300 health and wellness benefit quarterly
  • Lunch is on us every day you're in the office, and dinner is on us when you're working late
  • Regular team off‑sites and company events

Edison Scientific is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law. Compensation Range: $200K - $350K

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Senior Infrastructure Engineer
Member of Technical Staff, Senior Infrastructure Engineer

Edison Scientific, Inc. • New York (NY)

On-site
USD 175,000 - 240,000
Competitive salary & equity
Full healthcare coverage (premiums for
Parental leave & fertility coverage
+7
Member of Technical Staff, Infrastructure Engineer
Member of Technical Staff, Infrastructure Engineer

Edison Scientific, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 175,000 - 240,000
Equity
Healthcare
Parental leave
+6
Member of Technical Staff, Infrastructure Engineer
Member of Technical Staff, Infrastructure Engineer

Edison Scientific • San Francisco (CA)

On-site
USD 175,000 - 240,000
Health care coverage
Equity offered
Parental leave
Member of Technical Staff, Cloud Software Engineer
Member of Technical Staff, Cloud Software Engineer

Edison Scientific • San Francisco (CA)

On-site
USD 175,000 - 240,000
Full healthcare
Fertility coverage
Parental leave
+6
Member of Technical Staff, Full-stack Engineer
Member of Technical Staff, Full-stack Engineer

Edison Scientific, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 175,000 - 240,000
Competitive salary and equity
Full healthcare coverage
Parental leave stipend
+6
Member of Technical Staff, Applied AI Engineer, Agents
Member of Technical Staff, Applied AI Engineer, Agents

Edison Scientific • New York (NY)

On-site
USD 175,000 - 240,000
Competitive salary & equity
Full healthcare coverage
Parental leave
Member of Technical Staff, Agent Engineer
Member of Technical Staff, Agent Engineer

Edison Scientific • New York (NY)

On-site
USD 175,000 - 240,000
Competitive salary
100% healthcare coverage
Parental/family support stipend
+3
Member of Technical Staff, Applied AI Engineer
Member of Technical Staff, Applied AI Engineer

Edison • San Francisco (CA), New York (NY)

On-site
USD 150,000 - 210,000
Competitive salary & equity
Full healthcare coverage
Parental leave stipend
+6
Member of Technical Staff, Cloud Software Engineer
Member of Technical Staff, Cloud Software Engineer

Edison • San Francisco (CA)

On-site
USD 170,000 - 250,000
Competitive salary & equity
Full healthcare coverage
Parental leave (12 weeks)
+7
Engineering Manager, Product Engineering
Engineering Manager, Product Engineering

Edison Scientific, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 250,000 - 350,000
Competitive salary and equity
Full healthcare coverage
Parental leave (12 weeks)
+5