**Role overview**
An internal engineering platform team is hiring a Staff Platform Engineer to lead the software and infrastructure that enables product engineering teams to build, deploy, and operate services at scale. The role combines deep hands-on software engineering with cross-team technical leadership, focused on turning infrastructure, delivery, and reliability work into reliable, self-service platform products. The mission is to make engineering faster, safer, and more scalable by replacing manual operations with durable software and clear defaults.
**Responsibilities**
- Set technical direction and contribute to a long-term roadmap across cloud, networking, compute, storage, software delivery, and reliability engineering.
- Lead architecture reviews and technical engagements across the software development lifecycle, from early design through production operation.
- Design, build, and maintain platform services, APIs, CLIs, controllers, workflows, and automation in languages such as Go, TypeScript, Python, or similar, owning them through design, implementation, testing, deployment, and ongoing production support.
- Build self-service capabilities that let engineering teams provision environments, deploy services, manage infrastructure, and respond to operational conditions without manual platform intervention.
- Improve CI/CD and automated software delivery systems while increasing deployment safety and developer velocity.
- Measure and reduce operational toil, coordinate response to significant incidents, run effective postmortems, and turn recurring issues into permanent reliability improvements.
- Influence engineering standards across teams through design reviews, technical writing, prototypes, internal education, and mentorship.
**Requirements**
- 8+ years of professional experience in software engineering, infrastructure engineering, site reliability engineering, or an equivalent combination, including substantial ownership of production systems.
- Strong software engineering ability in at least one language such as Go, Python, TypeScript, or Rust, with experience building maintainable services, CLIs, APIs, or developer tools.
- Hands-on experience operating distributed systems in production, including failure diagnosis, capacity management, performance improvement, and high-availability design.
- Practical experience with a major cloud provider (AWS, GCP, or Azure) and infrastructure-as-code tooling such as Terraform or an equivalent system.
- Experience with containers and orchestration such as Kubernetes, EKS, GKE, AKS, Docker, or Nomad, including workload deployment, networking, security, and operational troubleshooting.
- Experience owning cloud network security controls, including firewalls, DNS filtering, and network segmentation.
- Track record of setting technical direction, leading architecture across multiple teams, and influencing decisions without relying on direct authority.
**Nice to have**
- Build tooling experience such as Bazel or Buck2.
- Background in capacity planning, failure recovery, and the operational tradeoffs of running critical systems at scale.
- Experience building secure-by-default systems and collaborating on identity, access control, secrets, supply-chain security, and policy enforcement.
- Ability to learn unfamiliar domains quickly, including the infrastructure characteristics of data-intensive, event-driven, or AI/ML workloads.
**Benefits and work setup**
- Base salary range of $168,000 to $240,000 for New York, with a discretionary annual bonus and a new-hire equity grant.
- Comprehensive health plans, 401(k) with company match, paid parental leave, and flexible time off.
- Hybrid work model at hub office locations, with a remote workforce option for employees outside hub cities and required in-person onboarding.