The Platform SRE is a member of the Platform SRE team, the group responsible for the operational health, resiliency, and modernization of our hosting and cloud infrastructure. This role sits at the intersection of systems engineering, site reliability engineering, and automation : increasingly AI-augmented automation : combining hands-on technical debt remediation with large-scale automation and reliability practice.
The Platform SRE team owns three core mandates
1. Technical debt reduction :
systematically identifying and remediating aging firmware, kernels, operating systems, and software across the managed hosting fleet and managed application environments
2. Automation of critical operational workflows :
provisioning, patching, remediation, and release processes across both managed apps and managed hosting fleets, with automated release planning and execution under human supervision (not “automation for automation’s sake,” but automation with a human checkpoint before production impact)
3. Reliability and incident support :
defining, instrumenting, and tracking SLIs/SLOs for platform engineering and operations, visualized through dashboards and reporting tools, and providing fast, expert frontline response and remediation during service-impacting events (incident command and process ownership sit with the Incident Management team; this team is the technical responder, not the incident owner).
4. AI-augmented operations :
operating agentic AI SRE tooling for proactive anomaly detection, automated investigation, and human-gated remediation across daily operations, so the team scales its impact without over-scaling headcount with the fleet.
This position serves as a senior technical point of contact for platform engineering, driving initiatives around scalability, fault tolerance, automation, and operational excellence across production infrastructure.
Key Responsibilities
Technical Debt & Platform Modernization
- Own the lifecycle of firmware, kernel, OS, and software patching across the managed hosting fleet and managed application environments
- Build a standing inventory and risk model of technical debt (end-of-life OS versions, unpatched firmware, deprecated software) and drive prioritized remediation plans
- Evaluate and implement infrastructure modernization initiatives, replacing manual or legacy processes with supportable, automated alternatives
Automation & Release Engineering
- Design and build automation for provisioning, deployment, patching, remediation, and configuration management across managed apps and managed hosting fleets
- Own the design of automated release pipelines : planning, staging, and executing releases with defined human-in-the-loop approval gates
- Develop self-healing and auto-remediation capability for common failure modes to reduce manual operational load
- Support and extend CI/CD workflows and infrastructure-as-code practices across the platform
Reliability Engineering, SLIs/SLOs & Observability
- Define SLIs and SLOs for platform engineering and operations in partnership with engineering and product stakeholders
- Instrument systems to measure SLIs accurately and build SLO tracking into standard reporting
- Build and maintain dashboards (e.g., Grafana, Datadog, or equivalent visualization tooling) to make SLI/SLO performance, error budgets, and platform health visible to engineering and leadership
- Continuously improve platform observability : monitoring, alerting, logging, and tracing : across distributed and containerized environments
- Serve as the frontline technical responder: acknowledge pages quickly, diagnose, and remediate platform-level issues
- Partner with the Incident Management team, who own incident command, severity classification, and customer communication : this role provides the technical hands and expertise, not incident ownership
- Contribute technical findings to blameless root cause analysis (RCA) and own follow-through on corrective actions for platform systems
- Maintain runbooks and on-call readiness for platform and infrastructure systems
- Track incident trends on platform systems and feed them back into the technical debt and automation roadmap
AI-Augmented Operation
- Operate and tune agentic AI SRE tooling (e.g., Harness AI SRE, HolmesGPT, K8sGPT, or equivalent platforms) for proactive anomaly detection, automated investigation, and root-cause drafting across the fleet
- Apply AIOps-style alert correlation and noise reduction to cut duplicate/low-value pages and protect on-call sustainability
- Use AI-assisted runbook and chaos-engineering tooling to convert manual procedures into self-executing, testable workflows
- Maintain the human-in-the-loop gate on all AI-suggested or AI-generated remediation before it reaches production : AI drafts and proposes, this role verifies and approves
- Evaluate new AI SRE tooling for fit, accuracy, and safety before adoption; retire tools that don’t earn their keep
Collaboration & Technical Leadership
- Partner with software engineering teams on platform architecture, operational readiness reviews, and scalability initiatives
- Support platform security, compliance, and operational governance requirements
- Mentor engineers and contribute to technical leadership and knowledge-sharing across the team
- Maintain clear operational documentation and contribute to team standards and process improvement
- Other duties as assigned
Requirements
- 3–5+ years of experience in platform engineering, systems engineering, SRE, or infrastructure operations (level based on experience and scope)
- Advanced Linux systems administration and troubleshooting expertise, including kernel and firmware-level familiarity
- Strong experience with Kubernetes, Docker, and container orchestration/distributed systems
- Hands-on automation and infrastructure-as-code experience (e.g., Terraform, Ansible, Puppet/Chef, or equivalent)
- Familiarity with agentic AI SRE/AIOps tooling for proactive detection, investigation, and human-gated remediation (e.g., Harness AI SRE, HolmesGPT, K8sGPT, or equivalent platforms)
- Experience building or maintaining CI/CD and automated release/deployment pipelines
- Experience defining and tracking SLIs/SLOs and working with observability/visualization tools (e.g., Grafana, Datadog, Prometheus, or equivalent)
- Experience supporting enterprise-scale, high-concurrency, or customer-impacting production environments
- Demonstrated experience as a technical responder in production incidents, including root cause analysis and corrective action follow-through
- Strong scripting ability (e.g., Python, Bash, Go) for automation and tooling
- Strong troubleshooting skills across compute, network, storage, and application layers
- Experience supporting cloud-hosted, managed hosting, or hybrid infrastructure environments
- Ability to lead technical initiatives, mentor others, and communicate clearly across teams
Preferred Qualifications
- Experience owning fleet-wide firmware/OS patch management programs at scale
- Experience designing human-in-the-loop release automation or progressive delivery systems (canary, blue/green)
- Experience operating or tuning AI-driven observability/AIOps platforms (alert correlation, autonomous RCA drafting, AI-assisted chaos engineering test generation)
- Experience evaluating or piloting emerging AI SRE agent platforms and setting guardrails for safe adoption
- Familiarity with error budgets and SLO-driven prioritization frameworks
- Experience with configuration/patch management at scale across heterogeneous hardware fleets
- The physical demands described here are representative of those that must be met by an individual to successfully perform the essential duties of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential duties.
- While performing the essential duties of this job, the individual is regularly required to speak and hear
- Works at a desk and computer screen for extended periods of time
- Works in a traditional climate-controlled office environment or from home
- Works in a highly stressful environment dealing with a wide variety of challenges, deadlines, and diverse employee population
- Require participation in an on-call rotation, including responding to incidents outside standard business hours