6905315
Location: Pleasanton,CA, US
Date Posted: 08-31-2026
Job Description
Experience: Senior Level
Salary: $300,000 - $350,000 per year
Job Details
---------
Responsibilities
- Implement and optimize an AI-powered Site Reliability Engineering (SRE) platform to meet customer needs across production and pre-production environments.
- Proactively monitor customer deployments to ensure customers maximize value from the platform.
- Identify latent reliability issues such as misconfigurations, deployment regressions, and scaling challenges within customer environments.
- Recommend best practices for implementing AI-powered SRE solutions.
- Plan, design, build, and maintain highly scalable, reliable, and efficient cloud infrastructure.
- Serve as the customer's technical advocate with internal engineering and product teams.
- Conduct post-incident reviews to identify root causes and implement preventative measures.
- Ensure security best practices are integrated into customer deployments.
- Train customer SRE, Operations, and Platform Engineering teams on platform usage and best practices.
- Lead enterprise migrations from legacy alerting, AIOps, and incident management platforms, including correlation rule migration, phased cutovers, and production go-live execution.
- Design, build, and optimize alert normalization and correlation policies using conditions, regular expressions, field extraction, and customized workflows.
- Integrate the platform with customer operational systems, including ITSM, collaboration, observability, source control, and documentation platforms.
- Validate and continuously improve AI investigation quality by tuning enrichment, root cause analysis accuracy, and investigation workflows.
- Build proactive monitoring for customer deployments to identify issues before they impact customers.
- Own customer-facing project communications, including executive status updates, SLA documentation, escalation management, and implementation tracking.
- Develop long-term technical relationships with senior engineering leadership.
- Own customer implementations from technical discovery through solution design, implementation, user acceptance testing, production go-live, stabilization, and ongoing optimization.
- Translate ambiguous customer requirements into clear technical designs, milestones, acceptance criteria, and execution plans.
- Design and implement AI-powered investigation and automation workflows with appropriate guardrails, governance, deterministic fallbacks, and human oversight.
- Develop reusable deployment modules, reference architectures, implementation guides, and operational runbooks to accelerate future deployments.
- Define customer success metrics, establish baselines, measure operational improvements, and demonstrate business value through KPIs such as MTTR reduction and operational efficiency.
- Capture customer feedback and recurring implementation learnings to influence future product development.
- Foster a culture of continuous improvement and technical excellence.
Qualifications
- Customer-focused with deep empathy for SRE, DevOps, Platform Engineering, and IT Operations teams.
- Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
- 6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure-focused roles, including technical leadership or end-to-end customer delivery.
- Experience in Forward Deployed Engineering, Solutions Engineering, Technical Customer Success, or Professional Services is highly preferred.
- Strong programming experience in at least one language such as Python, Go, or Java.
- Hands-on experience with public cloud platforms (AWS, Azure, or Google Cloud Platform).
- Strong knowledge of Kubernetes, Infrastructure as Code (Terraform, CloudFormation, Ansible), and CI/CD pipelines.
- Practical experience using Generative AI and machine learning technologies to improve engineering productivity.
- Experience with observability platforms, ITSM systems, and incident management tools, including systems integration and data mapping.
- Strong troubleshooting, analytical, and debugging skills, including alert correlation, normalization, and regular expression development.
- Excellent written and verbal communication skills.
- Demonstrated ownership of enterprise software implementations from discovery through production deployment.
- Strong integration experience with APIs, webhooks, event-driven architectures, authentication (SSO/SAML), data transformations, synchronization, and enterprise application integrations.
- Experience designing and deploying production-grade AI or automation workflows with governance and evaluation frameworks.
- Understanding of enterprise security concepts including RBAC, encryption, identity management, auditing, and secure networking.
- Ability to operate effectively in ambiguous, fast-paced customer environments while balancing architecture with execution.
- Self-motivated, adaptable, and capable of managing shifting priorities while driving successful customer outcomes.
Preferred Qualifications
- Experience supporting customers operating AI infrastructure or AI-enabled platforms.
- Experience migrating customers from legacy alerting, AIOps, or incident management platforms.
- Experience building internal automation and tooling using Python, Node.js, Bash, or similar scripting languages.
A bit about us
Backed by over $21M in capital from leading investors, they are building a next-generation AI product designed to transform how reliability engineering is done.
The founding team includes senior leaders and technical pioneers from industry giants like AWS, Cisco, VMware, and Gigamon - holding dozens of patents and having built critical systems at some of the most respected tech companies in the industry. This is a rare opportunity to join an early-stage team that’s solving tough technical problems in distributed systems, observability, and automation — all while shaping a product from the ground up.
Why join us?
Benefits
- Comprehensive medical, vision, and dental benefits.
- 401 (k) plans and commuter benefits.
- Free lunches, snacks, and top-of-the-line espressos!
- Equity that could change your life.
- High-impact role with plenty of mentorship opportunities from founders and other coworkers
- Collaborative coworkers with high IQ and high EQ. No politics. No bureaucracy. Open door policy.