Senior AI-Powered SRE Forward Deployed Engineer

Jobot

Pleasanton (CA)

On-site

USD 300,000 - 350,000

Full time

16 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical benefits
401(k) plans
Free lunches and snacks
Equity opportunity
Mentorship from founders
Collaborative culture

Job summary

Jobot is seeking an experienced Site Reliability Engineer to lead the design, deployment, and optimization of an AI-powered SRE platform across production and pre-production environments. You will proactively monitor deployments, identify reliability gaps, and drive scalable cloud infrastructure while partnering with engineering and product teams.

You will guide migrations from legacy systems, implement robust alerting and automation, and mentor customer teams on best practices.

Qualifications

  • 6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure-focused roles, including technical leadership or end-to-end customer delivery.
  • Strong programming experience in at least one language such as Python, Go, or Java.
  • Hands-on experience with public cloud platforms (AWS, Azure, or Google Cloud Platform).
  • Strong knowledge of Kubernetes, Infrastructure as Code (Terraform, CloudFormation, Ansible), and CI/CD pipelines.
  • Practical experience using Generative AI and machine learning technologies to improve engineering productivity.
  • Experience with observability platforms, ITSM systems, and incident management tools, including systems integration and data mapping.
  • Strong troubleshooting, analytical, and debugging skills, including alert correlation, normalization, and regular expression development.
  • Excellent written and verbal communication skills.
  • Demonstrated ownership of enterprise software implementations from discovery through production deployment.

Responsibilities

  • Implement and optimize an AI-powered Site Reliability Engineering (SRE) platform to meet customer needs across production and pre-production environments.
  • Proactively monitor customer deployments to ensure customers maximize value from the platform.
  • Identify latent reliability issues such as misconfigurations, deployment regressions, and scaling challenges within customer environments.
  • Plan, design, build, and maintain highly scalable, reliable, and efficient cloud infrastructure.
  • Serve as the customer's technical advocate with internal engineering and product teams.
  • Conduct post-incident reviews to identify root causes and implement preventative measures.
  • Ensure security best practices are integrated into customer deployments.
  • Train customer SRE, Operations, and Platform Engineering teams on platform usage and best practices.
  • Lead enterprise migrations from legacy alerting, AIOps, and incident management platforms, including correlation rule migration, phased cutovers, and production go-live execution.
  • Design, build, and optimize alert normalization and correlation policies using conditions, regular expressions, field extraction, and customized workflows.
  • Integrate the platform with customer operational systems, including ITSM, collaboration, observability, source control, and documentation platforms.
  • Validate and continuously improve AI investigation quality by tuning enrichment, root cause analysis accuracy, and investigation workflows.
  • Build proactive monitoring for customer deployments to identify issues before they impact customers.
  • Own customer-facing project communications, including executive status updates, SLA documentation, escalation management, and implementation tracking.
  • Develop long-term technical relationships with senior engineering leadership.
  • Own customer implementations from technical discovery through solution design, implementation, user acceptance testing, production go-live, stabilization, and ongoing optimization.
  • Translate ambiguous customer requirements into clear technical designs, milestones, acceptance criteria, and execution plans.
  • Design and implement AI-powered investigation and automation workflows with appropriate guardrails, governance, deterministic fallbacks, and human oversight.
  • Develop reusable deployment modules, reference architectures, implementation guides, and operational runbooks to accelerate future deployments.
  • Define customer success metrics, establish baselines, measure operational improvements, and demonstrate business value through KPIs such as MTTR reduction and operational efficiency.
  • Capture customer feedback and recurring implementation learnings to influence future product development.
  • Foster a culture of continuous improvement and technical excellence.

Skills

SRE experience
DevOps
Platform Engineering
Customer Delivery
Python
Go
Java
Cloud platforms
Kubernetes
CI/CD
Observability
Generative AI

Education

Bachelor's degree in CS/Engineering

Tools

Terraform
CloudFormation
Ansible
AWS
Azure
GCP
Kubernetes

Job description

Jobot is seeking an experienced Site Reliability Engineer to lead the design, deployment, and optimization of an AI-powered SRE platform across production and pre-production environments. You will proactively monitor deployments, identify reliability gaps, and drive scalable cloud infrastructure while partnering with engineering and product teams.

You will guide migrations from legacy systems, implement robust alerting and automation, and mentor customer teams on best practices.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of SRE Engineering for AI Security & Cloud
Director of SRE Engineering for AI Security & Cloud

Jobot • Chicago (IL)

Remote
USD 220,000 - 260,000
AI Security SRE Director: Platform Reliability
AI Security SRE Director: Platform Reliability

Jobot • Philadelphia

On-site
USD 220,000 - 260,000
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

Jobot • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plans
Free lunches and snacks
+3
AI-Driven SRE Engineer for Cloud & Automation
AI-Driven SRE Engineer for Cloud & Automation

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Site Reliability Engineer
Site Reliability Engineer

IntraEdge • Austin (TX)

On-site
USD 120,000 - 180,000
Senior SRE: AI-Driven Reliability & Automation (Hybrid)
Senior SRE: AI-Driven Reliability & Automation (Hybrid)

Splash • United States

Hybrid
USD 120,000 - 150,000
Senior Hybrid Cloud SRE — Reliability & Automation
Senior Hybrid Cloud SRE — Reliability & Automation

PathAI • Boston (MA)

On-site
USD 165,750 - 224,450
Director of SRE & Platform Engineering — Remote
Director of SRE & Platform Engineering — Remote

Jobot • New York (NY)

Remote
USD 220,000 - 260,000
Stock options
10% annual bonus
Senior AI SRE Enterprise Deployment Lead
Senior AI SRE Enterprise Deployment Lead

Ciroos • United States

Remote
USD 180,000 - 240,000
Medical benefits
401k
Free meals
+1