Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce

Pleasanton (CA)

On-site

USD 300,000 - 350,000

Full time

7 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical benefits
401(k) plan
Free meals and snacks
Equity opportunities
Mentorship from founders

Job summary

LeoForce is seeking a senior-level SRE leader to implement and optimize an AI-powered platform across production and pre-production environments in Pleasanton, CA. You will own end-to-end customer deployments, monitor deployments for value, and guide architectural decisions with a focus on reliability and security.

You will design scalable cloud infrastructure, lead migrations from legacy systems, and collaborate with product and engineering teams to drive operational excellence and measurable

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
  • 6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar roles with technical leadership.
  • Strong programming experience in Python, Go, or Java.
  • Hands-on experience with public cloud platforms (AWS, Azure, or Google Cloud).
  • Strong knowledge of Kubernetes, IaC (Terraform, CloudFormation, Ansible), and CI/CD pipelines.
  • Experience with observability platforms, ITSM systems, and incident management tools, including systems integration.
  • Experience designing and deploying production-grade AI or automation workflows with governance and evaluation frameworks.
  • Understanding of enterprise security concepts including RBAC, encryption, identity management, auditing, and secure networking.

Responsibilities

  • Implement and optimize an AI-powered SRE platform to meet customer needs across production and pre-production environments.
  • Proactively monitor customer deployments to ensure customers maximize value from the platform.
  • Identify latent reliability issues such as misconfigurations, deployment regressions, and scaling challenges within customer environments.
  • Plan, design, build, and maintain highly scalable, reliable, and efficient cloud infrastructure.
  • Lead enterprise migrations from legacy alerting, AIOps, and incident management platforms, including correlation rule migration, phased cutovers, and production go-live execution.
  • Design, build, and optimize alert normalization and correlation policies using conditions, regular expressions, field extraction, and customized workflows.
  • Integrate the platform with customer operational systems, including ITSM, collaboration, observability, source control, and documentation platforms.
  • Validate and continuously improve AI investigation quality by tuning enrichment, root cause analysis accuracy, and investigation workflows.
  • Own customer implementations from technical discovery through solution design, implementation, user acceptance testing, production go-live, stabilization, and ongoing optimization.
  • Develop reusable deployment modules, reference architectures, implementation guides, and operational runbooks to accelerate future deployments.

Skills

Python
Go
Java
CI/CD pipelines
APIs integration
RBAC and security concepts

Education

Bachelor's degree in Computer Science or Engineering (or equivalent practical experience)

Tools

Kubernetes
Terraform
CloudFormation

Job description

Location: Pleasanton,CA, US

Experience: Senior Level

Salary: $300,000 - $350,000 per year

Responsibilities
  • Implement and optimize an AI-powered Site Reliability Engineering (SRE) platform to meet customer needs across production and pre-production environments.
  • Proactively monitor customer deployments to ensure customers maximize value from the platform.
  • Identify latent reliability issues such as misconfigurations, deployment regressions, and scaling challenges within customer environments.
  • Recommend best practices for implementing AI-powered SRE solutions.
  • Plan, design, build, and maintain highly scalable, reliable, and efficient cloud infrastructure.
  • Serve as the customer's technical advocate with internal engineering and product teams.
  • Conduct post-incident reviews to identify root causes and implement preventative measures.
  • Ensure security best practices are integrated into customer deployments.
  • Train customer SRE, Operations, and Platform Engineering teams on platform usage and best practices.
  • Lead enterprise migrations from legacy alerting, AIOps, and incident management platforms, including correlation rule migration, phased cutovers, and production go-live execution.
  • Design, build, and optimize alert normalization and correlation policies using conditions, regular expressions, field extraction, and customized workflows.
  • Integrate the platform with customer operational systems, including ITSM, collaboration, observability, source control, and documentation platforms.
  • Validate and continuously improve AI investigation quality by tuning enrichment, root cause analysis accuracy, and investigation workflows.
  • Build proactive monitoring for customer deployments to identify issues before they impact customers.
  • Own customer-facing project communications, including executive status updates, SLA documentation, escalation management, and implementation tracking.
  • Develop long-term technical relationships with senior engineering leadership.
  • Own customer implementations from technical discovery through solution design, implementation, user acceptance testing, production go-live, stabilization, and ongoing optimization.
  • Translate ambiguous customer requirements into clear technical designs, milestones, acceptance criteria, and execution plans.
  • Design and implement AI-powered investigation and automation workflows with appropriate guardrails, governance, deterministic fallbacks, and human oversight.
  • Develop reusable deployment modules, reference architectures, implementation guides, and operational runbooks to accelerate future deployments.
  • Define customer success metrics, establish baselines, measure operational improvements, and demonstrate business value through KPIs such as MTTR reduction and operational efficiency.
  • Capture customer feedback and recurring implementation learnings to influence future product development.
  • Foster a culture of continuous improvement and technical excellence.
Qualifications
  • Customer-focused with deep empathy for SRE, DevOps, Platform Engineering, and IT Operations teams.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
  • 6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure-focused roles, including technical leadership or end-to-end customer delivery.
  • Experience in Forward Deployed Engineering, Solutions Engineering, Technical Customer Success, or Professional Services is highly preferred.
  • Strong programming experience in at least one language such as Python, Go, or Java.
  • Hands-on experience with public cloud platforms (AWS, Azure, or Google Cloud Platform).
  • Strong knowledge of Kubernetes, Infrastructure as Code (Terraform, CloudFormation, Ansible), and CI/CD pipelines.
  • Practical experience using Generative AI and machine learning technologies to improve engineering productivity.
  • Experience with observability platforms, ITSM systems, and incident management tools, including systems integration and data mapping.
  • Strong troubleshooting, analytical, and debugging skills, including alert correlation, normalization, and regular expression development.
  • Excellent written and verbal communication skills.
  • Demonstrated ownership of enterprise software implementations from discovery through production deployment.
  • Strong integration experience with APIs, webhooks, event-driven architectures, authentication (SSO/SAML), data transformations, synchronization, and enterprise application integrations.
  • Experience designing and deploying production-grade AI or automation workflows with governance and evaluation frameworks.
  • Understanding of enterprise security concepts including RBAC, encryption, identity management, auditing, and secure networking.
  • Ability to operate effectively in ambiguous, fast-paced customer environments while balancing architecture with execution.
  • Self-motivated, adaptable, and capable of managing shifting priorities while driving successful customer outcomes.
Preferred Qualifications
  • Experience supporting customers operating AI infrastructure or AI-enabled platforms.
  • Experience migrating customers from legacy alerting, AIOps, or incident management platforms.
  • Experience building internal automation and tooling using Python, Node.js, Bash, or similar scripting languages.
A bit about us:

Backed by over $21M in capital from leading investors, they are building a next-generation AI product designed to transform how reliability engineering is done.

The founding team includes senior leaders and technical pioneers from industry giants like AWS, Cisco, VMware, and Gigamon - holding dozens of patents and having built critical systems at some of the most respected tech companies in the industry. This is a rare opportunity to join an early-stage team that’s solving tough technical problems in distributed systems, observability, and automation — all while shaping a product from the ground up.

Why join us?
Benefits
  • Comprehensive medical, vision, and dental benefits.
  • 401 (k) plans and commuter benefits.
  • Free lunches, snacks, and top-of-the-line espressos!
  • Equity that could change your life.
  • High-impact role with plenty of mentorship opportunities from founders and other coworkers
  • Collaborative coworkers with high IQ and high EQ. No politics. No bureaucracy. Open door policy.

#techservices #python #aws #java #azure #splunk #datadog #prometheus #devops #gcp #sre #servicenow #dynatrace #oberservability #claude-code #bigpanda #moogsoft #tier3

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

Jobot • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) and commuter benefits
Free lunches and snacks
+3
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
AI Platform / SRE Lead
AI Platform / SRE Lead

Luxoft • Northern (KY)

Hybrid
USD 160,000 - 230,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Staff+ Software Engineer - Backend
Staff+ Software Engineer - Backend

ResolveAI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Medical Insurance
Housing Stipend
Unlimited PTO
+5
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Staff+ Software Engineer - Backend
Staff+ Software Engineer - Backend

Resolve AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive Medical, Dental, and Vision Insurance
Monthly Housing Stipend
Flexible (Unlimited) Paid Time Off
+5