Senior Software Engineer – Application Reliability, Kubernetes, GCP, SQL

Jobtailor

California (MO)

On-site

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior reliability engineering leader to own AI-powered application reliability across user-facing features. You will define SLI/SLOs, build observability with Looker on BigQuery/BigTable, and drive blameless postmortems while working with cross-functional teams.

The role emphasizes hands-on GenAI workflows, agent engineering, and incident response, requiring deep GCP experience, complex SQL skills, and strong collaboration across engineering leadership.

Qualifications

  • 10+ years of software engineering with reliability, observability, or production operations focus.
  • Strong Python development skills for tooling and agent-based systems.
  • Production GCP experience deploying on GKE (Kubernetes).
  • Deep SQL expertise with BigQuery, complex queries and design.
  • Experience designing/operating SLI/SLO frameworks and error budgets.
  • Hands-on Looker dashboards and LookML for observability.

Responsibilities

  • Own the reliability of AI-powered applications and user-facing features.
  • Define, implement, and enforce SLI, SLO, and error budgets for APIs and systems.
  • Build and maintain Looker dashboards on BigQuery and BigTable.
  • Lead incident response, root cause analysis, and blameless postmortems.
  • Develop agent evaluation harnesses for benchmarking and testing.
  • Write complex BigQuery SQL for usage trends and anomaly detection.

Skills

Python Development
BigQuery SQL Expertise
GCP Experience
Application-Level SLI/SLO Frameworks
Agent Evaluation Harnesses

Education

Bachelor's Degree in Computer Science
Master's Degree in Engineering

Tools

Looker
BigQuery
BigTable
GKE
Cloud Logging
Cloud Trace
Cloud Monitoring
LangGraph

Job description

Responsibilities
  • Own the reliability of AI-powered applications and features from the user’s perspective
  • Define, implement, and enforce feature-level SLIs, SLOs, and error budgets for APIs, RAG systems, AI agents, and user-facing applications
  • Build and maintain application observability systems using Looker dashboards on BigQuery and BigTable
  • Provide visibility into feature health, error patterns, and usage trends for developers, product managers, and leadership
  • Design and build LangGraph-based agents for automated issue identification and remediation
  • Implement anomaly detection, root cause diagnosis, auto-rollback, feature-flag kill switches, and self-healing workflows
  • Develop agent evaluation harnesses for benchmarking, multi-step workflow testing, non-deterministic outputs, and regression testing
  • Write complex BigQuery SQL for usage trend analysis, anomaly detection, and operational analytics
  • Design BigQuery table schemas optimized for observability and debugging
  • Analyze application usage trends and adoption metrics to identify reliability risks, capacity needs, and degraded user experiences
  • Partner with application development teams to embed deployment safety, structured logging, and distributed tracing practices
  • Lead application-level incident response, root cause analysis, and blameless postmortems
  • Build Python tooling and automation to reduce mean time to detect and resolve application-layer issues
  • Apply emerging AI techniques to improve platform reliability and developer productivity
  • Collaborate with application developers, data engineers, infrastructure SREs, security, compliance, product teams, and leadership
Requirements
  • 10+ years of experience in software engineering with significant focus on reliability, observability, or production operations
  • Bachelor's or Master's Degree in Computer Science, Engineering, or a related technical discipline
  • Strong Python development skills, including production tooling, automation, and agent-based systems
  • Production GCP experience deploying and managing applications on GKE (Kubernetes)
  • Deep SQL expertise with BigQuery, including complex queries, window functions, schema design, and cost optimization
  • Hands-on experience with BigTable or an equivalent high-throughput operational data system
  • Experience designing and operating application-level SLI/SLO frameworks, burn-rate alerting, and error budget policies
  • Strong application-layer debugging skills, including distributed tracing, profiling, structured log analysis, and dependency mapping
  • Experience building agent evaluation harnesses
  • Familiarity with A2A protocols, streaming architectures, and event-driven systems
  • Experience with feature flags, canary deployments, progressive rollouts, and automated rollback
  • Experience with GCP observability services such as Cloud Logging, Cloud Trace, and Cloud Monitoring
  • Exposure to AIOps concepts including ML-driven anomaly detection, automated root cause analysis, and intelligent alerting
  • Experience driving reliability culture across engineering teams
  • Active engagement with the evolving AI ecosystem
  • Hands‑on GenAI application development experience with LangGraph, agent engineering, prompt design, and agentic workflows
  • Experience building Looker dashboards and LookML models for operational observability
Core Competencies

Demonstrates expertise in application reliability and observability, leveraging strong Python development skills and deep SQL knowledge with BigQuery. Capable of implementing SLI/SLO frameworks and driving a reliability culture across engineering teams.

Highest-signal resume keywords
  • Python Development
  • BigQuery SQL Expertise
  • GCP Experience
  • Application-Level SLI/SLO Frameworks
  • Agent Evaluation Harnesses
Hard Skills
  • Application Reliability
  • Observability Systems
  • Anomaly Detection
  • Root Cause Analysis
  • Feature Flags
  • Automated Rollback
  • Complex SQL Queries
  • Schema Design
  • Distributed Tracing
  • Debugging Skills
Soft Skills
  • Collaboration
  • Leadership
  • Communication
Certifications & Qualifications
  • Bachelor's Degree in Computer Science
  • Master's Degree in Engineering
Industry Keywords
  • AI-Powered Applications
  • Production Operations
  • Reliability Engineering
  • AIOps
  • Event-Driven Systems
Tools & Technologies
  • Looker
  • BigQuery
  • BigTable
  • GKE
  • Cloud Logging
  • Cloud Trace
  • Cloud Monitoring
  • LangGraph
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Software Engineer – AI/ML Engineer
Distinguished Software Engineer – AI/ML Engineer

Jobtailor • Sunnyvale (CA)

On-site
USD 230,000 - 350,000
Senior Data Engineer
Senior Data Engineer

Jobtailor • Menomonee Falls (WI)

On-site
USD 120,000 - 180,000
Principal Engineer – Public Cloud Data
Principal Engineer – Public Cloud Data

Jobtailor • Arizona

On-site
USD 150,000 - 185,000
Senior Software Engineer – Technical Lead, Python, .NET/C#, Angular, AWS
Senior Software Engineer – Technical Lead, Python, .NET/C#, Angular, AWS

Jobtailor • Connecticut

On-site
USD 130,000 - 200,000
Principal Software Engineer – Backend
Principal Software Engineer – Backend

Jobtailor • Kirkland (WA)

On-site
USD 180,000 - 260,000
Senior AI Engineer
Senior AI Engineer

Jobtailor • Houston (TX)

On-site
USD 120,000 - 180,000
Senior AI Reliability Engineer – Kubernetes & GCP
Senior AI Reliability Engineer – Kubernetes & GCP

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Lead Software Engineer
Lead Software Engineer

Jobtailor • Colorado

On-site
USD 150,000 - 190,000
Senior AI Delivery & Operations Engineer
Senior AI Delivery & Operations Engineer

Jobtailor • Illinois

On-site
USD 140,000 - 190,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Jobtailor • North Carolina

On-site
USD 150,000 - 190,000