AI Site Reliability Engineer

Seekr

Austin (TX)

Hybrid

USD 140,000 - 200,000

Full time

28 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity Ownership – RSUs
Unlimited PTO
14 paid company holidays
Hybrid work environment
401(k) with Company Match
Comprehensive Health & Wellness
Parental Leave

Job summary

Seekr is seeking an AI Site Reliability Engineer to ensure reliability, scalability, and safe operation of AI-powered services and infrastructure. You will blend software engineering with AI/ML operational expertise to deploy, monitor, and operate models, APIs, data pipelines, and platform services in production.

You will partner with AI/ML, Platform, Security, and Product teams to build dependable, performant, resilient systems ready to scale, with a focus on SLIs/SLOs, automation, and

Qualifications

  • Bachelor's degree or equivalent practical experience in CS/engineering.
  • 4+ years in Site Reliability Engineering, DevOps, or similar roles.
  • Strong knowledge of distributed systems, cloud infra, networks, containers, and Kubernetes.
  • Experience building and operating CI/CD pipelines and progressive deployment strategies.
  • Proficiency with infrastructure-as-code and tools like Terraform.
  • Experience with monitoring, logging, tracing, alerting, and incident management (Prometheus, Grafana, OpenTelemetry).
  • Strong scripting and software development in Python or Go.
  • Knowledge of SLIs/SLOs, capacity planning, and production-readiness reviews.

Responsibilities

  • Build safe release processes with CI/CD automation, progressive deployments, validation, rollback, and health metrics.
  • Define and operate SLIs/SLOs for AI APIs and critical services.
  • Develop observability through metrics, logs, traces, dashboards, and SLO alerts.
  • Lead on-call rotation, incident response, and blameless postmortems.
  • Reduce toil with infrastructure as code, automation, and DR planning.
  • Design automated load, stress, and scalability tests for AI production traffic.
  • Set performance baselines, release thresholds, and capacity forecasts.
  • Validate autoscaling, rate limiting, backpressure, graceful degradation, and recovery.
  • Scale GPU-backed model inference on Kubernetes with batching and load balancing.
  • Optimize inference engines and infra across GPUs, compute, networking, storage, and cost.
  • Support reliable training workloads including distributed training and checkpointing.
  • Collaborate with AI/ML, Platform, Security, and Product teams for production readiness.

Skills

Distributed systems
Cloud infrastructure
Kubernetes
CI/CD pipelines
Terraform
Monitoring & tracing
Prometheus
Grafana
OpenTelemetry
Python
Go
SRE / DevOps
Blameless postmortems

Education

Bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience

Tools

Terraform
Prometheus
Grafana
OpenTelemetry
Kubernetes
NVIDIA Triton
PyTorch
DeepSpeed
FSDP
vLLM

Job description

We're looking for an AI Site Reliability Engineer to ensure the reliability, scalability, and safe operation of Seekr's AI-powered services and supporting infrastructure. You will combine software engineering and site reliability practices with AI/ML operational expertise to improve how models, APIs, data pipelines, and platform services are deployed, monitored, and operated in production. You will partner closely with AI/ML, Platform, Security, and Product teams to build dependable systems that are performant, resilient, and ready to scale.

The Impact
Duties and Responsibilities
  • Build safe release processes using CI/CD automation, progressive deployments, automated validation, rollback mechanisms, and deployment health metrics.
  • Define and operate SLIs/SLOs for AI APIs and critical services, covering availability, latency, errors, throughput, model quality, output safety, user-facing correctness, and cost.
  • Develop actionable observability through metrics, logs, traces, dashboards, and SLO-based alerts.
  • Participate in a sustainable on-call rotation; lead incident response, improve runbooks, and facilitate blameless postmortems.
  • Reduce operational toil and improve resilience through infrastructure as code, automation, and disaster-recovery planning.
  • Design automated load, stress, spike, soak, and scalability tests that model realistic AI production traffic.
  • Establish performance baselines, release thresholds, and capacity forecasts for latency, throughput, concurrency, resource utilization, and cost per inference.
  • Validate autoscaling, rate limiting, backpressure, graceful degradation, and recovery from infrastructure and dependency failures.
  • Design and scale GPU-backed model inference on Kubernetes using replicas, autoscaling, batching, caching, and load balancing to meet SLAs/SLOs.
  • Optimize and troubleshoot inference engines and infrastructure across models, Kubernetes, compute, networking, storage, GPU memory, and cost.
  • Support reliable training workloads, including distributed training, scheduling, checkpointing, observability, and failure recovery.
  • Partner with AI/ML, Platform, Security, and Product teams to establish production-readiness standards.
Qualifications and Skills Required
  • Bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, production infrastructure engineering, or a similar role.
  • Strong knowledge of distributed systems, cloud infrastructure, networking, containers, and Kubernetes.
  • Experience building and operating CI/CD pipelines, automated validation, and progressive deployment strategies.
  • Proficiency with infrastructure-as-code and configuration-management tools, such as Terraform.
  • Experience with monitoring, logging, distributed tracing, alerting, and incident-management practices, including tools such as Prometheus, Grafana, and OpenTelemetry.
  • Strong scripting and software-development skills in Python, Go, or a similar language.
  • Practical knowledge of SLIs, SLOs, error budgets, capacity planning, production-readiness reviews, and blameless postmortems.
  • Ability to troubleshoot complex systems across application, model-serving, and infrastructure layers.
  • Strong communication and collaboration skills, with a commitment to sustainable and blameless operations.
  • Experience operating machine-learning or generative-AI systems in production is preferred.
  • Familiarity with model serving, inference optimization, LLM gateways, vector databases, GPU infrastructure, ML observability, and model evaluation is preferred.
  • Experience effectively utilizing AI technologies and tools, including large language models, agents, or AI coding assistants, to enhance workflows and operational effectiveness.
  • Hands-on experience with performance-testing tools such as k6, Locust, JMeter, or Gatling.
  • Experience testing distributed systems and interpreting latency percentiles, saturation, throughput, concurrency, and resource-consumption metrics.
  • Familiarity with AI inference benchmarking, GPU profiling, autoscaling, and performance-versus-cost optimization.
  • Experience with inference engines such as vLLM, NVIDIA Triton, TensorRT-LLM, or similar platforms, and scaling GPU workloads on Kubernetes.
  • Working knowledge of GPU infrastructure, memory management, quantization, distributed training, and frameworks such as PyTorch, DeepSpeed, or FSDP.
About the Company:

Seekr is a leader in explainable and trustworthy artificial intelligence designed to power mission-critical decisions in enterprises, government, and regulated industries. SeekrFlow™, our end-to-end AI platform, provides secure, auditable AI solutions tailored to sectors where transparency, accuracy, and compliance are paramount. Available across cloud, on-premises, and edge environments, SeekrFlow reduces bias, strengthens data integrity, and simplifies model oversight so organizations can rely on trusted AI decisions in high-stakes settings that impact society’s most sensitive and vital systems. Trusted by leading enterprises and government agencies, we partner with defense, finance, telecom, and critical infrastructure leaders to enable AI solutions that drive real-world results with unmatched transparency and control. We are a team of strategic thinkers and problem-solvers tackling the toughest challenges facing critical infrastructure and global enterprises through best-in-class AI models and customer deployment. Our team operates with unwavering commitment to our core values and mission:

  • We are driven by outcomes—our customers' success is what we strive for every day.
  • We believe trust is earned, which is why we build explainability and transparency into the entire AI lifecycle.
  • We take our responsibility to deliver secure AI seriously.
  • We believe innovation drives progress—we are building the technologies that power the systems our society depends on.
Company Benefits:
  • Meaningful Mission & Impact - Work with a deeply talented, collaborative team solving some of the toughest AI challenges that matter.
  • Equity Ownership – RSUs that let you share directly in Seekr’s long‑term success and growth.
  • Time Off That Respects Real Life – Unlimited PTO plus 14 paid company holidays to truly recharge.
  • Work Your Way – A flexible hybrid work environment with offices in Reston, VA and Austin, TX.
  • Competitive Total Rewards – A role‑appropriate compensation structure that supports long‑term growth, including base salary, bonuses, or commission plans depending on role.
  • 401(k) with Company Match – Build your future with a retirement plan that includes employer matching.
  • Comprehensive Health & Wellness – Medical, dental, vision, and life insurance coverage starting day one—for you and your family.
  • Parental Leave – Paid parental leave to support employees as they welcome a new child through birth, adoption, or foster placement.
Important Notice:

To conform to U.S. Government international trade regulations, applicant must be a U.S. Citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State or U.S. Department of Commerce.
Seekr is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other characteristic protected by applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Engineer
Principal AI Engineer

Seekr • Austin (TX)

Hybrid
USD 180,000 - 270,000
Meaningful mission
Equity RSUs
Unlimited PTO
+4
Principal AI Engineer
Principal AI Engineer

Seekr • Menlo Park (CA)

Hybrid
USD 190,000 - 320,000
Meaningful Mission & Impact
Equity Ownership - RSUs
Unlimited PTO + 14 holidays
+4
Principal AI Engineer
Principal AI Engineer

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Meaningful Mission & Impact
Equity Ownership
Unlimited PTO
+4
AI Engineer/Scientist - Staff
AI Engineer/Scientist - Staff

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Meaningful Mission & Impact
Equity Ownership RSUs
Time Off That Respects Real Life
+5
AI Engineer/Scientist - Staff
AI Engineer/Scientist - Staff

Seekr • Austin (TX)

Hybrid
USD 180,000 - 280,000
Equity ownership (RSUs)
Unlimited PTO
Flexible hybrid work environment
+2
Senior DevOps Engineer
Senior DevOps Engineer

Seekr • Reston (VA)

Hybrid
USD 150,000 - 210,000
Equity RSUs
Unlimited PTO
14 paid company holidays
+1
AI Site Reliability Engineer
AI Site Reliability Engineer

Seekr • Reston (VA)

Hybrid
USD 140,000 - 190,000
Equity Ownership
Unlimited PTO
14 paid company holidays
+3
Senior Software Engineer
Senior Software Engineer

Seekr • Reston (VA)

Hybrid
USD 150,000 - 210,000
Meaningful Mission & Impact
Equity Ownership – RSUs
Unlimited PTO + 14 holidays
+4
Senior Software Engineer
Senior Software Engineer

Seekr • Austin (TX)

Hybrid
USD 140,000 - 190,000
Equity Ownership - RSUs
Unlimited PTO + 14 holidays
Flexible hybrid work: Reston & Austin
+4
Senior DevOps Engineer
Senior DevOps Engineer

Seekr • Austin (TX)

Hybrid
USD 140,000 - 190,000
Equity Ownership
Unlimited PTO
Flexible hybrid (Reston/Austin)
+3