Lead SRE

BMC Software, Inc.

United States

Remote

USD 153,000 - 255,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

BMC Software, Inc. is seeking a Senior AI Platform Engineer to build, operate, and improve AI/ML and Generative AI platforms in production, covering deployment, monitoring, incident management, and continuous improvement.

You will work with AI engineers, data scientists, DevOps, and security to ensure reliability, scalability, and observability across LLM, GenAI, and related workloads, with on-call responsibilities as needed.

Qualifications

  • SRE fundamentals; SLOs/SLIs, monitoring, incident management.
  • Linux, networking, distributed systems, troubleshooting, performance analysis.
  • CI/CD and DevSecOps practices; containers/orchestration (Docker, Kubernetes, OpenShift); cloud platforms.
  • Generative AI/LLM literacy; RAG/agents/tooling concepts as level requires.
  • MLOps/LLMOps tooling and AI observability as level requires.
  • Scripting/automation (Python, Bash); security-first, calm incident handling.
  • Linux/scripting fundamentals; exposure to monitoring or cloud (project/internship OK).
  • Curiosity about how LLM/agent systems fail in production.
  • Calm, careful work under guidance; readiness for eventual on-call with support.
  • Past experience: Contributed to monitoring, runbooks, or operational fixes within an existing workflow.

Responsibilities

  • Build and operate production-grade infrastructure for LLM, GenAI, RAG, and Agentic AI applications.
  • Define and defend SLOs/SLIs for agent and model behavior—latency, availability, quality, and cost.
  • Monitor production AI systems for performance, drift, hallucination rates, and quality regressions; act before customers are affected.
  • Lead incident response for AI systems: detection, triage, mitigation, and blameless post-mortems.
  • Build runbooks, automation, and self-healing for common AI operational issues; participate in on-call as required.
  • Operate model/agent lifecycle in production: versioning, rollout, rollback, and safe promotion with evaluation gates.
  • Manage LLM cost, capacity, and scaling; manage model serving/inference infrastructure.
  • Maintain observability/tracing for agents, tools, and prompts (Langfuse, OpenTelemetry, OpenSearch).
  • Operate AI/ML workloads on IBM Z where required; harden the AMI Platform for reliability, scalability, and security.
  • Partner with AI Engineers, Data Scientists, DevOps, and Security to embed reliability from design onward.

Skills

SRE fundamentals
Linux fundamentals
Networking
Distributed systems
Troubleshooting
CI/CD
DevSecOps
Containers & Orchestration
Docker
Kubernetes
OpenShift
Cloud platforms
Generative AI / LLM literacy
MLOps / LLMOps
Python / Bash scripting
Security-first incident handling
On-call readiness
Monitoring tools exposure

Tools

Langfuse
OpenTelemetry
OpenSearch
CI/CD tooling
Incident management tools

Job description

You may occasionally be required to travel for business

Secondary locations:

USA California - OFC at Home, USA Michigan - OFC at Home

Additional Locations:

This role can be based remotely in United States

Looking for more details about our benefits?

Description and Requirements

BMC empowers nearly 80% of the Forbes Global 100 to accelerate business value, faster than humanly possible. Our industry-leading portfolio unlocks human and machine potential to drive business growth, innovation, and sustainable success. BMC does this in a simple and optimized way by connecting people, systems, and data that power the world’s largest organizations so they can seize a competitive advantage.

About the Role

You build, operate, and continuously improve reliable, scalable, secure, and observable AI/ML and Generative AI platforms in production — supporting the full lifecycle from deployment and monitoring through incident management, optimization, and continuous improvement.

At this level you execute defined work with guidance and grow through review and mentorship.

Scope at this level: Completes defined operations/observability tasks with guidance; learns AI failure modes and on-call basics.

Organizational impact expected: Contributes reliable work to a project.

Key Responsibilities

Core responsibilities for this role:

  • Build and operate production-grade infrastructure and operational frameworks for LLM, GenAI, RAG, and Agentic AI applications.
  • Define and defend SLOs/SLIs for agent and model behaviour — latency, availability, quality, and cost.
  • Monitor production AI systems for performance, drift, hallucination rates, and quality regressions; act before customers are affected.
  • Lead incident response for AI systems: detection, triage, mitigation, and blameless post-mortems.
  • Build runbooks, automation, and self-healing for common AI operational issues; participate in on-call as required.
  • Operate model/agent lifecycle in production: versioning, rollout, rollback, and safe promotion with evaluation gates.
  • Manage LLM cost, capacity, and scaling; manage model serving/inference infrastructure.
  • Maintain observability/tracing for agents, tools, and prompts (Langfuse, OpenTelemetry, OpenSearch).
  • Operate AI/ML workloads on IBM Z where required; harden the AMI Platform for reliability, scalability, and security.
  • Partner with AI Engineers, Data Scientists, DevOps, and Security to embed reliability from design onward.

Must-Have Skills & Experience

  • SRE fundamentals — SLOs/SLIs, monitoring, alerting, incident management (depth scales with level).
  • Linux, networking, distributed systems, troubleshooting, performance analysis.
  • CI/CD and DevSecOps practices; containers/orchestration (Docker, Kubernetes, OpenShift); cloud platforms.
  • Generative AI/LLM literacy; RAG/agents/tooling concepts as level requires.
  • MLOps/LLMOps tooling and AI observability as level requires.
  • Scripting/automation (Python, Bash); security-first, calm incident handling.
  • Linux/scripting fundamentals; exposure to monitoring or cloud (project/internship OK).
  • Curiosity about how LLM/agent systems fail in production.
  • Calm, careful work under guidance; readiness for eventual on-call with support.
  • Past experience: Contributed to monitoring, runbooks, or operational fixes within an existing workflow; basic exposure to cloud/containers.
  • Delivery evidence: Reliable execution of defined ops tasks; clear notes; careful changes under guidance.
  • Shared expectation: Completes defined tasks with guidance and checks work carefully.

Nice-to-Have Skills

  • Agent frameworks and AgentOps for multi-agent systems.
  • LLM cost-optimization and inference-serving stacks (vLLM, Ollama).
  • Enterprise on-prem or hybrid deployments; Agile/Atlassian.

Our commitment to you!

BMC’s culture is built around its people. We have 6000+ brilliant minds working together across the globe. You won’t be known just by your employee number, but for your true authentic self. BMC lets you be YOU!

BMC is committed to equal opportunity employment regardless of race, age, sex, creed, color, religion, citizenship status, sexual orientation, gender, gender expression, gender identity, national origin, disability, marital status, pregnancy, disabled veteran or status as a protected veteran. If you need a reasonable accommodation for any part of the application and hiring process, visit the accommodation request page.

BMC Software maintains a strict policy of not requesting any form of payment in exchange for employment opportunities, upholding a fair and ethical hiring process.

The annual base salary range represents the low and high end of the BMC salary range for this position. Actual salaries depend on a wide range of factors that are considered in making compensation decisions, including but not limited to skill sets; experience and training, licensure, and certifications; and other business and organizational needs.

The range listed is just one component of BMC's employee compensation package. Other rewards may include a variable plan and country specific benefits.

At BMC, it is not typical for an individual to be hired at /near the top of the range. A reasonable estimate of the current range is $152,925 - $254,875

We use AI technology to support parts of our recruitment process, but people—not algorithms—make all final hiring decisions. AI may assist with tasks like scheduling, screening for role alignment, or helping us manage large volumes of applications more efficiently. However, candidates are reviewed by a member of our recruitment team, and interviews and hiring decisions are always made by people. We’re committed to ensuring that technology enhances fairness, efficiency, and the candidate experience—never replace genuine human judgment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Engineer
Principal AI Engineer

BMC Software, Inc. • California (MO)

On-site
USD 198,000 - 330,000
Principal QA Engineer
Principal QA Engineer

BMC Software, Inc. • United States

Remote
USD 153,000 - 255,000
Principal AI Engineer
Principal AI Engineer

BMC Software, Inc. • Santa Clara (CA)

On-site
USD 198,000 - 330,000
Lead AI Engineer
Lead AI Engineer

BMC Software • Santa Clara (CA)

On-site
USD 176,000 - 293,000
Lead Data Scientist
Lead Data Scientist

BMC Software, Inc. • United States

Remote
USD 176,000 - 293,000
Senior Data Scientist
Senior Data Scientist

BMC Software, Inc. • United States

Remote
USD 153,000 - 255,000
Principal QA Engineer
Principal QA Engineer

BMC Software, Inc • New York (NY)

On-site
USD 153,000 - 255,000
Principal AI Engineer - Office of the CTO
Principal AI Engineer - Office of the CTO

BMC Software, Inc • Santa Clara (CA)

On-site
USD 176,000 - 293,000
Lead Forward Deployed Engineer
Lead Forward Deployed Engineer

BMC Software, Inc. • United States

Remote
USD 153,000 - 255,000
Lead Software Engineer - AgenticAI
Lead Software Engineer - AgenticAI

BMC Software, Inc. • United States

On-site
USD 153,000 - 255,000