Site Reliability Engineer (SRE)

Nuvento Inc

Kansas City (MO)

On-site

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Imperial PFS is seeking an experienced Site Reliability Engineer (SRE) to ensure reliability, performance, and availability of production systems. You will pair SRE fundamentals with AI-powered automation using Azure SRE Agent to reduce toil and accelerate MTTR.

You will collaborate with DevOps, Platform Engineering, and development teams to instrument applications for AI-driven monitoring, define SLOs/SLIs, and implement governance for automated remediations and incident response.

Qualifications

  • Bachelor's degree in Computer Science, IT, or related field or equivalent experience.
  • 4+ years in Site Reliability Engineering, DevOps, or similar production ops roles.
  • Hands-on with Azure SRE Agent or similar AI-powered incident tooling preferred.
  • Strong knowledge of Azure services and observability platforms.

Responsibilities

  • Deploy and manage Azure SRE Agent across production workloads to automate investigations.
  • Monitor systems and respond to incidents with agent-guided diagnoses and approvals.
  • Build runbooks and hooks to automate common operational tasks and reduce toil.
  • Define and maintain SLOs/SLIs and error budgets for critical services.
  • Collaborate with DevOps and development teams to instrument apps for AI-driven monitoring.

Skills

Azure SRE Agent
Azure Cloud
SRE fundamentals
CI/CD tooling
Automation scripting

Education

Bachelor's degree in Computer Science or related field

Tools

Azure Monitor
Application Insights
Log Analytics
GitHub Actions
Terraform
Bicep

Job description

Job Title: Site Reliability Engineer (SRE)

Job Type: Contract / Contract-to-Hire

Position Description

As an SRE Engineer at Imperial PFS, you will help ensure the reliability, performance, and availability of production systems by pairing strong site reliability engineering fundamentals with AI-powered automation. You will work hands‑on with Azure SRE Agent as your primary tool for incident investigation, root cause analysis, and remediation, using it to reduce operational toil and accelerate mean time to resolution. You will partner closely with DevOps, Platform Engineering, and development teams to build a culture where AI agents and engineers work side by side to keep systems healthy and continuously improve reliability.

Essential Job Functions
  • Deploy, configure, and manage Azure SRE Agent across production workloads to automate incident investigation, root cause analysis, and remediation workflows
  • Monitor production systems and respond to incidents in partnership with Azure SRE Agent, reviewing agent‑proposed diagnoses and mitigations and approving actions per governance policy
  • Build and maintain runbooks, subagents, and agent hooks within Azure SRE Agent to automate common operational tasks and reduce manual toil
  • Connect Azure SRE Agent to observability, incident management, and source control tooling (Azure Monitor, Application Insights, Log Analytics, GitHub, PagerDuty, or similar) to enable end-to-end automated investigations
  • Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for critical services, and use agent-driven insights to track and improve them
  • Establish and enforce tool permissions, hooks, and governance controls for AI agent actions to ensure safe, auditable automation with appropriate human approval gates
  • Analyze incident trends and agent‑surfaced institutional knowledge to identify recurring issues and drive permanent fixes and reliability improvements
  • Participate in on-call rotation, leveraging Azure SRE Agent to accelerate triage, reduce mean time to detect/resolve (MTTD/MTTR), and minimize after-hours disruptions
  • Author and refine incident response plans, playbooks, and postmortem processes, incorporating agent‑generated documentation and learnings
  • Partner with DevOps, Platform Engineering, and development teams to instrument applications and infrastructure for effective AI‑driven monitoring and diagnostics
  • Continuously evaluate new Azure SRE Agent capabilities (connectors, private plugins, subagents) and pilot adoption to expand automated coverage
  • Contribute to Infrastructure as Code and automation scripts that support reliable, repeatable, and agent‑manageable environments
Experience in the following areas is required:
  • Bachelor's degree in Computer Science, Information Technology, or a related field required, or equivalent professional experience
  • 4+ years of experience in Site Reliability Engineering, DevOps, or a related production operations role
  • Hands‑on experience with Azure SRE Agent or comparable AI‑powered incident response/observability tooling strongly preferred
  • Strong working knowledge of Azure cloud services, including AKS, App Service, Azure Functions, and Azure networking
  • Experience with observability and monitoring platforms (Azure Monitor, Application Insights, Log Analytics, or similar)
  • Practical understanding of SRE fundamentals, including SLOs/SLIs, error budgets, incident management, and blameless postmortems
  • Familiarity with CI/CD tooling such as Azure DevOps Pipelines and GitHub Actions
  • Experience with Infrastructure as Code tools such as Bicep or Terraform
  • Comfort working with AI agents and automation frameworks, including reviewing and approving AI‑proposed remediations under governance controls
  • Strong scripting and automation skills (PowerShell, Bash, or Python)
  • Excellent troubleshooting skills and the ability to remain calm and methodical during production incidents
  • Strong communication skills, with the ability to document findings clearly for both technical and non‑technical audiences
  • Willingness to participate in an on‑call rotation.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI-Driven SRE on Azure — Reliability Engineer (Contract)
AI-Driven SRE on Azure — Reliability Engineer (Contract)

Nuvento Inc • Kansas City (MO)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Moultrie • Birmingham (AL)

On-site
USD 110,000 - 170,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Senior SRE - Azure
Senior SRE - Azure

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 100,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

IntraEdge • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000