Get more replies from employers
Send a job-specific resume in minutes.
FinThrive, Inc. is seeking a Site Reliability Engineer at an intermediate level to design, operate, and optimize cloud-native platforms with a strong Azure focus. You will drive automation-first initiatives, implement IaC, and contribute to reliable, scalable services across distributed systems.
The role emphasizes incident response, RCA leadership, and collaboration with cross-functional teams to reduce toil and improve availability and performance.
Posted Sunday, August 16, 2026 at 6:30 PM
Site Reliability Engineer with 3–5+ years of experience in designing, operating , and optimizing cloud-native platforms with a strong focus on Azure environments . Proven expertise in building highly available , scalable, and secure systems using Infrastructure as Code ( IaC ) and automation-first practices.
Experienced in managing application hosting architectures including Azure App Services, ASEv3, Application Gateway (AGW) and Azure Front Door , ensuring high performance and resilience across distributed systems.
Demonstrates an automation mindset by leveraging modern engineering tools and AI-assisted development platforms (e.g., GitHub Copilot, Microsoft Copilot) to accelerate delivery, reduce operational toil, and improve reliability standards — with careful validation of outputs for security and production readiness.
Understanding and experience in developing Azure function Apps, Azure logic Apps
Understanding of event triggers, event hub, service bus.
Cloud Architecture: High Availability, Fault Tolerance, Scalability Patterns
Incident Management and RCA
Incident Management, P1 troubleshooting, Change Management
Experienced in leading RCA and representing on the weekly call
SLA / SLO / Error Budget concepts
System Performance Optimization & Capacity Planning
Toil Reduction through Automation
Infrastructure as Code & Automation
API-based automation and orchestration
Observability & Monitoring
Azure Monitor, Log Analytics Workspace, Grafana, Site 24x7 (or similar SaaS based synthetic monitoring tool)
Application Insights
Alert tuning and signal-to-noise optimization
AI-Enabled Productivity (Not as Skill)
Code acceleration and script generation
Troubleshooting and log analysis assistance
Proven track record of workforce optimization leveraging AI tools.
Applying validation frameworks to ensure secure, accurate, and production-grade outputs
Deep understanding on version control
API integrations (REST, Postman, SoapUI)
Source control and release management
SRE & Reliability Engineering
Managed production environments ensuring high availability and reliability of cloud-hosted applications
Led incident response, performed deep root cause analysis , and implemented preventive measures to reduce recurrence
Improved system resilience through proactive monitoring and performance tuning strategies
Designed and supported application architectures using:
Azure App Services and App Service Plans
Azure App Service Environment v3 (ASEv3) for isolated, high-scale workloads
Azure Application Gateway (WAF-enabled) for L7 traffic management
Azure Front Door for global traffic routing and failover
Implemented secure and scalable cloud networking patterns , optimizing latency and throughput
Identified repetitive operational tasks and reduced manual effort through automation-first solutions
Developed automation using:
Terraform / Bicep / ARM templates
Azure Functions for event-driven workflows
Leveraged AI-assisted tools (GitHub Copilot, Copilot) to accelerate scripting and automation development, while ensuring strict validation for enterprise use
Built and enhanced observability using:
Created KQL-based queries and dashboards for proactive issue detection
Reduced false alerts by optimizing alert thresholds and improving signal quality
Analyzed application performance across distributed systems to identify bottlenecks
Implemented improvements through:
Scaling strategies (horizontal & vertical)
Network optimization (AGW / Front Door tuning)
Backend service improvements
Partnered with SRE, CloudOps, and development teams to design resilient systems
Contributed to runbooks, documentation, and operational standards
Enabled engineering teams by improving platform reliability and deployment pipelines
Bachelor’s Degree in Computer Science / Engineering or related field
Experience with microservices and distributed architectures
Working knowledge of AWS cloud services
AZ-700 Designing and Implementing Microsoft Azure Networking Solutions
FinThrive is advancing the healthcare economy. For the most recent information on FinThrive’s vision for healthcare revenue management visit finthrive.com/why-finthrive
At FinThrive we’re proud of our agile and committed culture, which makes FinThrive an exceptional place to work. Explore our latest workplace recognitions at https://finthrive.com/careers#culture
FinThrive is an Equal Opportunity Employer and ensures its employment decisions comply with principles embodied in Title VII, the Age Discrimination in Employment Act, the Rehabilitation Act of 1973, the Vietnam Veterans Readjustment Assistance Act of 1974, Executive Order 11246, Revised Order Number 4, and applicable state regulations.