Senior Site Reliability Engineer

Cisco

Bengaluru

On-site

INR 4,000,000 - 6,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cisco in Bengaluru seeks a Senior Site Reliability Engineer to lead cloud cost optimization and efficiency initiatives for the ThousandEyes cloud platform. You will partner with engineering, finance, and platform teams to drive FinOps practices, cost visibility, and governance across AWS services.

You will analyze usage across compute, storage, databases, and data pipelines, guiding capacity planning and autoscaling efforts while improving reliability and performance.

Qualifications

  • Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
  • 8–12 years of experience in Site Reliability Engineering, Cloud Infrastructure, or related fields.
  • Strong AWS experience (EC2, S3, RDS, EMR, Lambda, CloudWatch, OpenSearch, ElastiCache, IAM, VPC).
  • Practical experience in cloud cost optimization, FinOps, budgeting, forecasting, and cost governance.
  • Experience with IaC and automation tools (Terraform, CloudFormation, Puppet, Ansible).
  • Scripting in Python/Go/Shell for automation and tooling.
  • Experience with observability/monitoring platforms (CloudWatch, Prometheus, Grafana, Datadog, OpenSearch).
  • Strong Linux and distributed systems knowledge.
  • Incident management and production reliability experience.
  • Ability to translate data into engineering recommendations.
  • Excellent cross-functional communication in Agile environments.

Responsibilities

  • Own AWS cost optimization and efficiency initiatives across ThousandEyes infrastructure.
  • Analyze cloud usage and identify savings opportunities across compute, storage, and data platforms.
  • Collaborate with engineering to improve performance while reducing cloud waste.
  • Drive FinOps practices including tagging hygiene, budgeting, forecasting, and cost governance.
  • Build dashboards, guardrails, and automation for cost governance and visibility.
  • Provide senior technical leadership and mentor engineers across teams.
  • Balance reliability, scalability, performance, and cost in platform decisions.

Skills

Cloud infrastructure
FinOps
Automation
Python

Education

Bachelor's degree or higher in Engineering/CS

Tools

Terraform
CloudFormation
Puppet/Ansible
OpenSearch/Observability tools

Job description

Meet the Team

The ThousandEyes Efficiency & Performance team is responsible for improving the reliability, scalability, performance, and cost efficiency of the ThousandEyes cloud platform. The team partners closely with engineering, platform, finance, product, and leadership team members to drive infrastructure optimization, cloud cost governance, performance improvements, and operational excellence across AWS environments.

As a Senior Site Reliability Engineer in the Efficiency & Performance team, you will play a key role in managing and optimizing ThousandEyes cloud infrastructure cost, improving AWS efficiency, driving FinOps practices, and helping engineering teams make data-driven decisions around cost, performance, and scalability.

Your Impact
  • Own and drive AWS cost optimization and efficiency initiatives across ThousandEyes infrastructure.
  • Analyze cloud usage across compute, storage, databases, observability, networking, and data platform workloads to identify savings opportunities.
  • Partner with engineering teams to improve application and infrastructure performance while reducing cloud wastage.
  • Drive FinOps practices including cost visibility, tagging hygiene, showback/chargeback, budgeting, forecasting, anomaly detection, and cost allocation.
  • Improve infrastructure efficiency through better resource utilization, autoscaling, right-sizing, and capacity planning.
  • Build automation, dashboards, reports, and guardrails to improve cost governance and operational visibility.
  • Support ThousandEyes cost reviews, OKR tracking, leadership updates, and stakeholder communications.
  • Collaborate with product, finance, engineering, and platform teams to align cost optimization with business priorities.
  • Identify and reduce underutilized infrastructure, idle resources, over-provisioned workloads, and inefficient service usage.
  • Improve observability, alerting, and monitoring for cost, performance, and platform health indicators.
  • Influence engineering teams to adopt cost-aware architecture, reliable design patterns, and performance-efficient implementation practices.
  • Provide senior-level technical leadership, mentor engineers, and independently drive cross-functional initiatives to closure.
  • Balance reliability, scalability, performance, and cost efficiency in all platform decisions.
Minimum Qualifications
  • Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
  • 8–12 years of relevant experience in Site Reliability Engineering, Cloud Infrastructure, Platform Engineering, DevOps, Production Engineering, or Performance Engineering.
  • Strong experience with AWS services such as EC2, S3, RDS, EMR, Lambda, CloudWatch, OpenSearch, ElastiCache, IAM, VPC, and related cloud-native services.
  • Practical experience in cloud cost optimization, FinOps, AWS billing analysis, cost allocation, tagging, budget tracking, forecasting, and cost governance.
  • Experience with performance analysis, capacity planning, infrastructure optimization, and reliability improvements for large-scale cloud platforms.
  • Experience with infrastructure-as-code and automation tools such as Terraform, CloudFormation, Puppet, Ansible, or similar.
  • Strong scripting or programming skills in Python, Go, Shell, or similar languages for automation, reporting, and operational tooling.
  • Experience with observability and monitoring platforms such as ThousandEyes, CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar.
  • Strong Linux systems knowledge and understanding of distributed systems.
  • Experience in incident management, production support, reliability engineering, and operational excellence.
  • Ability to analyze large-scale infrastructure, performance, and cost data and convert findings into actionable engineering recommendations.
  • Strong communication skills with the ability to present technical, performance, and cost insights to engineering teams, finance, leadership, and multi-functional stakeholders.
  • Experience working in Agile/Scrum environments and managing priorities across multiple teams.
Preferred Qualifications
  • Experience supporting or optimizing large-scale SaaS platforms in AWS.
  • FinOps certification or equivalent experience with cloud financial management practices.
  • Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle optimization, workload right-sizing, and commitment planning.
  • Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloud-ability, Cloud-Health, or similar platforms.
  • Experience with performance tuning of cloud infrastructure, distributed systems, databases, data pipelines, and high-scale services.
  • Experience with data platforms or large-scale infrastructure components such as EMR, Kafka, OpenSearch, RDS, Airflow, Spark, Redis, Cassandra, or similar.
  • Experience driving cross-team cost optimization programs, efficiency OKRs, governance reviews, executive reporting, and measurable savings outcomes.
  • Ability to influence engineering teams toward cost-conscious architecture, performance-aware design, and operational standard methodologies.
  • Exposure to AI/ML infrastructure cost optimization, capacity planning, GPU/accelerator cost governance, or AI-driven efficiency tooling is a plus.
Why Cisco?

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.

We are Cisco, and our power starts with you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer, Data Platform - AI/ML Infrastructure
Lead Site Reliability Engineer, Data Platform - AI/ML Infrastructure

Cisco Systems, Inc. • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Software Engineer- Site Reliability Engineer (8+ Yrs)
Software Engineer- Site Reliability Engineer (8+ Yrs)

Cisco Systems, Inc. • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Software Engineer
Software Engineer

Cisco • Hyderabad

On-site
INR 1,500,000 - 2,700,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Cisco • Hyderabad

On-site
INR 2,500,000 - 3,500,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Cisco • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Data Engineer
Data Engineer

Cisco Systems, Inc. • Bengaluru

On-site
INR 900,000 - 1,500,000
Hybrid work model
Volunteer time off (80 hours/yr)
Leader, Software Engineering - Go/Python - Container/Kubernetes - SaaS/Public cloud - AI - 12 t[...]
Leader, Software Engineering - Go/Python - Container/Kubernetes - SaaS/Public cloud - AI - 12 t[...]

Cisco • Bengaluru

On-site
INR 2,500,000 - 3,800,000
DevOps Engineer – Any database stack | AWS | Terraform | Ansible | Jenkins (Mandate Skils) (4–8[...]
DevOps Engineer – Any database stack | AWS | Terraform | Ansible | Jenkins (Mandate Skils) (4–8[...]

Cisco • Bengaluru

On-site
INR 1,200,000 - 2,500,000
Data Engineer
Data Engineer

Cisco • Telangana

On-site
INR 1,200,000 - 1,800,000
AI Operations Engineering Technical Leader
AI Operations Engineering Technical Leader

Cisco • Bengaluru

On-site
INR 3,500,000 - 5,500,000