Site Reliability Engineer

AssetMark Global Wealth

Hyderabad

On-site

INR 1,800,000 - 3,600,000

Full time

18 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AssetMark Global Wealth is expanding with a Global Capability Center in Hyderabad, seeking a Site Reliability Engineer to join the engineering team to ensure high availability, performance, and recoverability across production systems. This role emphasizes automation, observability, and reliability-driven design rather than traditional operations.

The SRE will influence architecture, manage incidents, and drive reliability improvements across cloud-native and legacy environments, partnering with

Qualifications

  • Bachelor's degree in computer science, software engineering, or related technical field.
  • 4–6 years of software engineering experience in Site Reliability Engineering, DevOps, Platform Engineering, or production operations.
  • Proven experience troubleshooting and improving production system reliability.
  • Experience supporting 24/7 systems, batch processing, and mission-critical workloads.
  • Strong collaboration across engineering, security, and infrastructure teams.
  • Experience working in Agile/Scrum environments.
  • Experience building APIs, services, and platform components.
  • Understanding of enterprise integration patterns, SOA, and large-scale system design.
  • Experience with DevOps practices, cross-functional collaboration, and agile/scrum development methodologies.

Responsibilities

  • Design, implement, and improve reliability, availability, and performance of critical AssetMark systems.
  • Define and operationalize SLIs, SLOs, and error budgets with engineering and product teams.
  • Participate in on-call rotations, incident response, and major incident management.
  • Lead blameless post-incident reviews, root cause analysis, and reliability improvements.
  • Proactively identify reliability risks and lead remediation efforts.
  • Build and maintain end-to-end observability across applications, infrastructure, and integrations.
  • Implement actionable monitoring and alerting to reduce noise and improve signal quality.
  • Partner with application teams to instrument services using best-in-class observability practices.
  • Ensure visibility into system health, capacity, performance, and failure modes across environments.
  • Embed security, compliance, and risk controls into operational practices.

Skills

SRE principles
Distributed systems
Azure/AWS/GCP
CI/CD pipelines
Infrastructure as Code
Multi-language dev

Education

Bachelor's degree in CS/Software Eng or related

Tools

Grafana
Azure Log Analytics
Prometheus
Datadog
Splunk
Kubernetes
Docker
JIRA
PagerDuty

Job description

AssetMark is a leading wealth management platform dedicated to empowering financial advisors and the investors they serve. As we continue to expand our global capabilities, we are establishing a Global Capability Center (GCC) in Hyderabad to enhance our operational excellence, strengthen strategic capabilities, and support our long-term growth.

We are seeking a Site Reliability Engineer (SRE) to join our engineering team. This role sits at the center of platform resilience — ensuring high availability, performance, recoverability, and operational maturity across AssetMark’s production systems.

This is not a traditional operations role. Our SREs are engineers first: designing automation, building observability frameworks, improving deployment safety, defining reliability standards, and reducing operational toil through code. You will influence architectural decisions, strengthen incident management practices, and raise the reliability bar across both legacy and cloud-native systems.

At AssetMark, reliability is a first-order expression of client obsession. Our SRE team plays a critical role in delivering the consistent, trusted technology experience that advisors depend on to run their businesses.

Key Responsibilities
Reliability Engineering & Operations
  • Design, implement, and continuously improve the reliability, availability, and performance of critical AssetMark systems (batch, APIs, integrations, and customer-facing platforms)
  • Define and operationalize SLIs, SLOs, and error budgets for critical services in partnership with engineering and product teams
  • Participate in on-call rotations, incident response, and major incident management
  • Lead and contribute to blameless post-incident reviews, driving root cause analysis and measurable reliability improvements
  • Proactively identify reliability risks and lead remediation efforts before they impact clients
Observability & Monitoring
  • Build and maintain end-to-end observability across applications, infrastructure, and integrations (metrics, logs, traces, alerts)
  • Implement actionable monitoring and alerting to reduce noise and improve signal quality
  • Partner with application teams to instrument services using best-in-class observability practices
  • Ensure visibility into system health, capacity, performance, and failure modes across environments
  • Hands-on experience with Grafana, Azure Log Analytics, Prometheus, Datadog, Splunk, or equivalents.
Automation & Toil Reduction
  • Identify repetitive operational tasks and automate them through code
  • Improve deployment reliability through automation, self-service tooling, and safe rollout patterns
  • Reduce manual intervention in batch processing, integrations, and operational workflows
  • Apply Infrastructure-as-Code and configuration automation to improve consistency and repeatability
Cloud, Platform & Infrastructure Reliability
  • Support reliability of Azure-based infrastructure, containerized workloads, and hybrid environments
  • Partner with platform, DevOps, and infrastructure teams to improve resilience, scalability, and recovery
  • Contribute to capacity planning, performance tuning, and cost-aware reliability decisions
  • Ensure systems meet RTO/RPO, backup, and disaster recovery expectations
Secure & Compliant Operations
  • Embed security, compliance, and risk controls into operational practices
  • Work closely with Security and Compliance teams to meet financial services regulatory requirements
  • Ensure production systems follow least privilege, secure configuration, and auditability standards
  • Support vulnerability remediation and secure operational processes
  • Deep understanding of environment reliability and stability engineering needs in High Availability, Fault Tolerance, Capacity Planning, Performance Engineering, Disaster Recovery, Multi-Region Architecture, Scalability, and Chaos Engineering
  • Partner with application engineering teams to improve production readiness and operational maturity
  • Influence system design by advocating for reliability-first architectural decisions
  • Provide guidance on operational best practices, deployment safety, and observability standards
  • Document operational patterns, runbooks, and reliability guidelines in Confluence
  • Act as a reliability advocate across AssetMark engineering teams
Knowledge, Skills, Abilities
  • Strong software engineering skills in .NET / C#, Python, Java, NodeJS, or similar
  • Experience operating distributed systems in production
  • Deep understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, incident management
  • Experience with Azure (or AWS/GCP), including compute, networking, and managed services
  • Knowledge of containerization and orchestration (Docker, Kubernetes preferred)
  • Experience with monitoring, logging, tracing, and alerting tools
  • Familiarity with CI/CD pipelines, automation, and Infrastructure-as-Code
  • Deep understanding of system administration, networking (TCP/IP, DNS, HTTPS, SSL / TLS), Storage, Load Balancing, and File Systems
  • Understanding of security best practices in regulated enterprise environments
  • Experience with JIRA, PagerDuty, or equivalent
  • Experience supporting financial services or highly regulated systems (preferred)
Education & Experience
  • Bachelor's degree in computer science, Software Engineering, or related technical field
  • 4-6 years of software engineering experience of experience in Site Reliability Engineering, DevOps, Platform Engineering, or production operations
  • Proven experience in troubleshooting and improving production system reliability
  • Experience supporting 24/7 systems, batch processing, and mission-critical workloads
  • Strong collaboration skills across engineering, security, and infrastructure teams
  • Experience working in Agile/Scrum environments
  • Experience building APIs, services, and/or platform components
  • Understanding of enterprise integration patterns, service-oriented architecture, and large-scale system design
  • Experience with DevOps practices, cross-functional collaboration, and agile/scrum development methodologies
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Associate Site Reliability Engineer
Associate Site Reliability Engineer

AssetMark Global Wealth • Hyderabad

Hybrid
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AssetMark Global Wealth • Hyderabad

On-site
INR 2,000,000 - 3,600,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SMC Squared • Hyderabad

On-site
INR 3,500,000 - 6,000,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Sr. Site Reliability Engineer I
Sr. Site Reliability Engineer I

MetLife • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

HighRadius • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Technologies Pvt. Ltd. • Pune District

On-site
INR 900,000 - 1,400,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000