Associate Architect - Site Reliability

Highradius

Hyderabad

On-site

INR 2,500,000 - 4,200,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Highradius is seeking an Associate Architect (9-12 Years) to join our Cloud Engineering team in Hyderabad. You will design and refine cloud infrastructure with focus on reliability, security, and scalability, balancing live production work with automation and system improvements.

The role emphasizes collaborating with software engineers, implementing monitoring, IaC, and automation, and guiding designs toward resilience while reducing manual effort.

Qualifications

  • Experience with at least one major cloud platform and cloud security basics.
  • Proven hands-on Linux administration (RHEL/CentOS/Rocky) and Windows support.
  • Experience with IaC tools and automation to enforce reliability.

Responsibilities

  • Design, build, and maintain resilient cloud infrastructure for scalable apps.
  • Lead reliability practices with monitoring and alerting; apply SLI/SLO/SLA.
  • Automate provisioning and standardize infrastructure; develop SOPs.
  • Participate in incident response and root-cause analyses.
  • Collaborate with cross-functional teams to incorporate cloud solutions and POCs.

Skills

Cloud platforms
SRE concepts (SLI/SLO/SLA)
Communication & collaboration
DevOps mindset & change advocacy

Tools

Docker
Kubernetes
Terraform
CloudFormation
Ansible
Puppet
Prometheus
Grafana
ELK Stack
OpenTelemetry
Rundeck/Jenkins
Git
ArgoCD
Crossplane

Job description

Job Summary:

We are looking for a highly skilled and adaptableAssociate Architect (9 - 12 Years) to become a key member of our Cloud Engineering team. In this crucial role, you will be instrumental in designing and refining our cloud infrastructure with a strong focus on reliability, security, and scalability. As an Architect, you'll apply software engineering principles to solve operational challenges, ensuring the overall operational resilience and continuous stability of our systems. This position requires a blend of managing live production environments and contributing to engineering efforts such as automation and system improvements.

Key Responsibilities:
  • Cloud Infrastructure Architecture and Management: Design, build, and maintain resilient cloud infrastructure solutions to support the development and deployment of scalable and reliable applications.
    This includes managing and optimizing cloud platforms for high availability, performance, and cost efficiency.
  • Enhancing Service Reliability: Lead reliability best practices by establishing and managing monitoring and alerting systems to proactively detect and respond to anomalies and performance issues. Utilize SLI, SLO, and SLA concepts to measure and improve reliability. Identify and resolve potential bottlenecks and areas for enhancement
  • Driving Automation and Efficiency: Contribute to the automation, provisioning, and standardization of infrastructure resources and system configurations. Identify and implement automation for repetitive tasks to significantly reduce operational overhead. Develop Standard Operating Procedures (SOPs) and automate workflows using tools like Rundeck or Jenkins.
  • Incident Response and Resolution: Participate in and help resolve major incidents, conduct thorough root cause analyses, and implement permanent solutions. Effectively manage incidents within the production environment using a systematic problem-solving approach.
  • Collaboration and Innovation: Work closely with diverse stakeholders and cross-functional teams, including software engineers, to integrate cloud solutions, gather requirements, and execute Proof ofConcepts (POCs). Foster strong collaboration and communication. Guide designs and processes with a focus on resilience and minimizing manual effort. Promote the adoption of common tooling andcomponents, and implement software and tools to enhance resilience and automate operations. Be open to adopting new tools and approaches as needed.
Required Skills and Experience:
  • Cloud Platforms: Demonstrated expertise in at least one major cloud platform (AWS, Azure, or GCP).
  • Extensive experience with containerization (Docker) and orchestration (Kubernetes) technologies.
  • Automation & IaC: Proficiency in scripting languages (shell and Python). Experience with configuration management tools (Ansible or Puppet). Must have exposure to Infrastructure as Code (IaC) tools(Terraform or CloudFormation).
  • Monitoring & Observability: Experience setting up and configuring monitoring tools (Prometheus, Grafana, or the ELK stack). Hands-on experience implementing OpenTelemetry for observability.
  • Familiarity with monitoring and logging tools for cloud-based applications.
  • Service Reliability Concepts: A strong understanding of SLI, SLO, SLA, and error budgeting.
  • Infrastructure Management: Proven proficiency in on-premises hosting and virtualization platforms (VMware, Hyper-V, or KVM). Solid understanding of storage internals (NAS, SAN, EFS, NFS) and protocols (FTP, SFTP, SMTP, NTP, DNS, DHCP). Experience with networking and firewall technologies.
  • Strong hands‑on experience with Linux internals and operating systems (RHEL, CentOS, Rocky Linux).
  • Experience with Windows operating systems to support varied environments.
  • Soft Skills & Mindset: Excellent communication and interpersonal skills for effective teamwork. We value proactive individuals who are eager to learn and adapt in a dynamic environment. Must possess apragmatic and adaptable mindset, with a willingness to step outside comfort zones and acquire new skills. Ability to consider the broader system impact of your work. Must be a change advocate for reliability initiatives.
Desired/Bonus Skills:
  • Experience with DevOps toolchain elements like Git, Jenkins, Rundeck, ArgoCD, or Crossplane.
  • Experience with database management, particularly MySQL and Hadoop.
  • Knowledge of cloud cost management and optimization strategies.
  • Exposure to Gen AI.
  • Understanding of cloud security best practices, including data encryption, access controls, and identity management.
  • Experience implementing disaster recovery and business continuity plans.
  • Familiarity with ITIL (Information Technology Infrastructure Library) processes
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ASSOCIATE ARCHITECT - Reliability Analysis
ASSOCIATE ARCHITECT - Reliability Analysis

Happiest Minds Technologies • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Site Reliability Engineer
Site Reliability Engineer

HighRadius • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Director Cloud & Infrastructure Architect (Multi-Cloud | Datacenter | SRE)
Director Cloud & Infrastructure Architect (Multi-Cloud | Datacenter | SRE)

Mancer Consulting Services • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Associate Architect
Associate Architect

Bitbybit Solutions • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Cloud & Enterprise Architect (Cloud Services & AI-Augmented Operations)
Senior Cloud & Enterprise Architect (Cloud Services & AI-Augmented Operations)

Bosch Global Software Technologies • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NCR Voyix • Chennai District

On-site
INR 3,000,000 - 5,400,000
Senior SRE Engineer
Senior SRE Engineer

EPAM Systems India Pvt Ltd • Chennai District

On-site
INR 1,800,000 - 3,200,000
Technical Architect (Cloud & Platform Engineering)
Technical Architect (Cloud & Platform Engineering)

Bosch Global Software Technologies • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Staff Engineer - Cloud Backend Engineering
Staff Engineer - Cloud Backend Engineering

Coupang • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Arcesium • Hyderabad, Bengaluru

Hybrid
INR 6,000,000 - 9,000,000