Senior Server Administrator

Rhysley

New Delhi

On-site

INR 2,000,000 - 2,800,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rhysley in New Delhi, India seeks a highly experienced Senior Server Administrator to own production server infrastructure end-to-end. The role emphasizes availability, security, and automated operations, with proactive incident handling and disaster recovery planning. You will lead deployment, monitoring, and capacity planning, ensuring zero-downtime releases where possible.

Qualifications

  • 8+ years of hands-on experience in Senior Server Administration / System Administration.
  • Strong Linux server administration and production server ownership experience.
  • Experience with High Availability and failover architectures.
  • Experience with NGINX / HAProxy or equivalent load-balancing technologies.
  • Experience with server monitoring tools (Prometheus, Grafana, ELK) and alerting.
  • Scripting skills in Bash and/or Python for automation.
  • Server security hardening, SSH, firewalls, and patching.

Responsibilities

  • Own and manage production servers end-to-end with high availability and minimal downtime.
  • Design and maintain highly available multi-server environments with load balancing.
  • Implement monitoring, configure real-time alerts, and perform RCA for major incidents.
  • Perform security hardening, patching, and secure access controls (MFA/RBAC).
  • Oversee backups, disaster recovery, and regular restore testing.
  • Coordinate deployments with development teams for reliable releases.
  • Automate repetitive admin tasks using Bash/Python to improve reliability.

Skills

Linux server administration
High Availability
Load balancing
Incident management
Monitoring & alerting
Automation Scripting
Security hardening

Tools

NGINX
HAProxy
Prometheus
Grafana
ELK

Job description

Role Overview

We are looking for a highly experienced and hands-on Senior Server Administrator who will take end-to-end

ownership of our production server infrastructure.

This is not a routine support or maintenance role. The role is focused on ensuring server availability, stability,

security, performance, scalability, and reliability. The ideal candidate should be capable of independently

managing production infrastructure, handling critical incidents, implementing high-availability solutions, and

proactively preventing infrastructure failures.

The candidate should have strong hands-on server administration experience along with an SRE-oriented

approach to reliability, monitoring, automation, and proactive problem-solving.

Key Responsibilities
1. Production Server Ownership
  • Own and manage production servers end-to-end.
  • Ensure high availability, stability, performance, and minimum downtime.
  • Take complete responsibility for server health and reliability.
  • Monitor server capacity, performance, and resource utilization.
  • Manage server configuration, upgrades, patching, and lifecycle activities.
2. High Availability & Failover
  • Design and maintain highly available multi-server environments.
  • Configure and manage load balancing using NGINX / HAProxy.
  • Implement automatic failover and redundancy for critical services.
  • Ensure infrastructure remains available during server or service failures.
  • Regularly test failover and recovery procedures.
3. Monitoring & Incident Management
  • Implement and manage server monitoring using Prometheus, Grafana, ELK, or equivalent tools.
  • Configure real-time alerts for server, application, network, and service issues.
  • Respond quickly to production incidents and restore services with minimum impact.
  • Perform Root Cause Analysis (RCA) for major incidents.
  • Identify recurring issues and implement permanent solutions.
4. Server Security & Hardening
  • Perform server security hardening and secure configuration.
  • Manage SSH, firewalls, Fail2ban, access permissions, and related security controls.
  • Perform regular security patching and system updates.
  • Implement secure access mechanisms, including MFA/RBAC where required.
  • Identify and address server-level security vulnerabilities.
5. Backup & Disaster Recovery
  • Manage and monitor server backup systems.
  • Ensure backup integrity and availability.
  • Maintain disaster recovery procedures for critical systems.
  • Conduct regular backup restoration and recovery testing.
  • Ensure critical production services can be restored effectively after failures.
6. Deployment & Reliability
  • Manage production deployments with minimal or zero downtime.
  • Implement deployment validation and rollback strategies.
  • Coordinate with development teams for reliable production releases.
  • Identify performance and reliability issues before they affect users.
  • Continueously improve infrastructure stability and availability.
7. Performance & Capacity Management
  • Monitor CPU, memory, disk, network, and other server resources.
  • Identify and resolve performance bottlenecks.
  • Plan server capacity based on business and application requirements.
  • Recommend infrastructure scaling and optimization initiatives.
8. Automation
  • Automate repetitive server administration and operational activities.
  • Use Bash, Python, or equivalent scripting for automation.
  • Reduce manual intervention and operational dependency.
  • Improve efficiency, reliability, and consistency through automation.
9. Documentation & Operational Excellence
  • Maintain server architecture documentation, SOPs, recovery procedures, and configuration records.
  • Document major incidents, RCA findings, and preventive actions.
  • Maintain clear operational processes for production infrastructure.
  • Ensure knowledge is properly documented for business continuity.
Required Technical Skills
  • 8+ years of hands-on experience in Senior Server Administration / System Administration / Infrastructure Administration.
  • Strong hands-on experience managing production servers.
  • Strong experience with Linux server administration.
  • Experience with High Availability and failover architecture.
  • Hands-on experience with NGINX / HAProxy or equivalent load-balancing technologies.
  • Experience with Prometheus, Grafana, ELK, or equivalent monitoring tools.
  • Strong troubleshooting skills across server, system, and networking issues.
  • Good understanding of TCP/IP, DNS, HTTP/HTTPS, SSL/TLS, firewalls, and networking fundamentals.
  • Experience with server security hardening, SSH, firewall, Fail2ban, patching, and access management.
  • Strong understanding of backup and disaster recovery.
  • Experience with production deployments, rollback, and minimal-downtime strategies.
  • Good scripting skills in Bash and/or Python.
  • Understanding of server performance, capacity planning, and infrastructure scalability.
Mandatory Experience

Candidates must have demonstrated hands-on experience in:

  • Production server ownership and administration
  • Linux server administration
  • High Availability and automatic failover
  • NGINX / HAProxy or equivalent load balancing
  • Infrastructure/server monitoring and alerting
  • Production incident management and RCA
  • Server security hardening
  • Backup and Disaster Recovery
  • Server automation using Bash/Python
  • Performance optimization and capacity management
  • Independent production troubleshooting
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Server Maintenance Engineer
Server Maintenance Engineer

Rimes International • Chennai District

On-site
INR 700,000 - 1,100,000
Senior DevOps Engineer
Senior DevOps Engineer

FYN Tune Solution • Navi Mumbai, Mumbai, Thane

On-site
INR 1,800,000 - 3,000,000
System, Network & Data Security Administrator
System, Network & Data Security Administrator

Sharaan Infosystems • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Lead System Administrator
Lead System Administrator

SoftTech Engineers Ltd • Pune District

On-site
INR 1,200,000 - 1,800,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Server Administrator
Senior Server Administrator

Indecomm Global Services (India) • Bengaluru

On-site
INR 1,200,000 - 2,200,000
Linux Systems Administrator Ubuntu Red Hat
Linux Systems Administrator Ubuntu Red Hat

Uvation • India

Remote
INR 1,200,000 - 1,800,000
Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
Production and Support Engineer
Production and Support Engineer

Fulcrum Worldwide Software • Pune District

Hybrid
INR 1,200,000 - 1,800,000