Turn this role into an interview — a resume and cover letter built around what this employer wants.
Tilaka Nusa Teknologi is seeking a skilled IT Infrastructure Engineer to manage, maintain and ensure the availability of on-premises and virtual servers. The role focuses on patching, security updates, and routine upgrades following change management.
You will build monitoring dashboards with Grafana, configure alerts via Prometheus, manage Kubernetes deployments, and lead incident response with Graylog. Strong documentation and postmortem reporting are essential for continual improvement.
Bachelor's degree (S1) in Information Technology or a related field.
Minimum 2 years of work experience in the applied position or a related role.
Proficient in Linux server installation, configuration, and administration.
Experienced in server performance troubleshooting, including CPU, memory, disk I/O, and network.
Expertise in server hardening to enhance system security.
Proficient in managing backup & restore processes, as well as conducting periodic restore testing.
Proficient in application deployment and management within a Kubernetes environment.
Experienced in troubleshooting Kubernetes pods, services, and networking.
Understanding and proficiency in building Grafana dashboards and configuring alerts based on Prometheus metrics or other data sources.
Expertise in managing Graylog (Input, Stream, Pipeline, Index Set) and conducting log analysis for issue investigation.
Ability to prepare clear and structured technical documentation, runbooks, and postmortem reports.
Job Description
Manage, maintain, and ensure the availability of servers, both bare-metal and virtual machines.
Perform patching, security updates, and system upgrades on a scheduled basis according to change management procedures.
Execute and verify routine server backups, and perform periodic restore tests.
Build and maintain monitoring dashboards to track infrastructure and application health.
Formulate and refine alerting rules to ensure relevance and prevent alert fatigue.
Manage Graylog, including log aggregation, parsing, retention, and storage capacity.
Handle infrastructure incidents in compliance with SLAs, conduct troubleshooting, and restore services.
Perform root cause analysis and produce postmortem reports along with remediation plans.
Implement hardening standards across servers, containers, and Kubernetes clusters.
Remediate findings from vulnerability scans and security audits.
Automate routine tasks using scripts, configuration management tools, or AI.