A complete application in a minute — tailored resume and cover letter, ready to send.
Siri InfoSolutions Inc in Morristown, NJ seeks a Platform Reliability Engineer to apply SRE practices to enterprise platforms, ensuring availability, security, and scalable ops.
You will own incident response, release discipline, patching, and capacity planning, collaborating with IT ops, development, and vendors to deliver reliable, observable platforms.
Skill: Platform Reliability Engineer
Must Have Technical/Functional Skills:
Required Skills and Qualifications:
3 7 years of experience in platform engineering, IT operations, production support, DevOps, SRE, service management, or related technology roles.
Strong knowledge of platform health, release management, patching, upgrades, regression testing, rollback planning, and production readiness.
Hands-on understanding of SRE concepts such as SLIs, SLOs, error budgets, incident command, postmortems, toil reduction, and reliability reviews.
Experience troubleshooting production issues using logs, metrics, traces, alerts, telemetry, dependency maps, and configuration data.
Knowledge of observability, dashboards, alert tuning, anomaly detection, capacity management, performance analysis, and invocation frequency trends.
Strong documentation, automation mindset, analytical thinking, communication, ownership, and cross-functional collaboration skills.
Roles & Responsibilities:
The Platform Reliability Engineer operates and improves enterprise platforms that support IT operations, automation, observability, AI-enabled workflows, and business-critical services.
The role ensures reliable, secure, scalable, and well-governed platforms through strong administration, SRE practices, release discipline, vendor coordination, self-service enablement, and continuous improvement.
Core Responsibilities:
Administer platforms, access, integrations, configuration, onboarding, agent registration, inventory, dependency mapping, and lifecycle records.
Maintain platform availability, performance, capacity, recoverability, observability, security posture, and operational readiness.
Plan and execute upgrades, patching, releases, regression testing, rollback readiness, and post-release validation.
Manage vendor escalations, licensing consumption, platform utilization, plugin governance, roadmap alignment, and service improvement actions.
Enable self-service through workflows, catalogs, knowledge articles, automation, runbooks, and simplified support processes.
Support ITSM practices for incidents, changes, problems, requests, assets, configuration, releases, and knowledge management. make a jd for job post