Site Reliability Engineer/Production Support

Smart IT Frame LLC

New York (NY)

On-site

USD 120,000 - 160,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Smart IT Frame LLC in New York City seeks a Production Support/Site Reliability Engineer to ensure high availability of production apps and infrastructure. You will troubleshoot issues, analyze incidents, and collaborate with development teams to improve reliability.

Responsibilities include on-call duties, RCA execution, rollout support, and building automation. The role emphasizes observability, monitoring, and incident management to drive continuous improvement.

Qualifications

  • Strong Linux/Unix administration and troubleshooting skills.
  • Cloud platform experience (AWS/Azure/GCP).
  • Experience with monitoring/observability tools (Prometheus, Grafana, ELK, Splunk, Datadog, AppDynamics).
  • Hands-on Docker and Kubernetes experience.
  • CI/CD pipelines with Jenkins, GitLab CI, GitHub Actions, Azure DevOps, or similar.
  • Scripting in Python, Shell/Bash, or PowerShell.
  • Git and basic source-control practices.
  • Experience troubleshooting REST APIs, microservices, web apps, distributed systems.
  • Networking fundamentals (DNS, HTTP/HTTPS, TCP/IP, load balancers, firewalls).
  • SQL/Databases like PostgreSQL, MySQL, Oracle, SQL Server.
  • Familiarity with incident management tools (ServiceNow, Jira, PagerDuty).

Responsibilities

  • Monitor production apps, infra, services, and dbs for high availability and performance.
  • Provide L2/L3 production support and troubleshoot issues.
  • Analyze alerts, logs and metrics to resolve incidents.
  • Participate in 24x7 on-call and incident response as needed.
  • Perform RCA for major incidents and drive preventive actions.
  • Manage incidents, problems and service requests per SLAs.
  • Collaborate with dev teams to troubleshoot defects and production issues.
  • Support deployments, releases, rollbacks and changes.
  • Develop automation to reduce manual operational tasks.
  • Maintain monitoring, alerting, dashboards and observability solutions.
  • Identify bottlenecks and recommend improvements.
  • Contribute to disaster recovery, backup, failover, and business continuity.
  • Maintain runbooks, SOPs and troubleshooting guides.
  • Continuously improve reliability, scalability, resilience and efficiency.

Skills

Linux/Unix administration
Cloud environments
Monitoring/observability
Docker & Kubernetes
CI/CD pipelines
Python/Shell/PowerShell
Git & source control
REST APIs & microservices
Networking basics
SQL databases
Incident management tools

Tools

Prometheus
Grafana
ELK/Elastic Stack
Splunk
Datadog
AppDynamics

Job description

Job Title: Production Support/Site Reliability Engineer
Location: New York City, NY
Key Responsibilities
  • Monitor production applications, infrastructure, services, and databases to ensure high availability and performance.
  • Provide L2/L3 production support and troubleshoot application, infrastructure, and deployment issues.
  • Analyze alerts, logs, metrics, and traces to identify and resolve production incidents.
  • Participate in 24x7 on‑call / shift‑based support, including incident response and escalation when required.
  • Perform Root Cause Analysis (RCA) for major incidents and implement preventive actions.
  • Manage incidents, problems, and service requests according to defined SLAs.
  • Work with development teams to troubleshoot application defects and production issues.
  • Support application deployments, releases, rollbacks, and production changes.
  • Develop automation/scripts to eliminate repetitive manual operational activities.
  • Implement and maintain monitoring, alerting, dashboards, and observability solutions.
  • Identify performance, capacity, and reliability bottlenecks and recommend improvements.
  • Participate in disaster recovery, backup, failover, and business continuity activities.
  • Maintain operational documentation, runbooks, SOPs, and troubleshooting guides.
  • Continuously improve system reliability, scalability, resilience, and operational efficiency.
Required Technical Skills
  • Strong experience in Linux/Unix administration and troubleshooting.
  • Good knowledge of AWS / Azure / GCP cloud environments.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, ELK/Elastic Stack, Splunk, Datadog, or AppDynamics.
  • Hands‑on experience with Docker and Kubernetes.
  • Good understanding of CI/CD pipelines using Jenkins, GitLab CI, GitHub Actions, Azure DevOps, or similar tools.
  • Scripting/programming experience in Python, Shell/Bash, or PowerShell.
  • Strong knowledge of Git and source‑control practices.
  • Experience troubleshooting REST APIs, microservices, web applications, and distributed systems.
  • Good understanding of networking concepts such as DNS, HTTP/HTTPS, TCP/IP, load balancers, and firewalls.
  • Working knowledge of SQL and databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
  • Experience with incident management and ticketing tools such as ServiceNow, Jira, or PagerDuty.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Site Reliability Engineer & Production Support – On-Call
Site Reliability Engineer & Production Support – On-Call

Smart IT Frame LLC • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Atlanta (GA)

On-site
USD 110,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Glove • Cleveland (OH)

On-site
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Resolve Tech Solutions • Irving (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Jersey City (NJ)

Hybrid
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

System One • Dallas (TX)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000