Cloud Reliability Engineer

Versant Media, LLC

New York (NY)

Hybrid

USD 135,000 - 165,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Access to on-site fitness center
Free employee parking
Shuttle service from Manhattan

Job summary

Versant Media, LLC is looking for a Cloud Reliability Engineer in New York City. This role is vital for ensuring the availability and performance of Versant's cloud platforms and services, working closely with engineering teams and ensuring operational excellence.

The ideal candidate will have strong AWS experience and skills in automation and incident management. The position offers a hybrid work model and opportunities for career growth within a leading media company.

Qualifications

  • 3-7 years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or related roles.
  • Strong hands-on experience with enterprise-scale AWS environments.
  • Experience with monitoring and observability platforms.

Responsibilities

  • Design and maintain reliability practices for cloud infrastructure.
  • Monitor service-level objectives and operational metrics.
  • Participate in incident response and root cause analysis activities.

Skills

AWS
Incident Management
Automation
Infrastructure as Code
Monitoring
Scripting

Education

Bachelor’s degree in Computer Science, Engineering, Information Systems

Tools

Terraform
CloudFormation
CI/CD platforms

Job description

  • Compensation: USD 135,000 - USD 165,000 - yearly
Company Description

VERSANT (Nasdaq: VSNT) is an industry-changing media and entertainment business and home to trusted brands that shape culture, inform audiences, and build lasting connections. It operates across four core markets: political news and opinion, business news and personal finance, golf, and sports and genre entertainment. These markets are served through a powerful portfolio of iconic and innovative brands, including CNBC, MS NOW, USA Network, Golf Channel, Oxygen, E!, SYFY, and Versant's sports division USA Sports, along with complementary digital assets including Fandango, Rotten Tomatoes, GolfNow and GolfPass.

Job Description

The Cloud Reliability Engineer is responsible for ensuring the availability, performance, scalability, and operational excellence of VERSANT’s cloud platforms and services.

This role works closely with cloud engineering, application development, networking, security, and operations teams to build and maintain highly reliable systems across a large multi-account AWS environment. The engineer will leverage automation, observability, and reliability engineering practices to improve platform resilience, reduce operational risk, and enhance the customer experience.

As a leading media company, VERSANT operates digital products, streaming platforms, content delivery systems, and media workflows that demand high levels of uptime and performance. The Cloud Reliability Engineer will help ensure these services remain resilient, scalable, and operationally mature.

The ideal candidate has strong experience with AWS, monitoring and observability platforms, incident management, automation, infrastructure as code, and operational best practices. Experience with AWS Organizations, Control Tower, Identity Center, Terraform, and modern cloud operations tooling is highly desirable.

Responsibilities
  • Design, implement, and maintain reliability practices for cloud infrastructure and platform services.
  • Define and monitor service-level objectives (SLOs), service-level indicators (SLIs), and operational metrics.
  • Identify reliability risks and implement solutions that improve availability, scalability, and resilience.
  • Drive continuous improvement initiatives focused on operational excellence and system stability.
Monitoring, Observability & Performance
  • Design and maintain monitoring, logging, alerting, and observability solutions across AWS environments.
  • Develop dashboards and reporting that provide visibility into platform health and performance.
  • Analyze system behavior, identify bottlenecks, and implement performance improvements.
  • Establish proactive monitoring practices that detect issues before they impact customers.
Incident Response & Operational Excellence
  • Participate in incident response, troubleshooting, and root cause analysis activities.
  • Lead post-incident reviews and identify corrective actions to prevent recurrence.
  • Improve operational processes, runbooks, and recovery procedures.
  • Support disaster recovery and business continuity initiatives.
AWS Platform Reliability
  • Support the reliability and operational health of large-scale AWS environments utilizing AWS Organizations, Control Tower, and Identity Center.
  • Partner with cloud engineering teams to improve platform architecture, resiliency, and operational consistency.
  • Assist in maintaining secure, scalable, and highly available cloud services.
Automation & Infrastructure as Code
  • Develop automation that reduces operational toil and improves system reliability.
  • Support infrastructure-as-code solutions using Terraform, CloudFormation, and related technologies.
  • Automate operational workflows, monitoring, remediation, and recovery activities.
  • Contribute to CI/CD pipelines and deployment automation initiatives.
Media & Digital Platform Reliability
  • Support the reliability of streaming platforms, content delivery systems, media workflows, APIs, and customer-facing applications.
  • Collaborate with engineering teams to improve application reliability and operational readiness.
  • Assist in capacity planning and scaling efforts for high-traffic events and media workloads.
  • Partner with cloud, networking, security, and application teams to identify and address operational risks.
  • Promote reliability engineering best practices throughout the organization.
  • Contribute to documentation, standards, and operational procedures.
  • Evaluate emerging technologies and recommend improvements to platform reliability and observability.
Qualifications

Bachelor’s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.

3–7 years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, Infrastructure Engineering, or related roles.

Strong hands-on experience with AWS cloud services and enterprise-scale AWS environments.

Experience with
  • Monitoring and observability platforms
  • Incident management and root cause analysis
  • Operational troubleshooting and performance tuning
  • AWS Organizations, Control Tower, and Identity Center
  • Infrastructure as Code: Terraform, CloudFormation
  • CI/CD platforms and deployment automation
  • Scripting and automation using Python, PowerShell, Bash, or similar languages
  • AWS networking, resiliency, and cloud architecture concepts
  • Logging, metrics, tracing, and alerting technologies

Strong troubleshooting, communication, and collaboration skills.

Additional Information

Location: New York City, NY or Englewood Cliffs, NJ – (Hybrid – 3 days onsite)

Employees based in our Englewood Cliffs office have access to a variety of convenient on-site services and amenities, including free employee parking, complimentary electric vehicle charging stations, an on-site fitness center with locker rooms and towel service, and dry‑cleaning drop‑off and pick‑up. To support an easier commute, shuttle service is also available to and from the Englewood Cliffs campus, with pickup locations across Manhattan’s East Side and West Side, Brooklyn, Hoboken, Jersey City, Newark, and Secaucus.

Equal Employment Opportunity

VERSANT Media’s policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.

If you are a qualified individual with a disability or a disabled veteran and require support throughout the application and/or recruitment process as a result of your disability, you have the right to request a reasonable accommodation. You can submit your request to candidateaccessibility@versantmedia.com.

VERSANT Media is committed to fair and equitable compensation practices. We include a good faith pay range for each position to comply with applicable state and local pay transparency laws and to promote equity across our organization. Actual compensation will be based on factors such as the candidate's skills, qualifications, experience, and location and may include additional forms of compensation and benefits such as health insurance, retirement plans, paid time off, etc.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Reliability Engineer
Cloud Reliability Engineer

Versant • New Jersey

Hybrid
USD 95,000 - 125,000
Free employee parking
On-site fitness center
Shuttle services available
Software Engineer
Software Engineer

Versant Media, LLC • New York (NY)

Hybrid
USD 120,000 - 150,000
On-site fitness center
Free employee parking
Electric vehicle charging stations
Sr. Platform Ops Engineer
Sr. Platform Ops Engineer

Versant Media, LLC • Englewood Cliffs (NJ)

On-site
USD 120,000 - 150,000
Health insurance
Retirement plans
Paid time off
Senior Software Engineer - Platforms
Senior Software Engineer - Platforms

Versant • New York (NY)

On-site
USD 110,000 - 160,000
Onsite fitness center
Gourmet cafeteria
Free shuttle transportation
Senior Software Engineer - Platforms
Senior Software Engineer - Platforms

Versant Media, LLC • New York (NY)

On-site
USD 110,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
SRO Lead
SRO Lead

Versant Media • Englewood Cliffs (NJ)

On-site
USD 140,000 - 180,000
Principal Software Engineer
Principal Software Engineer

Versant Media, LLC • Englewood Cliffs (NJ)

On-site
USD 165,000 - 225,000
Sr. Director, Platform Engineering
Sr. Director, Platform Engineering

Versant Media, LLC • Englewood Cliffs (NJ)

On-site
USD 220,000 - 240,000
Principal Software Engineer (Cloud Native)
Principal Software Engineer (Cloud Native)

Versant • New York (NY)

Hybrid
USD 170,000 - 200,000
Medical, dental, and vision insurance
401(k)
Paid leave
+1
Copy of Principal Software Engineer (Cloud Native)
Copy of Principal Software Engineer (Cloud Native)

Versant • New York (NY)

Hybrid
USD 170,000 - 200,000
Medical, dental, and vision insurance
401(k)
Paid leave
+2