SRO Lead

Versant Media

Englewood Cliffs (NJ)

On-site

USD 140,000 - 180,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Versant Media in Englewood Cliffs, NJ is seeking a hands-on System Reliability Engineering Lead to elevate the reliability, performance, and scalability of our software, production, and platform systems. This player-coach role partners with Software, Platform, and Infrastructure teams to implement reliability best practices and end-to-end testing.

You will build testing frameworks, improve observability, and drive proactive risk reduction across services while mentoring engineers in reliability

Qualifications

  • 5+ years in SRE/DevOps/Infrastructure roles.
  • Hands-on experience with production systems troubleshooting.
  • Experience with integration and E2E testing frameworks.
  • Familiarity with observability, monitoring, and tracing tools.
  • Experience with cloud platforms (AWS, GCP, Azure) and distributed systems.
  • Experience with CI/CD pipelines and automation.

Responsibilities

  • Improve system reliability, availability, and performance across services.
  • Define and implement SLIs, SLOs, and reliability standards.
  • Identify and remediate reliability gaps in design and ops.
  • Contribute code, tooling, and automation for resilience.
  • Design and implement end-to-end testing workflows.
  • Build and maintain integration testing frameworks for cross-service deps.
  • Execute load and performance testing for peak conditions.
  • Integrate automated testing into CI/CD pipelines.
  • Establish practical testing standards and adoption.
  • Support performance benchmarking and capacity planning.
  • Analyze performance and bottlenecks across layers.
  • Optimize throughput and latency with infra/platform teams.
  • Implement and improve monitoring, logging, alerting.
  • Ensure systems are observable, debuggable, instrumented.
  • Participate in incident response and RCA.
  • Contribute to post-incident reviews.
  • Collaborate with multiple teams to embed reliability practices.
  • Provide guidance and hands-on support for testing and observability.
  • Standardize tools, frameworks, and processes.
  • Mentor engineers on reliability fundamentals and testing.

Skills

Site Reliability Engineering
DevOps
Systems Engineering
Infrastructure engineering
Production systems troubleshooting
Observability tools
Cloud platforms (AWS, GCP, Azure)
CI/CD pipelines & automation
Kubernetes
Docker
Terraform
CloudFormation
Incident management

Tools

Kubernetes
Docker
Terraform
CloudFormation
Jenkins / CI tools

Job description

Company Description

VERSANT (Nasdaq: VSNT) is an industry-changing media and entertainment business and home to trusted brands that shape culture, inform audiences, and build lasting connections. It operates across four core markets: political news and opinion, business news and personal finance, golf, and sports and genre entertainment. These markets are served through a powerful portfolio of iconic and innovative brands, including CNBC, MS NOW, USA Network, Golf Channel, Oxygen, E!, SYFY, and Versant's sports division USA Sports, along with complementary digital assets including Fandango, Rotten Tomatoes, GolfNow and GolfPass.

Job Description

The System Reliability Engineering (SRE) Lead is a hands‑on technical leader responsible for improving the reliability, performance, and scalability of VERSANT’s software, production, and platform systems.

Reporting to the VP of Infrastructure, this role works closely with Software Engineering, Production Engineering, Platform Engineering, and Infrastructure teams to implement reliability best practices, drive end-to-end testing, and ensure systems perform under real-world conditions.

This is a player‑coach role focused on execution - building testing frameworks, improving observability, and helping teams proactively identify and resolve system weaknesses before they impact production.

Key Responsibilities
Reliability & Engineering Practices
  • Partner with engineering teams to improve system reliability, availability, and performance.
  • Help define and implement SLIs, SLOs, and basic reliability standards across services.
  • Identify reliability gaps and work with teams to address risks in system design and operations.
  • Contribute directly to code, tooling, and automation that improves system resilience.
E2E & System Testing Execution
  • Design and implement end-to-end (E2E) testing workflows across distributed systems.
  • Build and maintain integration testing frameworks validating cross-service dependencies.
  • Execute and scale load and performance testing to validate systems under peak conditions.
  • Partner with teams to integrate automated testing into CI/CD pipelines.
  • Help establish practical testing standards and ensure adoption across teams.
Performance & Capacity
  • Support performance benchmarking and system capacity planning efforts.
  • Analyze system performance and identify bottlenecks across application and infrastructure layers.
  • Partner with infrastructure and platform teams to optimize system throughput and latency.
Observability & Operations
  • Implement and improve monitoring, logging, and alerting across services.
  • Help ensure systems are observable, debuggable, and well‑instrumented.
  • Participate in incident response and support root cause analysis efforts.
  • Contribute to post‑incident reviews and track follow‑up actions to improve reliability.
Cross‑Team Collaboration
  • Work closely with software, platform, enterprise and production engineering teams to embed reliability practices into day‑to‑day development.
  • Provide guidance and hands‑on support for testing, observability, and performance improvements.
  • Help standardize tools, frameworks, and processes used across teams.
  • Mentor engineers on reliability engineering fundamentals and testing best practices.
Qualifications
Basic Requirements
  • 5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Infrastructure roles.
  • Strong hands‑on experience operating and troubleshooting production systems.
  • Experience implementing integration testing, E2E testing, or performance/load testing frameworks.
  • Familiarity with observability tools (metrics, logging, tracing) and monitoring systems.
  • Experience with cloud platforms (AWS, GCP, or Azure) and distributed systems.
  • Experience working with CI/CD pipelines and automation.
  • Strong debugging and problem‑solving skills in complex systems.
Desired Characteristics
  • Experience supporting media, broadcast, or real‑time production systems.
  • Familiarity with high‑throughput or low‑latency systems.
  • Exposure to SRE concepts such as SLIs/SLOs and incident management practices.
  • Experience with containerized environments (Kubernetes, Docker) is a plus.
  • Familiarity with infrastructure as code (Terraform, CloudFormation).
  • Strong collaboration skills and ability to work across multiple engineering teams.
Additional Information

VERSANT Media's policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.

If you are a qualified individual with a disability or a disabled veteran and require support throughout the application and/or recruitment process as a result of your disability, you have the right to request a reasonable accommodation. You can submit your request to candidateaccessibility@versantmedia.com.

VERSANT Media is committed to fair and equitable compensation practices. We include a good faith pay range for each position to comply with applicable state and local pay transparency laws and to promote equity across our organization. Actual compensation will be based on factors such as the candidate's skills, qualifications, experience, and location and may include additional forms of compensation and benefits such as health insurance, retirement plans, paid time off, etc.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRO Lead
SRO Lead

Versant • New Jersey

On-site
USD 150,000 - 210,000
Director, Platform SRO
Director, Platform SRO

Versant Media • New York (NY)

On-site
USD 180,000 - 240,000
Sr. Director, Platform Engineering
Sr. Director, Platform Engineering

Versant Media, LLC • Englewood Cliffs (NJ)

On-site
USD 220,000 - 240,000
Cloud Reliability Engineer
Cloud Reliability Engineer

Versant • New Jersey

Hybrid
USD 95,000 - 125,000
Free employee parking
On-site fitness center
Shuttle services available
Software Engineer
Software Engineer

Versant Media, LLC • New York (NY)

Hybrid
USD 120,000 - 150,000
On-site fitness center
Free employee parking
Electric vehicle charging stations
Senior Manager, Strategic Operations & AI Enablement
Senior Manager, Strategic Operations & AI Enablement

Versant Media • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Principal Software Engineer
Principal Software Engineer

Versant • New Jersey

On-site
USD 120,000 - 180,000
Operations Technical Manager
Operations Technical Manager

Versant • California (MO)

On-site
USD 90,000 - 120,000
Competitive compensation package
Comprehensive health benefits
Bonus eligibility
+2
Senior Software Engineer
Senior Software Engineer

Versant Media • Orlando (FL)

Remote
USD 100,000 - 130,000
Health insurance
Retirement plans
Paid time off
Senior Manager, Strategic Operations & AI Enablement
Senior Manager, Strategic Operations & AI Enablement

Versant • Town of Florida (NY)

Hybrid
USD 120,000 - 180,000