Operational Support Engineer

Russell Tobin

Atlanta (GA)

Hybrid

USD 82,656 - 89,544

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Russell Tobin is looking to hire an Operational Support Engineer for a hybrid role in Atlanta, GA. The successful candidate will take ownership of customer-impacting incidents and lead their resolution while working directly on production systems.

This position requires a minimum of 5 years of operational experience in video streaming platforms, strong troubleshooting skills, and familiarity with cloud infrastructure. The role demands participation in a 24/7 on-call rotation and offers a pay rate of $60 to $65 per hour based on experience.

Qualifications

  • 5+ years of experience in operational or support roles.
  • Ability to own complex incidents autonomously.
  • Calm and structured demeanor under high-pressure situations.

Responsibilities

  • Take ownership of customer incidents and drive resolutions.
  • Troubleshoot high-impact production issues.
  • Collaborate with Engineering teams for incident management.

Skills

Operational expertise in production video streaming platforms
Strong troubleshooting skills across distributed systems
Experience with monitoring and alerting tools
Knowledge of HLS, DASH, and CDN architectures

Tools

Terraform
Kubernetes
Grafana
Prometheus

Job description

Our client, a American technology corporation Client, is looking to hire an Operational Support Engineer in Atlanta, GA (Hybrid role).

Pay Rate Range: $60/h - $65/h on W2, depending on experience.

6 months W2 contract.

Role Summary

The team is responsible for the stability, availability, and operational excellence of our 24/7 live video streaming, ads, player, and real-time delivery platforms. As an Operational Support Engineer (L2), you take end‑to‑end ownership of customer‑impacting production incidents once they are triaged by Level 1 support. You operate directly on production systems, lead live incident resolution, and act as the operational bridge between Support, Engineering, DevOps, and customers, particularly during high‑impact live events. This is a hands‑on, customer‑facing role focused on incident ownership, production operations, automation, and operational scalability, not just reactive troubleshooting.

Incident & Operational Support
  • Take ownership of escalated customer issues from Level 1 Support and drive them to resolution.
  • Troubleshoot and resolve complex, high‑impact production incidents affecting live streams, VOD playback, ad insertion, DRM, and real‑time WebRTC services.
  • Operate directly on production environments, including configuration changes, CDN adjustments, and corrective actions, following established operational procedures, including executing mitigations and emergency changes during live incidents when customer impact requires immediate action.
  • Lead or actively contribute to live incident bridges involving customers, internal teams, and partners.
  • Provide clear, timely communication during incidents, including status updates and customer‑facing explanations.
Infrastructure as Code (IaC) & Production Operations
  • Work fluently with Infrastructure as Code (IaC) to understand, troubleshoot, and safely modify production environments.
  • Leverage tools and frameworks such as:
    • Terraform
    • Helm
    • Kubernetes manifests
    • GitOps workflows
    • CI/CD and deployment pipelines
  • Use IaC as the primary mechanism for safe, auditable, and repeatable operational changes.
  • Collaborate with Engineering and DevOps to improve deployment reliability and operational safety.
  • Validate and execute infrastructure or configuration changes through codified workflows.
AI‑Driven Operations & Automation
  • Leverage AI tools and automation to enhance operational efficiency and incident response.
  • Contribute to and use:
    • AI‑assisted incident triage and classification
    • Automated runbook execution
    • AI‑based pattern detection across incidents
    • Intelligent alert correlation and noise reduction
  • Use AI to:
    • Generate or improve incident communications
    • Accelerate troubleshooting workflows
    • Identify recurring patterns and systemic issues
  • Drive adoption of automation‑first and AI‑augmented operational practices.
Pre‑Event Planning & Operational Readiness
  • Participate in pre‑event readiness planning for critical customer events.
  • Validate system readiness through:
    • Runbook checks
    • Monitoring coverage validation
    • Risk identification and mitigation planning
  • Define and rehearse incident response strategies for high‑risk scenarios.
  • Collaborate with customers and internal teams to ensure smooth event execution.
On‑Call & 24/7 Operations
  • Participate in a 24/7 on‑call rotation, including nights, weekends, and holidays, as part of a global support model.
  • Ensure smooth handovers between shifts and regions.
  • Respond to critical alerts within defined SLAs for stream health, player errors, and delivery infrastructure.
Root Cause Analysis (RCA) & Continuous Improvement
  • Perform or contribute to root cause analysis (RCA) for production incidents.
  • Document findings, corrective actions, and preventive measures.
  • Identify recurring issues and work with Engineering and Product teams to eliminate them permanently.
  • Contribute to and improve runbooks, operational playbooks, and knowledge bases for all OptiView products (Player, ads, live and real‑time streaming).
Collaboration & Engineering Feedback Loop
  • Work closely with Engineering teams to elevate defects, validate fixes, and support production deployments.
  • Provide feedback on system observability, tooling gaps, and operational risks.
  • Act as the operational voice during post‑incident reviews.
Required Skills & Experience
Technical Skills
  • 5+ years of relevant experience in operational, support, or similar customer‑facing roles.
  • Proven ability to own complex problems end‑to‑end and operate with a high degree of autonomy.
  • Strong experience supporting production video streaming platforms, OTT services, and live systems.
  • Solid troubleshooting skills across distributed systems (APIs, microservices, cloud infrastructure).
  • Familiarity with HLS, DASH, CMAF, WebRTC, DRM, and CDN architectures.
  • Experience working with monitoring, alerting, and logs to diagnose live incidents (Grafana, Kibana/ELK, Prometheus, Loki).
  • Correlate backend streaming metrics, player telemetry, and CDN signals to diagnose live customer issues end‑to‑end.
  • Comfort performing controlled changes in production environments.
  • Working knowledge of incident management and on‑call operations.
Operational Mindset
  • Proven ability to remain calm, structured, and decisive during high‑pressure incidents.
  • Strong sense of ownership and accountability for customer outcomes.
  • Excellent written and verbal communication skills, including customer‑facing communication during incidents.
Equal Employment Opportunity

Russell Tobin is an equal opportunity employer. We do not discriminate on the basis of the race, religious creed, color, national origin, ancestry, physical disability, mental disability, reproductive health decision making, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, age, sexual orientation, veteran or military status, or any other characteristic protected by applicable federal, state, or local law.

Fair Chance Employment

Russell Tobin is a Fair Chance employer. We consider all qualified applicants, including those with criminal histories, in a manner consistent with applicable state and local Fair Chance laws and ordinances, including the California Fair Chance Act and all applicable local Fair Chance ordinances.

Accommodations

We are committed to providing reasonable accommodations to applicants and employees with disabilities. If you require a reasonable accommodation to participate in the application or interview process, or to perform the essential functions of this role, please contact us.

Applicable for San Francisco Candidates

Under the San Francisco Lactation in the Workplace Ordinance, we will provide written notice of lactation accommodation rights, and this notice will automatically be given upon hiring, any inquiry of parental leave or lactation accommodation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Operational Support Engineer
Staff Operational Support Engineer

Dolby • Atlanta (GA)

On-site
USD 136,000 - 188,000
Flex Work
Bonuses
Health benefits
Staff Operational Support Engineer
Staff Operational Support Engineer

Cypress HCM • Atlanta (GA)

On-site
USD 120,000 - 180,000
Staff Operational Support Engineer
Staff Operational Support Engineer

Dolby Laboratories • Atlanta (GA)

On-site
USD 136,000 - 188,000
Sr Staff Operational Support Engineer
Sr Staff Operational Support Engineer

Dolby • Atlanta (GA)

On-site
USD 152,000 - 210,000
Flexible work approach
Excellent compensation and benefits
Live Streaming Operator
Live Streaming Operator

Advanced Systems Group, LLC • Los Gatos (CA)

On-site
USD 50,000 - 80,000
Streaming Operations Specialist
Streaming Operations Specialist

Advanced Systems Group • Los Gatos (CA)

On-site
USD 60,000 - 75,000
Staff Operational Support Engineer Atlanta, Georgia,United States Posted a month ago
Staff Operational Support Engineer Atlanta, Georgia,United States Posted a month ago

Via Licensing Corporation • Atlanta (GA)

On-site
USD 136,000 - 188,000
Bonus
Benefits
Equity
Sr Staff Operational Support Engineer
Sr Staff Operational Support Engineer

Via Licensing Corporation • Atlanta (GA)

On-site
USD 152,000 - 210,000
Bonus
Equity
Comprehensive Benefits
Media Engineer
Media Engineer

Russell Tobin • New York (NY)

Hybrid
USD 77,000 - 92,000
Healthcare coverage (medical, dental,
Vision plans
401(k) retirement savings
+4
Live Command Center Operator (LCC Operator)
Live Command Center Operator (LCC Operator)

Advanced Systems Group • Los Gatos (CA)

On-site
USD 60,000 - 90,000