Staff Software Engineer, Reliability - Command|Alert

CommandLink, LLC

Argentina

Presencial

ARS 180.802.000 - 271.203.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo: un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Room to grow
Flexible time off
Employee referral bonuses
Team events

Descripción de la vacante

Command|Link is seeking a Staff Software Engineer for reliability to own the alerting pipeline end to end and to drive architectural decisions across security tooling, monitoring telemetry, and LLM-driven investigation.

The role requires deep SRE/DevOps experience, strong Kafka and ML-based detection capabilities, and a track record delivering highly reliable, scalable systems in multi-cloud environments.

Formación

  • Background operating as a site reliability engineer, DevOps engineer, or similar production-ownership role with fluency in SLOs, SLIs, and error budgets.
  • Experience designing, building, or operating high-reliability alerting or notification systems in production, including ML-based detection at scale.
  • Strong Kafka experience with webhook reliability, idempotency, and delivery guarantees under load.
  • Experience building observability and automation for a high-volume production system with low toil.
  • Working command of telemetry and protocol data and the ability to turn it into real network/topology insight.
  • Mastery of Go and/or Python with container orchestration and multi-cloud footprint; able to make org-level architecture calls.

Responsabilidades

  • Own the reliability of the alerting pipeline end to end, from OpenSearch evaluation to Kafka delivery and downstream notifications.
  • Define SLOs and SLIs and use error budgets to guide reliability investments.
  • Lead incident response for critical production issues with blameless post-mortems and systemic fixes.
  • Build observability and automation to run the pipeline at Fortune 1000 scale with low operational toil.
  • Own capacity planning as ingestion and customer base grow.
  • Set architecture for rule-based thresholds, ML scores, and correlation logic for investigation/remediation.
  • Mentor engineers and represent Command|Alert's direction to stakeholders outside engineering.

Conocimientos

SLOs/SLIs
Blameless incident response
Go/Python
Observability/automation
Multi-cloud

Herramientas

Kafka
OpenTelemetry
OpenSearch
NetFlow/sFlow
SNMP

Descripción del empleo

Staff Software Engineer, Reliability - Command|Alert

About Command|Link

Command|Link is a global SaaS Platform providing network, voice services, and IT security solutions, helping corporations consolidate their core infrastructure into a single vendor and layering on a proprietary single pane of glass platform. Command|Link has revolutionized the IT industry by tackling the problems our competitors create. In recognition for our unprecedented innovation and dedication, Command|Link was recognized as the SD-WAN Product of the Year, ITSM Visionary Spotlight, UCaaS Product of the Year, NaaS Product of the Year, Supplier of the Year, and the AT&T Strategic Growth Partner. Command|Link has built the only IT platform for scale that solves ISP vendor sprawl and IT headaches. We make it easy for our customers to get more done, maximize uptime and improve the bottom line.

Command|Alert is CommandLink's signal-processing core, the engine that turns raw security, monitoring, and customer-defined telemetry into alerts customers actually trust. Alert fatigue and noise are the top complaint across every competitor in this space, and this role exists to make sure our alerts are the ones people don't tune out.

As a Staff Software Engineer on Command|Alert, you'll own the reliability of the alerting pipeline end to end and drive the org's most consequential decisions on how it's architected. This role suits someone who thinks like a site reliability engineer as much as a software engineer: fluent in SLOs, error budgets, and blameless incident response, with the technical range to reason across security tooling, monitoring telemetry, syslog, OpenTelemetry, and L2-L4 network protocols.

Key Responsibilities:

  • Own the reliability of the alerting pipeline end to end, from OpenSearch alert evaluation through Kafka delivery via OpenSearch callbacks to downstream notification, including idempotency guarantees and soak-tested behavior under sustained load.
  • Define SLOs and SLIs for the pipeline's availability, latency, and delivery guarantees, and use error budgets to guide how much investment goes into reliability work versus new capability.
  • Lead incident response for the pipeline's most critical production issues, running blameless post-mortems and driving the systemic fixes and automation that come out of them.
  • Build the observability and automation needed to run the pipeline at Fortune 1000 scale with low operational toil, and own capacity planning as ingestion volume and customer count grow.
  • Set the architecture for how Command|Alert evaluates rule-based thresholds, ML anomaly scores, and correlation logic, turning diverse telemetry into usable network and system topologies that power LLM-driven investigation and remediation.
  • Mentor engineers across the teams you touch and represent Command|Alert's technical direction to stakeholders outside engineering.
  • Takes on additional responsibilities and projects as needed to support the success of the team and organization.

What you'll need for success:

Required
  • Background operating as a Site Reliability Engineer, DevOps engineer, or in a similar production-ownership role, with fluency in SLOs, SLIs, error budgets, and blameless incident response.
  • Demonstrated experience designing, building, or operating high-reliability alerting or notification systems in production, including rule-based and ML-based detection at scale.
  • Strong Kafka experience, and a track record building systems where webhook reliability, idempotency, and delivery guarantees under load are non-negotiable.
  • Experience building the observability and automation that let a high-volume production system run with low operational toil.
  • A working command of telemetry and protocol data (security tooling output, syslog, OpenTelemetry, NetFlow/sFlow, SNMP, ICMP, firewall logs) and the ability to turn it into real network and system topologies.
  • Recognized mastery of Go and/or Python, with range across container orchestration and a multi-cloud footprint, and a demonstrated ability to make org-level architecture calls.
Nice to Have
  • Experience with chaos engineering or fault-injection testing (e.g., Gremlin, Chaos Mesh).
  • Familiarity with SLO and error-budget tooling and practice (e.g., Nobl9, Google's SRE workbook approach).
  • Experience with Temporal or a comparable workflow orchestration platform.
  • Experience operating in a multi-tenant, cloud-native environment with secrets management and TLS at scale.

Why you'll love life at Command|Link:

  • Room to grow at a high-growth company
  • An environment that celebrates ideas and innovation
  • Your work will have a tangible impact
  • Flexible time off
  • Fun events at cool locations
  • Employee referral bonuses to encourage the addition of great new people to the team

At CommandLink, we’re committed to creating a fair, consistent, and efficient hiring experience. As part of our process, we use AI-assisted tools to help review and analyze applications. These tools support our recruiting team by identifying qualifications and experience that align with the requirements of each role.

AI tools are used only to assist in the evaluation process — they do not make final hiring decisions. Every application is reviewed by a member of our recruiting or hiring team before any decisions are made.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Software Engineer, Alerting Platform
Senior Software Engineer, Alerting Platform

CommandLink, LLC • Argentina

Presencial
ARS 3.906.000 - 5.580.000
Career growth
Innovative environment
Impactful work
+2
Staff Software Engineer, Security Plaftorm
Staff Software Engineer, Security Plaftorm

CommandLink, LLC • Argentina

Presencial
ARS 180.802.000 - 241.069.000
Room to grow
Flexible time off
Team events
+1
Staff Software Engineer
Staff Software Engineer

CommandLink, LLC • Argentina

Presencial
ARS 20.088.000 - 35.712.000
Flexible time off
Employee referral bonuses
Team events and growth opportunities
Senior Software Engineer
Senior Software Engineer

CommandLink, LLC • Argentina

Presencial
ARS 1.800.000 - 2.400.000
Room to grow at a high-growth company
Flexible time off
Fun events at cool locations
+1
Staff Software Engineer, Platform Engineering
Staff Software Engineer, Platform Engineering

CommandLink, LLC • Argentina

Presencial
ARS 3.000.000 - 6.000.000
Flexible time off
Employee referral bonuses
Senior Software Engineer, Alerting Platform — Scale-Ready
Senior Software Engineer, Alerting Platform — Scale-Ready

CommandLink, LLC • Argentina

Presencial
ARS 3.906.000 - 5.580.000
Career growth
Innovative environment
Impactful work
+2
Staff Reliability Engineer — End-to-End Alerting at Scale
Staff Reliability Engineer — End-to-End Alerting at Scale

CommandLink, LLC • Argentina

Presencial
ARS 180.802.000 - 271.203.000
Room to grow
Flexible time off
Employee referral bonuses
+1
Staff Go Engineer - Full-Stack Security Platform Architect
Staff Go Engineer - Full-Stack Security Platform Architect

CommandLink, LLC • Argentina

Presencial
ARS 20.088.000 - 35.712.000
Flexible time off
Employee referral bonuses
Team events and growth opportunities
Devops / Sre / Devsecops Engineer (Aws) - Remote, Latin America
Devops / Sre / Devsecops Engineer (Aws) - Remote, Latin America

Bluelight Consulting, Llc • Municipio de Profesor Salvador Mazza

A distancia
ARS 1.200.000 - 1.800.000
Competitive salary and bonuses
Generous paid-time-off policy
Work remotely
+1
Devops / Sre / Devsecops Engineer (Aws) - Remote, Latin America
Devops / Sre / Devsecops Engineer (Aws) - Remote, Latin America

Bluelight Consulting, Llc • Rosario

Híbrido
ARS 1.800.000 - 3.200.000
Competitive salary and bonuses
Generous paid-time-off
Remote work option
+1