Get more replies from employers
Send a job-specific resume in minutes.
Cvent, Inc. is seeking a Lead Site Reliability Engineer to shape reliability strategy across cloud architectures and large-scale systems in a fast-growing environment.
You will own incident management, SLI/SLO definitions, and automation efforts, mentoring engineers and guiding CI/CD pipelines, Terraform, and AWS-based deployments.
Collaborate with product and engineering teams to balance cost, performance, and resilience while driving blameless postmortems and security compliance.
Lead Site Reliability Engineer
Cventisaleadingmeetings,events,andhospitalitytechnologyproviderwithmorethan5,500+employeesand30,000+customersworldwide,including60%oftheFortune500.Foundedin1999,Cventdeliversacomprehensiveeventmarketingandmanagementplatformformarketersandeventprofessionalsandofferssoftwaresolutionstohotels,specialeventvenuesanddestinationstohelpthemgrowtheirgroup/MICEandcorporatetravelbusiness.Ourtechnologybringsmillionsofpeopletogetherateventsaroundtheworld.Inshort,we’retransformingthemeetingsandeventsindustrythroughinnovativetechnologythatpowersthehumanconnection.
Set the direction and long-term strategy for solving complex problems; communicate timeline, scope, risks, and the technical roadmap to leadership and stakeholders.
Keep abreast of emerging cloud technologies, running POCs to assess suitability and value.
Lead design and implementation of large-scale AWS architectures optimized for reliability, scalability, and cost.
Develop and refine CI/CD pipelines and automated deployments for complex environments.
Drive error budgets; define and track SLIs/SLOs; own organization-wide reliability targets.
Lead deep-dive RCAs for major incidents using Datadog, Prometheus, Grafana, and ELK; run blameless postmortems.
Establish best practices for containerization (Docker, Kubernetes) and infrastructure as code (Terraform, AWS CDK).
Mentor, coach, and provide technical escalation for SRE/DevOps engineers; foster a learning and ownership culture.
Direct AI/automation initiatives to improve deployment, monitoring, troubleshooting, and developer efficiency (e.g., ChatGPT, Copilot, generative AI for runbooks and incident management).
Partner with engineering, product, and business to shape reliability strategy, incident process, architecture reviews, and roadmap.
Ensure security, regulatory, and operational compliance across cloud architecture and automation.
continuously communicating timeline, scope, risks, and technical road map.