Get more replies from employers
Send a job-specific resume in minutes.
Cvent, Inc. is seeking a Lead Site Reliability Engineer to shape reliability strategy and drive large-scale AWS architectures across regional deployments. You will own CI/CD pipelines, incident response, and automation efforts while mentoring the team to deliver resilient, scalable infrastructure.
The role emphasizes leadership, collaboration with product and engineering, and ensuring security and regulatory compliance in cloud environments.
Lead Site Reliability Engineer
Cvent is a leading meetings, events, and hospitality technology provider with more than 5,500+ employees and 30,000+ customers worldwide, including 60% of the Fortune 500. Founded in 1999, Cvent delivers a comprehensive event marketing and management platform for marketers and event professionals and offers software solutions to hotels, special event venues and destinations to help them grow their group/MICE and corporate travel business. Our technology brings millions of people together at events around the world. In short, we're transforming the meetings and events industry through innovative technology that powers the human connection.
Cvent's strength lies in its people, fostering a culture where everyone is encouraged to think like entrepreneurs, taking risks and making decisions confidently. We value diverse perspectives and celebrate differences, working together with colleagues and clients to build strong connections.
Set the direction and long-term strategy for solving complex problems; communicate timeline, scope, risks, and the technical roadmap to leadership and stakeholders.
Keep abreast of emerging cloud technologies, running POCs to assess suitability and value.
Lead design and implementation of large-scale AWS architectures optimized for reliability, scalability, and cost.
Develop and refine CI/CD pipelines and automated deployments for complex environments.
Drive error budgets; define and track SLIs/SLOs; own organization-wide reliability targets.
Lead deep-dive RCAs for major incidents using Datadog, Prometheus, Grafana, and ELK; run blameless post-mortems.
Establish best practices for containerization (Docker, Kubernetes) and infrastructure as code (Terraform, AWS CDK).
Mentor, coach, and provide technical escalation for SRE/DevOps engineers; foster a learning and ownership culture.
Direct AI/automation initiatives to improve deployment, monitoring, troubleshooting, and developer efficiency (e.g., ChatGPT, Copilot, generative AI for runbooks and incident management).
Partner with engineering, product, and business to shape reliability strategy, incident process, architecture reviews, and roadmap.
Ensure security, regulatory, and operational compliance across cloud architecture and automation.
Continuously communicating timeline, scope, risks, and technical road map.