Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.
Elastic, the Search AI Company, is seeking a Senior Site Reliability Engineer (Capacity) to optimize compute resources for Elastic Cloud Hosted and Serverless workloads. You will collaborate with control plane and platform teams across EMEA and NASA to ensure scalable capacity.
You will develop capacity models, implement autoscaling, and monitor performance to prevent bottlenecks. This role emphasizes cross-team collaboration and solving real-world cloud scaling challenges.
Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the precision of search and the intelligence of AI to enable everyone to accelerate the results that matter. By taking advantage of all structured and unstructured data — securing and protecting private information more effectively — Elastic’s complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI.
What is The Role
As a Principal Platform Engineer focused on capacity, you will play a crucial role in managing and optimizing our compute resources, ensuring that our Elastic Cloud Hosted and Serverless workloads can scale seamlessly. You’ll collaborate closely with our control plane and cross-functional platform engineering teams, addressing the real-world challenges of cloud scaling and resource allocation. Our diverse team spans EMEA and NASA, and together we tackle the complexities of building scalable systems, making a meaningful impact on how we serve our customers.
What You Will Be Doing
Assess current and future capacity requirements based on workload demands to ensure seamless scaling of resources. Develop and maintain accurate capacity models that predict resource needs and align with business objectives. Collaborate with teams to implement proactive measures that prevent capacity shortages and bottlenecks.
Implement effective strategies for optimizing resource usage across our cloud environments. Ensure that compute resources are utilized efficiently to enhance performance and support seamless scalability.
Analyze capacity metrics and trends to guide effective resource allocation decisions. Develop insightful reporting tools that provide clear visibility into capacity and performance, helping to optimize our compute resources.
Operate an autoscaling framework that accommodates various customer workloads seamlessly. Optimize infrastructure performance across over 60 regions in Elastic Cloud. Collaborate with development teams to implement scaling best practices effectively .
What You Bring
5+ years with cloud infrastructure and capacity management
Knowledge of performance monitoring and optimization techniques
Understanding of cloud scaling challenges and solutions
Proficiency with incident investigation and troubleshooting processes
experience with compute auto-scaling processes and capacity reservations across the three major CSPs
solid software and platform engineering background
worked with the three major cloud service providers and navigated compute capacity scaling issues
Additional Information - We Take Care of Our People As a distributed company, diversity drives our identity. Whether you’re looking to launch a new career or grow an existing one, Elastic is the type of company where you can balance great work with great life. Your age is only a number. It doesn’t m...