Site Reliability Engineer

Fulcrum Digital Inc

Portugal

Presencial

EUR 50 000 - 80 000

Tempo integral

Há 5 dias
Torna-te num dos primeiros candidatos

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Fulcrum Digital Inc is seeking an experienced DevOps/SRE engineer to plan, manage, and optimize production environments across multiple geographies. You will define monitoring, automation, and deployment strategies, and work with global teams to improve reliability and performance.

You will engage in the full service lifecycle, from design through operation, and play a key role in capacity planning, incident response, and continual improvement of platform resiliency.

Qualificações

  • Experience planning and managing production environments.
  • Ability to define monitoring and alerting strategies for production systems.
  • Strong incident response and post-incident analysis skills.
  • Experience supporting CI/CD pipelines and automating deployment processes.
  • Understanding of DevOps practices and platform reliability.

Responsabilidades

  • Plan, manage, and oversee all aspects of a Production Environment.
  • Define strategies for Application Performance Monitoring and Optimisation in Prod.
  • Respond to Incidents and improvise a platform based on feedback and measure reduction of incidents.
  • Support the deployment of code into multiple lower environments with emphasis on automation.
  • Design, develop, and standardize monitoring and alerting for supported applications.
  • Take a holistic approach to problem solving during production events across the tech stack.
  • Engage in the lifecycle of services from inception to refinement and optimization.
  • Provide feedback on operational gaps or resiliency concerns from ITSM activities.
  • Support pre-go-live activities like capacity planning and launch reviews.
  • Lead DevOps automation and best practices in promoting software to higher environments.
  • Maintain service health by monitoring availability and latency and overall system health.
  • Scale systems via automation and changes that improve reliability and velocity.

Conhecimentos

Production environment management
APM / performance optimization
Incident response
CI/CD automation
DevOps practices

Ferramentas

Shell scripting
ITIL/ITSM
SQL
Application troubleshooting
Splunk
Dynatrace

Descrição da oferta de emprego

Requirements
Must Have
  • Plan, manage, and oversee all aspects of a Production Environment
  • Define strategies for Application Performance Monitoring, Optimization in Prod environment
  • Respond to Incidents and improvise a platform based on feedback and measure the reduction of incidents over time.
  • Support the deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.
  • Design, develop, and standardize the monitoring and Alerting mechanism for the supported applications.
  • Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize the mean time to recover.
  • Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
  • Analyze ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
  • Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
  • Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
  • Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
  • Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.
  • Work with a global team spread across tech hubs in multiple geographies and time zones.
  • Ability to share knowledge and explain processes and procedures to others.
  • Share knowledge and mentor junior resources
  • Able to perform on-call duties on a rotational basis.
  • Occasional off-hours work required.
  • Candidate should have an inclination for Training and should be a good trainer and ready to mentor others
  • Shell Scripting
  • ITIL / ITSM
  • SQL - Basic / Good to have
  • Application Troubleshooting
  • Any Monitoring tool (Preferred Splunk/Dynatrace)
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

System Reliability Engineer (Application Support + Automation)
System Reliability Engineer (Application Support + Automation)

Fulcrum Digital Inc • Lisboa

Presencial
EUR 42 000 - 65 000
System Reliability Engineer (Application Support + Automation)
System Reliability Engineer (Application Support + Automation)

Fulcrum Digital • Lisboa

Presencial
EUR 55 000 - 75 000
Site Reliability Engineer
Site Reliability Engineer

EPAM Systems • Portugal

Presencial
EUR 40 000 - 60 000
Application Production Support Engineer
Application Production Support Engineer

Inetum • Lisboa

Presencial
EUR 52 000 - 78 000
Site Reliability Engineer
Site Reliability Engineer

Fulcrum Digital Inc • Lisboa

Presencial
EUR 50 000 - 70 000
IT Operations Engineer (DevOps & Production Services)
IT Operations Engineer (DevOps & Production Services)

Inetum • Lisboa

Presencial
EUR 42 000 - 64 000
Application Support (Senior)
Application Support (Senior)

Caixa Mágica Software • Lisboa

Presencial
EUR 32 000 - 46 000
Health and Life Insurance
Social events and team buildings
Training in the latest technologies
+1
Application Support
Application Support

HN Services Portugal • Lisboa

Presencial
EUR 32 000 - 52 000
Application Support
Application Support

Caixa Mágica Software • Lisboa

Presencial
EUR 32 000 - 52 000
Health and Life Insurance
Social events and team buildings
Training in latest technologies
+1
Cloud Support Infrastructure Expert
Cloud Support Infrastructure Expert

act digital • Lisboa

Presencial
EUR 42 000 - 66 000