On-Site SRE Engineer - Scale Reliable AI Systems

Air Apps

Helsinki

On-site

EUR 75,000 - 110,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Apple hardware ecosystem for work.
Annual Bonus
Top-tier Health and Life Insurance
Transportation Budget
Coverflex benefits package
Childcare support
Air Conference
Pension Fund
Urban Sports Club membership
Meals 100% free at the hub

Job summary

Air Apps is hiring a Site Reliability Engineer (SRE) to ensure the reliability, availability, and scalability of our AI-powered PRP platform. You will work at the intersection of software development and operations, implementing automation, monitoring, and performance optimization strategies to minimize downtime.

This onsite Lisbon-based role offers relocation assistance and a chance to shape critical infrastructure for a fast-growing, globally distributed product.

Qualifications

  • 4+ years experience in Site Reliability Engineering (SRE), DevOps, or System Engineering.
  • Strong knowledge of cloud platforms (AWS, Azure, or GCP).
  • Experience with observability and monitoring tools (Prometheus, Grafana, ELK, Datadog, New Relic).
  • Proficiency in Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Pulumi.
  • Hands-on experience with containerization and orchestration (Docker, Kubernetes, Helm).
  • Strong Linux system administration and networking fundamentals.
  • Experience with incident management, debugging, and root cause analysis.
  • Proficiency in scripting (Bash, Python, or Go) for automation and system monitoring.
  • Knowledge of load balancing, failover strategies, and distributed systems.
  • Understanding of security best practices, access control, and compliance requirements.
  • Strong communication skills and collaboration across cross-functional teams.

Responsibilities

  • Design and implement scalable, reliable, and fault-tolerant systems across cloud environments.
  • Develop and maintain observability tools, including monitoring, logging, and alerting (e.g., Prometheus, Grafana, Datadog, ELK).
  • Automate infrastructure provisioning, deployment, and incident response using IaC tools like Terraform or CloudFormation.
  • Optimize system performance, scalability, and incident response workflows to improve uptime.
  • Work closely with development and DevOps teams to improve system design for reliability.
  • Conduct root cause analysis (RCA) and implement preventative measures to minimize failures.
  • Ensure high availability by designing and maintaining load balancing, failover, and disaster recovery strategies.
  • Improve CI/CD pipelines to enhance deployment speed while maintaining stability.
  • Optimize cloud cost and resource utilization for AWS, Azure, or GCP.
  • Participate in on-call rotations to quickly address system failures and minimize downtime.

Skills

SRE
Cloud platforms
Observability
Terraform
Kubernetes
Linux
Scripting
Incident management
Load balancing
Security

Tools

Prometheus
Grafana
Datadog
ELK
CloudFormation
Pulumi
Docker
Kubernetes
Helm
New Relic

Job description

Air Apps is hiring a Site Reliability Engineer (SRE) to ensure the reliability, availability, and scalability of our AI-powered PRP platform. You will work at the intersection of software development and operations, implementing automation, monitoring, and performance optimization strategies to minimize downtime.

This onsite Lisbon-based role offers relocation assistance and a chance to shape critical infrastructure for a fast-growing, globally distributed product.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Air Apps • Helsinki

On-site
EUR 75,000 - 110,000
Apple hardware ecosystem for work.
Annual Bonus
Top-tier Health and Life Insurance
+7
Mobile AI/ML Engineer - Onsite in Lisbon (Relocation)
Mobile AI/ML Engineer - Onsite in Lisbon (Relocation)

airapps • Helsinki

On-site
EUR 55,000 - 80,000
Apple hardware ecosystem
Annual Bonus
Top-tier Health and Life Insurance
+7
Senior AI Tools Lead for R&D Operations
Senior AI Tools Lead for R&D Operations

RELEX Solutions • Helsinki

Hybrid
EUR 90,000 - 130,000
An international career
Flexible work locations
Annual leave
+2
People Operations Specialist
People Operations Specialist

airapps • Helsinki

On-site
EUR 40,000 - 55,000
Apple hardware ecosystem
Annual Bonus
Top-tier Health and Life Insurance
+7
Data Platform SRE: Cloud, Automation & Observability
Data Platform SRE: Cloud, Automation & Observability

SEB • Pargas

On-site
EUR 72,000 - 99,000
Senior R&D Operations Specialist
Senior R&D Operations Specialist

RELEX Solutions • Helsinki

Hybrid
EUR 90,000 - 130,000
An international career
Flexible work locations
Annual leave
+2
Head of Engineering, Agentic AI Platform
Head of Engineering, Agentic AI Platform

RELEX Solutions • Helsinki

On-site
EUR 150,000 - 190,000
Ownership of strategic bets
Autonomy over technical strategy
AI frontier capabilities
Senior DevOps Engineer
Senior DevOps Engineer

Smartlyio • Helsinki

On-site
EUR 60,000 - 80,000
Generous healthcare packages
Mental health services
Flexible work-life balance
+2
React Native Engineer - AI-Powered Mobile Apps (Lisbon)
React Native Engineer - AI-Powered Mobile Apps (Lisbon)

airapps • Helsinki

On-site
EUR 40,000 - 60,000
Apple hardware ecosystem
Flexible Paid Time Off (PTO)
Annual Bonus
+8
People Ops & Recruitment Specialist - Lisbon (Onsite)
People Ops & Recruitment Specialist - Lisbon (Onsite)

airapps • Helsinki

On-site
EUR 40,000 - 55,000
Apple hardware ecosystem
Annual Bonus
Top-tier Health and Life Insurance
+7