Site Reliability Engineer

Socket.dev

Overland Park (KS)

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) with Profit Sharing
Flexible Time Off
Office Dog!!

Job summary

Ad Astra is seeking a Site Reliability Engineer (SRE) to ensure the performance, reliability, and scalability of our cloud-based systems in an in-office setting in Overland Park, Kansas. You will automate, monitor, and optimize infrastructure, while collaborating with engineering and product teams to achieve high availability and security.

You will design and implement scalable systems, participate in on-call rotations, and drive post-incident reviews to prevent recurrence.

Qualifications

  • 2+ years in Site Reliability Engineering, Systems Engineering, or DevOps with strong automation.

Responsibilities

  • Write automation and production code to improve system reliability and performance.
  • Design, build, and maintain highly available, scalable systems across cloud environments (AWS, Azure, GCP).
  • Own reliability and scalability considerations for a multi-tenant SaaS platform, including tenant isolation.
  • Maintain and extend logging, monitoring, and alerting systems to enhance observability and proactive incident response.
  • Bridge development and operations by automating workflows, deployments, and infrastructure provisioning.
  • Proactively monitor and respond to alerts and incidents, ensuring system uptime and performance.
  • Support security and compliance initiatives.
  • Eliminate manual toil across legacy systems and data integration with automation.
  • Collaborate with engineering, product, and operations to capacity plan and reduce cloud costs.

Skills

Creative Problem Solving
Collaborative Communication
Adaptability & Flexibility
Sense of Urgency with Quality
Attention to Detail
Technical Aptitude

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

GitHub
Jenkins
Terraform
CloudFormation
Docker
Kubernetes
Ansible
AWS

Job description

Competitive Compensation & Benefits Package
  • 401(k) with Profit Sharing
  • Flexible Time Off
  • Office Dog!!
About Us

By combining our unparalleled domain expertise with leading-edge technology, Ad Astra is helping higher education in its mission to advance timely student completions. We are building a cloud-based software platform that will provide the foundation for our next generation of industry-leading solutions and analytics. Simply put, we’re helping students graduate faster.

Our Core Values
  • We recognize talent. We recognize and appreciate the unique God-given talents that our people bring to Ad Astra. Aligning these individual gifts with our work sets team members up to succeed.
  • We’re unpretentious. There’s no room for ego. We admit our imperfections and have the humility to know what we don’t know.
  • We’re passionate. We aren’t satisfied with the status quo. We’re on a mission together to protect the value of degree completion and to transform the higher education industry.
  • We’re pioneering. We’re pioneering and aren’t afraid of failing—in fact, we celebrate it. We love it when our people boldly experiment with innovative solutions.
  • We love fun. The health of our relationships is strengthened by working with people who stretch our thinking—and by enjoying the lighter side of life together. We don’t take ourselves too seriously, but we do take fun seriously.
  • We have grit. Beyond talent and intelligence, our people have stick-to-itiveness. We push through challenges to make goals a reality.
Position Summary

The Site Reliability Engineer (SRE) will ensure the performance, reliability, and scalability of our systems as we continue to grow. This role bridges the gap between software development and operations, applying software engineering principles to automate, optimize, and enhance the reliability of our infrastructure and production systems. Your role includes identifying recurring failure patterns, implementing automated solutions, and continuously improving platform performance. Leveraging your intellectual curiosity and expertise in operations and development, you will also play a pivotal role in monitoring security and reliability threats, while actively advocating effective solutions.

This role spans a genuinely wide range of work, from deep automation and greenfield infrastructure projects to legacy system support and cross-team collaboration. You’ll thrive here if you enjoy variety and can move between priorities without missing a beat, and if you’re motivated by helping shape reliability practices as we grow toward higher availability targets.

Essential Functions / Core Responsibilities
  • Write automation and production code to improve system reliability and performance.
  • Design, build, and maintain highly available, scalable systems across cloud environments (e.g., AWS, Azure, or GCP).
  • Own reliability and scalability considerations unique to our multi-tenant SaaS platform, including tenant isolation and blast radius containment.
  • Maintain and extend logging, monitoring, and alerting systems to enhance observability and proactive incident response.
  • Bridge development and operations by automating workflows, deployments, and infrastructure provisioning.
  • Proactively monitor and respond to alerts and incidents, ensuring system uptime and performance.
  • Support security and compliance initiatives.
  • Eliminate manual toil across legacy systems (currently being migrated away from), client onboarding, and data integration, with the autonomy to build lasting automation.
  • Collaborate with engineering, product, and operations teams to capacity plan, drive cloud cost reduction, and enhance the overall reliability and efficiency of our products
  • Support production systems, including participation in on-call rotations and performing limited after-hours maintenance.
  • Lead and contribute to post-incident reviews, driving root cause analysis and long-term solutions.
  • Document reliability patterns, runbooks, and learnings to build operational maturity
  • Other duties as assigned
Position Requirements
  • Bachelor’s degree in Computer Science, Engineering, or related field preferred; equivalent experience in supporting distributed software systems accepted.
  • 2+ years of experience in Site Reliability Engineering, Systems Engineering, or DevOps, with strong systems/infrastructure fundamentals and comfort using modern tooling, including AI-assisted development, backed by the judgment to design sound automation, not just generate it
  • Strong understanding of networking concepts including load balancing, DNS, IPSec, and VPNs.
  • Experience with source version control, CI/CD, and Infrastructure as Code tools (e.g., GitHub, Jenkins, Terraform, CloudFormation).
  • Working knowledge of Linux operating systems.
  • Proficiency with relational or NoSQL database technologies (both preferred)
  • Proficiency in at least one scripting or programming language (Node.js, Python, Go, Bash, PowerShell, etc.).
  • Experience with containerization and orchestration (Docker, ECS, Kubernetes).
  • Familiarity with observability tools (Graylog, New Relic, Prometheus, Grafana, ELK Stack, etc.).
  • Strong collaboration, problem-solving, and communication skills.
Essential Competencies
  • Creative Problem Solving
  • Collaborative Communication
  • Adaptability & Flexibility
  • Sense of Urgency with Quality
  • Attention to Detail
  • Technical Aptitude
Additional Preferred Qualifications
  • Expertise in git, docker, terraform, ansible and AWS
  • Experience with blue/green or canary deployment strategies and zero-downtime releases.
  • Understanding of security best practices in cloud-native environments.
  • Background in automating large-scale infrastructure management.
  • Experience working in an agile or SaaS-based environment.
How Performance Is Measured For This Role
  • Meaningful contributions to the SRE high value/team stories.
  • Timely response to infrastructure alerts and ensuring system reliability.
  • Regular preventative maintenance.
  • Contribution to the overall success of the Cloud Ops team.
  • Incident Mean time to Acknowledge (MTTA) < 15 minutes.
  • Drive availability improvements to exceed 99.95% uptime.

This full time, in-office position is located in Overland Park, Kansas. Ad Astra does not pay for relocation expenses.

Ad Astra is proud to be an equal opportunity employer. We are committed to fostering an inclusive workplace where all individuals are treated with respect and fairness—regardless of race, color, national origin, sex, gender identity or expression, sexual orientation, religion, age, political affiliation, disability, veteran status, or any other characteristic protected by law.

All applicants must be legally authorized to work in the United States. Please note that Ad Astra is unable to provide work visa sponsorship for this position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Cloud Platform & Flexible Time Off
Site Reliability Engineer — Cloud Platform & Flexible Time Off

Socket.dev • Overland Park (KS)

On-site
USD 110,000 - 150,000
401(k) with Profit Sharing
Flexible Time Off
Office Dog!!
Senior SRE / DevOps Engineer
Senior SRE / DevOps Engineer

Waystar, Inc • Louisville (KY)

On-site
USD 90,000 - 130,000
Customizable benefits package
Paid parental leave
Education assistance
+2
Site Reliability Engineer - 7 Month Contract
Site Reliability Engineer - 7 Month Contract

Orion Health group • Fort Worth (TX), Town of Texas (WI)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

Hybrid
USD 210,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Axle • Frederick (MD)

On-site
USD 140,000 - 155,000
Paid Time Off
401K match
Educational Benefits
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Govcio LLC • United States

Hybrid
USD 210,000 - 230,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Site Reliability Engineer
Site Reliability Engineer

Bolt Graphics, Inc. • Sunnyvale (CA)

On-site
USD 145,000 - 165,000
100% covered medical, dental, and vision premiums
Equity - Stock Options
401(k) match
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Core Scientific, Inc • Austin (TX)

Hybrid
USD 140,000 - 190,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Oaks (PA)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare coverage
401(k) matching
Tuition reimbursement
+1