Site Reliability Engineer

Nextpoint

Chicago (IL)

On-site

USD 120,000 - 170,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nextpoint, a Chicago-based legal-tech company, is seeking a Site Reliability Engineer to ensure reliability, security, and scalability of our AWS-based platform. You’ll be the first line of defense for production infrastructure and help expand AI capabilities on Bedrock.

You will own automation, monitoring, and incident response, while collaborating with engineering to bake reliability into new features. A strong on-call mindset and clear runbooks are essential.

Qualifications

  • 5+ years in SRE/DevOps or Infra roles, preferably in a B2B SaaS environment.
  • Hands-on AWS experience (EC2, S3, RDS, Lambda, VPC, IAM, CloudWatch).
  • Terraform, CloudFormation, or CDK; configuration management.
  • Scripting in Python or Go for automation.
  • Containerization with Docker/Kubernetes; CI/CD pipelines.
  • Monitoring stacks with Datadog, CloudWatch, Prometheus/Grafana.
  • Security/compliance familiarity (SOC 2, HIPAA) and encryption practices.
  • Experience with AI/ML infra (Bedrock) is a plus.
  • Strong communication and on-call readiness.

Responsibilities

  • Infrastructure & Automation: maintain and extend IaC and CI/CD pipelines.
  • Reliability & Operations: monitor uptime, participate in on-call, triage incidents, write post-incident reports, improve monitoring.
  • Cross-Functional Collaboration: work with engineers, document runbooks and decisions, communicate status.
  • Security & Compliance: support SOC 2, encryption, access controls, audit logging, security reviews.

Skills

AWS expertise
Terraform / IaC
Python / Go
Docker & Kubernetes
CI/CD pipelines
Monitoring & observability
Security & compliance

Education

Bachelor's degree in CS/Engineering

Tools

Datadog
CloudWatch
Prometheus
Grafana
Amazon Bedrock

Job description

ABOUT NEXTPOINT

Nextpoint builds transformative software and services for the legal industry — making eDiscovery, case management, and litigation prep simple, fluid, and affordable for law firms of all sizes. Our secure, cloud-based platform lets teams start document review in minutes, backed by powerful analytics, an intuitive interface, and best-in-class security at every point.

We're problem solvers, simplifiers, and challenge seekers, united by a shared goal: a great team culture and satisfied clients. We're headquartered in Chicago's Ravenswood neighborhood and proud to have been named one of Built In's Best Startups to Work for in Chicago five years running (2022–2026).

ABOUT THE ROLE

Nextpoint's platform handles massive, unpredictable volumes of sensitive legal data — unlimited-upload document review, AI-assisted analysis, and secure electronic production — all running on AWS with zero downtime tolerance for firms in active litigation. We're looking for a Site Reliability Engineer to help maintain the reliability, scalability, and security posture of that platform as we expand our AI capabilities (built on Amazon Bedrock) and grow our customer base.

This is a hands-on role for someone who wants to be the first line of response for production infrastructure at a company where reliability is a customer-trust issue, not just an engineering metric.

RESPONSIBILITIES

Infrastructure & Automation (50%)

  • Maintain and extend existing infrastructure-as-code (Terraform/CloudFormation/CDK) following established patterns and standards
  • Support and operate CI/CD pipelines; implement improvements as directed
  • Monitor cloud cost trends and flag optimization opportunities for review

Reliability & Operations (25%)

  • Monitor uptime, latency, and performance SLOs/SLIs for production systems supporting document upload, processing, review, and production workflows
  • Participate in the on-call rotation and serve as first responder for production incidents during US business hours
  • Triage, troubleshoot, and resolve incoming infrastructure requests and incidents; **escalate** and coordinate on complex root-cause work
  • Write clear post-incident reports and help investigate recurring incidents and cost overruns
  • Maintain and extend existing monitoring, alerting, and observability tooling

Cross-Functional Collaboration (15%)

  • Partner with engineering teams to build reliability, scalability, and observability into new features from design through launch
  • Document runbooks, architecture decisions, and operational procedures for the broader engineering team
  • Communicate incident status and technical issues clearly to engineering and non-technical stakeholders
  • Participate in design reviews to flag reliability or operational concerns early

Security & Compliance (10%)

  • Support SOC 2 compliance activities and help uphold encryption, access-control, and audit-trail standards across all environments
  • Implement and maintain security best practices for infrastructure handling confidential legal and client data, including AI workloads on Amazon Bedrock
  • Support security reviews, vulnerability management, and patching cadences across production systems
  • Maintain permissions-based access controls and comprehensive audit logging in line with client security commitments
QUALIFICATIONS
  • 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles, ideally in a B2B SaaS environment
  • Deep hands-on experience with AWS (EC2, S3, RDS, Lambda, VPC, IAM, CloudWatch, or equivalent services)
  • Working knowledge of infrastructure-as-code (Terraform, CloudFormation, or CDK) and configuration management
  • Proficiency in at least one scripting/programming language (Python, Go, or similar) for automation and tooling
  • Experience with containerization and orchestration (Docker, Kubernetes, or ECS)
  • Track record of operating CI/CD pipelines
  • Experience with monitoring/observability stacks (Datadog, CloudWatch, Prometheus/Grafana, or similar)
  • Familiarity with security compliance frameworks (SOC 2, HIPAA, or similar) and encryption/access-control best practices
  • Experience supporting systems handling large-scale, variable-volume data processing is a plus
  • Exposure to AI/ML infrastructure (e.g., Amazon Bedrock, model-serving pipelines) is a plus, given our growing AI feature set
  • Experience using ClaudeCode, Kiro, OpenCode or similar Agentic AI
  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)
  • Strong written and verbal communication skills, with the ability to write clear runbooks and explain technical tradeoffs to non-technical stakeholders
  • Comfortable being part of an on-call rotation

EQUAL OPPORTUNITY EMPLOYER

Nextpoint is an equal opportunity employer. We actively work to build a diverse team and encourage candidates of all backgrounds to apply. All applicants are considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, disability, or any other characteristic protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Nextpoint, Inc. • Chicago (IL), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer – AI-Ready Cloud & Security
Site Reliability Engineer – AI-Ready Cloud & Security

Nextpoint • Chicago (IL)

On-site
USD 120,000 - 170,000
Product Support Specialist
Product Support Specialist

Nextpoint • Chicago (IL)

On-site
USD 60,000 - 65,000
Flexible hybrid schedule
Health insurance
PTO & holidays
+3
Site Reliability Engineer - AI-Driven Cloud Platform
Site Reliability Engineer - AI-Driven Cloud Platform

Nextpoint, Inc. • Chicago (IL), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Seek Now • Atlanta (GA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Health, dental, and vision coverage
401(k) with company match
+1
Site Reliability Engineer
Site Reliability Engineer

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Inclusion Services S.A • Chicago (IL)

On-site
USD 90,000 - 130,000
100% company-covered health insurance
401k plan with 4% match
15 days paid time off
+3
Site Reliability Engineer Austin, TX
Site Reliability Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Site Reliability Engineer
Site Reliability Engineer

Skill • Southlake (TX)

On-site
USD 66,000 - 73,000
Health insurance
Vision insurance
Dental insurance
+2
Product Reliability Engineer
Product Reliability Engineer

PointOne • New York (NY)

On-site
USD 100,000 - 130,000
Comprehensive health, dental, and vision insurance
Meals in office
Regular team events