Software Engineer (Infrastructure)

Rad AI

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

100% health, dental, & vision
Flexible PTO
WFH stipend
401k plan
Stock options
Socials & off-sites

Job summary

Rad AI is expanding its Platform Engineering organization, focusing on a scalable, reliable infrastructure powering Reporting, Impressions, and Continuity. You will shape architecture and reliability practices for cloud-native systems on AWS, implementing robust CI/CD, monitoring, and security patterns across multiple product teams.

You will collaborate with Platform leadership, product engineering, data, and ML teams to design compliant, observable systems, mentor engineers, and help drive

Qualifications

  • 4+ years of hands-on infrastructure/platform development in cloud-native environments.
  • Experience owning critical production systems and improving reliability.
  • Experience with HIPAA/regulatory considerations is a plus.

Responsibilities

  • Design and evolve cloud infrastructure across AWS, containers, serverless, and data stores.
  • Own observability, monitoring, and incident response for core platform services.
  • Mentor engineers and drive reliability, security, and maintainability improvements.

Skills

Linux fundamentals
Networking
CI/CD practices
Observability
Security best practices
Mentoring
Communication
Problem solving

Tools

AWS
Kubernetes
Terraform
Docker
EC2
OpenTelemetry

Job description

  • The Platform Engineering organization at Rad AI builds the foundations that power all of our products—Reporting, Impressions, and Continuity—and enables product teams to ship reliably, safely, and at scale
  • Within Platform, the Infrastructure team owns our core cloud infrastructure, platforms, and reliability practices. We’re hiring a multiple Infrastructure Engineers to help us design and operate robust, scalable systems
  • In this role, you’ll contribute to infrastructure architecture, reliability practices, and thoughtful improvements to our workflows
  • Influence the technical direction for infrastructure and platform capabilities that support our rapidly growing AI product suite
  • Architect and evolve our cloud infrastructure (primarily on AWS) across container orchestration (Kubernetes, Elastic Container Service), serverless (e.g., Lambda), virtual machines (e.g., EC2), and data stores to support current and future products
  • Work closely with Platform leadership, product engineering, data, and ML teams to design systems that are robust, observable, and compliant in a healthcare environment
  • Define and drive infrastructure strategy for the Platform org—partnering with engineering leadership to align roadmaps, set standards, and sequence work for maximum business impact
  • Secure networking, identity, and access patterns across environments
  • Improve reliability and operational excellence by defining SLOs, SLIs, and error budgets for core platform services
  • Leading and participating in blameless post-incident reviews and translating learnings into systemic improvements
  • Own observability and monitoring strategy across logging, metrics, and tracing, ensuring we can detect, debug, and prevent issues efficiently
  • Mentor and level up engineers across Platform and product teams—reviewing design docs, guiding architecture decisions, and modeling high standards for reliability, security, and maintainability
  • Partner with security and compliance stakeholders to ensure our infrastructure and operational practices meet HIPAA and other healthcare requirements
  • Advocate for and implement developer experience improvements, such as better CI/CD workflows, faster feedback loops, and tooling that reduces cognitive load for product teams
  • This posting includes multiple open headcount, spanning from Senior to Principal level
Benefits
  • 100% health, dental, & vision
  • Flexible PTO
  • WFH stipend
  • 401k plan
  • Stock options
  • Socials & off-sites

Communicate clearly and empathetically with both technical and non-technical partners, and enjoy mentoring engineers at multiple levelsHave demonstrable experience leading complex, cross-team initiatives from design through rollout—communicating tradeoffs, aligning stakeholders, de-risking launches, and measuring impactExtensive experience building tooling and automation for other engineersBring 4+ years of hands-on infrastructure / platform development experience (or equivalent practical experience) in modern, cloud-native environments, with a track record of owning critical systems in productionHave deep expertise with AWS (preferred) and/or GCP, including core networking, compute, storage, and managed servicesIf you’re passionate about building resilient platforms and enjoy collaborating across functions, we’d love to hear from youPossess solid Linux fundamentals and are comfortable debugging issues at the OS, networking, and application layersAre highly proficient in at least one programming/scripting language used for infrastructure work (Python preferred)Are comfortable with Infrastructure as Code (Terraform preferred, Pulumi, or similar) and Git-based workflowsHave strong experience with Kubernetes, containers (Docker), and container orchestration, and understand how to operate these systems reliably at scaleTake a data-informed, pragmatic approach to decision-making—balancing ideal architecture with business needs, delivery timelines, and team capacityFamiliarity with observability stacks (CloudWatch, New Relic, Grafana, OpenTelemetry, etc.)If you’re passionate about driving innovation and delivering impactful healthcare solutions, we’d love to hear from you!Prior experience at a fast-growing startup where you’ve helped scale infrastructure, processes, and teamsBackground in platform or security engineering, especially around access control, encryption, auditability, and complianceExperience designing or operating internal developer platforms, SDKs, or reusable frameworks that standardize how services are built and deployedExperience in regulated environments (e.g., HIPAA) or prior work in healthcare or health techExperience working closely with ML / data teams or with ML platforms (e.g., Airflow, Ray, ML pipelines, model serving stacks)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Infrastructure
Staff Software Engineer, Infrastructure

Transformcap • United States

On-site
USD 120,000 - 150,000
Comprehensive Medical, Dental, Vision & Life insurance
401(k)
Flexible PTO policy
+1
Software Engineer, Infrastructure (All Levels)
Software Engineer, Infrastructure (All Levels)

Rad AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Comprehensive Medical, Dental, Vision & Life insurance
401(k)
Flexible PTO policy
+2
DevOps Engineer
DevOps Engineer

Jobless • Norwalk (CT), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

Transformcap • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Flexible hours
+1
Senior DevOps Engineer
Senior DevOps Engineer

Qualified Health PBC • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical/Dental/Vision insurance
Hybrid work options
+1
Senior Platform Engineer
Senior Platform Engineer

Linuxconfig • Northern (KY)

Hybrid
USD 140,000 - 190,000
CloudOps / SRE Engineer
CloudOps / SRE Engineer

Capital Factory • Austin (TX)

On-site
USD 140,000 - 190,000
100% employer-paid health, vision, and
dental insurance
401k & retirement plans
Senior DevOps Engineer
Senior DevOps Engineer

Hi Marley • Boston (MA)

On-site
USD 140,000 - 190,000
Open vacation policy
Matching 401k program
Medical, dental, vision, disability, &
+1
Staff Software Engineer, Platform
Staff Software Engineer, Platform

Socket.dev • San Francisco (CA)

On-site
USD 130,000 - 180,000
Platform Engineer
Platform Engineer

Synergy • Chicago (IL)

On-site
USD 100,000 - 150,000