Get more replies from employers
Send a job-specific resume in minutes.
Ciroos, Inc. is an early-stage AI startup focused on Site Reliability Engineering (SRE) tooling. We seek a QA / Automation Engineer with 4-8 years of experience to define system-level validation and build scalable automation for distributed systems in production-like environments.
You will work on Kubernetes-based validation, AI-assisted incident tooling, and CI/CD reliability gates, partnering with SRE and platform teams to ensure provable reliability and quality across the stack.
Ciroos is an early-stage AI company founded by operators who’ve built businesses from $0 to $5B+. We’re tackling one of the hardest unsolved problems in enterprise technology — and we’re doing it fast. If you want your work to matter, ship quickly, and build alongside people who’ve done this before, you’re in the right place.
Mission That Matters
You're not building features. You're preventing the kind of failures that make headlines. Every engineer at Ciroos works on problems that have a direct line to customer outcomes.
Move Fast. See the Impact.
Early-stage means your work ships quickly. You'll see it in production, hear about it from customers, and iterate — not wait three years for your roadmap to land.
Build With the Best
Our founding team holds 84 patents and has scaled businesses to $5B+. You'll work alongside people who've built and operated systems at scale — and who have strong opinions about how to do it right.
Grow Fast Here
We invest in people who are curious, coachable, and hungry. Our Growth Mindset value isn't a poster on the wall — it's how we operate every day. If you're learning, you're contributing.
Ciroos is built on the belief that the best teams move fast, stay honest, and never stop learning. We’re customer‑obsessed in the truest sense — everyone here owns customer problems, not just the people with “customer” in their title. We treat time as a forcing constraint: if something matters, we move on it. We debate hard, disagree openly, and commit fully once a decision is made — because that’s how high‑trust teams actually operate. We hire people who ask “what did I learn today?” not “did I avoid making a mistake?” Growth mindset isn’t a value we aspire to — it’s the baseline we hire for. We’re early‑stage, which means the work is harder and the stakes are higher. If that sounds like a place you’d thrive, it probably is.
What drew me to Ciroos was the ‘Sun’ philosophy—the idea of bringing clarity to complex systems. Coming from an SRE background, I wanted to work somewhere where the culture was as innovative as the code. Between our commitment to the community and our belief that AI should empower rather than replace, we are redefining the SRE experience to be sustainable, collaborative, and human‑first.
Comprehensive medical, vision, and dental benefits. 401k plans and commuter benefits (US employees). Free lunches, snacks, and top‑of‑the‑line espressos!
Equity
Equity that could change your life, not just look nice on paper. QSBS eligibility for US employees.
Teamwork
Collaborative coworkers with high IQ and high EQ. No politics. No bureaucracy. No permission‑seeking.
Stewardship
A career‑defining, high‑impact role with plenty of mentorship opportunities from founders and other coworkers.
Full time
On‑site
About Us
We are an early‑stage AI startup focused on Site Reliability Engineering (SRE). Rather than being another observability platform, its goal is to act as an AI SRE teammate that works alongside SRE, DevOps, Platform Engineering, Cloud Operations, and IT Operations teams to investigate incidents, determine root causes, and automate remediation.
Our team includes experienced entrepreneurs and engineers who have built multiple billion‑dollar products from scratch. As a well‑funded US‑based company backed by top‑tier VCs, we have offices in the US, India, and Europe. Join us in our fast‑paced environment where you’ll have a front‑row seat to shape the future of AI‑driven Observability solutions.
Role Overview
We are looking for a QA / Automation Engineer having 4-8 years of industry experience who operates at the intersection of reliability engineering and system‑level QA. You will define how reliability is validated, not just monitored. This includes building automation systems that continuously test, break, and verify complex distributed systems in production‑like environments.
What You’ll Work On
Kubernetes‑based distributed systems at scale
Observability and alerting pipelines
AI‑assisted incident investigation systems
Reliability validation, chaos testing, and failure injection systems
Infrastructure and deployment automation pipelines
Key Responsibilities
Define and own the strategy for system‑level validation
Design and build scalable automation frameworks for:
API, integration, and end‑to‑end testing
Kubernetes and system‑level validation
Regression and reliability pipelines
Build systems that proactively detect failures before they reach production
Drive chaos engineering and failure injection practices
Establish CI/CD reliability gates with strong validation coverage
Partner with SRE, platform, and backend teams to ensure systems are both observable and testable
Lead incident analysis with a focus on improving validation and preventing recurrence
Mentor engineers and raise the bar for system reliability and quality
Technical Expectations
Kubernetes internals, debugging, and multi‑cluster systems
Distributed systems behavior and failure modes
Observability stacks and alerting frameworks
Production incident handling and root cause analysis
Strong Hands‑on Experience With
EKS, GKE, or managed Kubernetes platforms
Networking concepts: VPC, load balancers, service communication, IAM
Chaos testing and reliability engineering practices
Designing large‑scale automation and validation systems
Programming
Strong coding skills in Python and Go (mandatory)
Experience building automation frameworks and system‑level tooling
Proficiency in Shell scripting and infrastructure automation
What Makes This Role Different
You are responsible for ensuring systems are provably reliable, not just operational
Deep QA and validation engineering
Focus on testing distributed systems, not just application features
Work on failure scenarios, not just happy paths
What Success Looks Like
A robust validation layer that continuously tests system reliability
Significant reduction in production incidents and faster recovery times
Strong alignment between observability signals and real system behavior
Clear ownership of both reliability and quality across the platform
Scope
Own platform‑wide reliability and validation architecture
Drive cross‑team initiatives across SRE, platform, and engineering
Act as a technical leader in reliability, automation, and system validation