About the Team
Spectro Cloud, a recognized leader in enterprise Kubernetes management with its flagship Palette platform, is seeking a hands-on Senior Software Engineer to own the resolution of complex, customer-impacting issues in production. This role is the engineering backbone behind customer success, translating field escalations into durable code fixes, backports, and long-term product improvements across commercial and federal deployments.
As a Senior Software Engineer, you will operate at the intersection of Support, Product Engineering, and Customer Reliability. You will debug deep into Kubernetes internals, Go microservices, and distributed systems; deliver patches and hotfixes on supported release branches; and drive systemic fixes that prevent recurrence. You will be a technical authority for high‑severity escalations and a trusted partner to customers, TAMs, and field engineers.
About the Role
Your core mandate is to keep customer environments stable, supported, and current — while continuously improving the maintainability of the shipping product.
- Root‑cause analysis and code‑level fixes for customer‑reported defects across the Palette platform, Kubernetes control plane and VM deployments.
- Go‑based microservice debugging, patching, and backporting across multiple supported release branches.
- Kubernetes‑native troubleshooting spanning CAPI, controllers, operators, CNI, CSI, and workload runtime behavior.
- Reproduction environments and test harnesses that convert customer scenarios into repeatable engineering artifacts.
- Hotfix delivery, patch releases, and CVE remediation on supported versions with strict quality and regression discipline.
- Feedback loops into Product and Engineering that convert recurring escalations into permanent product improvements.
Success in this role requires a builder‑and‑debugger mindset. You should be equally comfortable reading a stack trace, reading Go source, and engaging with a customer. You will use AI responsibly to accelerate triage, log analysis, and reproduction — while maintaining strong engineering accountability for every fix that ships.
Responsibilities
1. Escalation Ownership & Root Cause Analysis
- Own high‑severity customer escalations end to end — from initial reproduction through code‑level root cause, fix, verification, and customer confirmation.
- Debug complex Kubernetes and distributed‑systems issues across control plane, cluster lifecycle, networking, storage, and workload runtime.
- Produce clear, technically rigorous RCA documents that hold up to customer, field, and executive scrutiny.
- Partner with Support, TAMs, and Customer Success to keep customers informed and unblock production impact quickly.
2. Code Fixes, Backports & Patch Delivery
- Deliver production‑quality fixes in Go across the Palette codebase and Kubernetes‑native components.
- Backport fixes cleanly across multiple supported release branches with strong regression discipline.
- Drive hotfix, patch, and CVE releases through the Software release train, coordinating with QA, Release Engineering, and Product.
- Maintain high code‑review standards on code changes — small, safe, well‑tested, and well‑documented.
3. Reproduction, Test Coverage & Regression Prevention
- Build reproduction environments and minimal test cases that convert one‑off customer scenarios into permanent engineering assets.
- Expand unit, integration, and end‑to‑end test coverage to prevent regression of every fix that ships.
- Partner with QA to harden test suites against the failure modes seen in the field.
- Identify systemic gaps in observability, error handling, or upgrade paths and drive them to closure.
4. Product Feedback & Long‑Term Hardening
- Identify recurring escalation patterns and drive engineering changes that eliminate their root causes.
- Partner with Product Management and feature teams to feed development insights into roadmap and design reviews.
- Improve upgrade, rollback, and day‑2 operations based on real‑world customer signals.
- Contribute to supportability improvements — logs, diagnostics, must‑gather tooling, and self‑service remediation.
5. AI‑Accelerated Software Development (Responsible Innovation)
- Apply generative AI tools responsibly to accelerate log triage, stack trace analysis, reproduction scaffolding, and RCA drafting.
- Use effective prompt‑engineering practices to improve the consistency and quality of AI‑assisted debugging workflows.
- Validate all AI‑generated artifacts before use. Increased velocity must never compromise engineering correctness, security, or compliance.
Clarity, Precision, Documentation‑Driven Engineering (DDE)
- Create and maintain clear, Markdown‑based RCAs, fix write‑ups, and knowledge‑base entries so root cause and remediation intent are documented before implementation.
- Own escalations end to end — from customer symptom through code fix, backport, verification, and follow‑through on preventative work.
- Drive continuous improvement of the software development function through measurable reductions in escalation age, backport lead time, and repeat‑defect rate.
- Use sustaining KPIs such as time‑to‑RCA, time‑to‑fix, backport coverage, escaped‑defect rate, and customer‑confirmed resolution to guide improvements.
- Collaborate effectively across Software Development, Support, Product Engineering, QA, Release Engineering, and Security teams.
Minimum Qualifications
- 3+years of hands‑on software engineering experience Engineering, Escalation Engineering, or a comparable production‑focused role.
- Hands‑on experience with Kubernetes - operating, debugging, and modifying Kubernetes and Kubernetes‑native components (controllers, operators, CAPI, CRDs) in complex production environments.
- Strong development experience - you must be comfortable reading, writing, and shipping production code, not only debugging it.
- Proven experience owning high‑severity customer escalations end to end, including code‑level root cause and fix delivery.
- Experience backporting fixes across multiple supported release branches with strong regression discipline.
- Experience with debugging skills across networking, storage, and control‑plane behavior.
- Working knowledge of at least one major cloud platform: AWS, Azure, or GCP.
- Excellent written and verbal communication — able to hold technical authority in front of customers and executives.
Preferred Qualifications
- Go (Golang) development experience — building, debugging, and patching Go microservices and Kubernetes‑native components.
- Certified Kubernetes Administrator (CKA); CKAD or CKS a plus.
- Experience with Cluster API (CAPI), controller‑runtime, or writing/maintaining Kubernetes operators.
- Experience with edge, bare‑metal, or virtualization platforms (VMware, KubeVirt, Hyper‑V).
- Experience delivering CVE remediation and patch releases in a regulated, enterprise SaaS, or federal environment.
- Experience with MongoDB, Terraform, Ansible, or CI/CD pipelines (GitHub Actions).
- Experience using AI‑assisted debugging or engineering workflows with validation and governance controls.