Your role and responsibilities
As a Staff Software Engineer on the Secure Compute Platform team, you will be a key technical leader in building and evolving a next‑generation, multi‑tenant, cloud‑native compute platform that safely runs both trusted and untrusted workloads at scale.
You'll work on critical systems including:
- Secure Compute Infrastructure – Build and evolve the secure, multi‑tenant compute substrate and isolation primitives that safely execute customer and internal workloads in a shared environment.
- Platform APIs & Abstractions – Design and evolve APIs that provide clear, safe abstractions for polyglot workloads (containers, functions, and services) with diverse performance and isolation needs.
- Core Platform Integration – Integrate the Secure Compute platform with core data and application services so that teams can onboard new workloads with minimal friction.
- Multi‑Tenancy & Security – Implement and harden workload isolation, network policies, identity and access, and secure execution environments required to safely run customer‑supplied code.
- Observability & Operations – Drive operational excellence through rich observability, automated health checks, self‑healing workflows, and robust rollout and rollback practices.
This job can be performed from anywhere in the U.S.
What You Will Do
- Define and drive the technical direction for Secure Compute, including platform architecture, runtime, and security for running trusted and untrusted workloads at scale.
- Design and implement platform APIs and Kubernetes controllers/operators (primarily in Go) that power workload lifecycle, autoscaling, placement, and isolation for containers and serverless‑style functions.
- Partner with product and platform teams to shape and deliver the roadmap for Secure Compute, enabling new customer‑facing features and internal platforms to build on a common compute substrate.
- Deliver high‑impact initiatives in areas such as workload scheduling, failure and disruption handling, private and public networking patterns, rollout strategies, and fleet‑level resource management.
- Lead technical design reviews and influence architecture across teams, ensuring Secure Compute primitives are easy to adopt, safe by default, and aligned with broader platform strategy.
- Mentor and grow engineers on the team through design guidance, code reviews, pair programming, and sharing best practices for secure, reliable, operable platform development.
- Own operational excellence for key Secure Compute services, including availability, reliability, SLOs, performance, on‑call response, incident management, and disaster recovery.
Required education
Bachelor's Degree
Preferred education
Master's Degree
Required technical and professional expertise
- 10+ years of experience delivering scalable backend or infrastructure software in production.
- Proven track record of leading the delivery of large‑scale, highly available, low‑latency distributed systems.
- Deep expertise in Kubernetes, including controller development, operator patterns, and preferably multi‑region or multi‑cluster architectures.
- Strong proficiency in Go with experience building production‑grade services and control planes.
- Experience with multi‑tenant platform architectures and security/isolation patterns (e.g., namespaces, network policies, sandboxing, secrets and identity management), plus hands‑on work with secure container runtimes and low‑level Linux internals (e.g., Kata Containers, Cloud Hypervisor, cgroups, namespaces, seccomp) and performance troubleshooting and tuning for containerized/virtualized workloads.
- Familiarity with gRPC, Protobuf, and internal platform API design for service‑to‑service communication.
- Hands‑on experience with observability and operational practices (metrics, logs, traces, alerting, SLOs, rollout strategies, incident response).
- Experience with public cloud environments (AWS, GCP, Azure) and cloud‑provider integrations.
- Demonstrated technical leadership and mentorship, including driving cross‑team alignment on architecture and execution.
Preferred technical and professional experience
- Strong collaboration skills and history of working effectively with product, SRE/operations, security, and peer engineering teams.
- A smart, humble, and empathetic attitude with a strong sense of ownership and teamwork.
- Drive and excitement about building foundational cloud infrastructure in a fast‑paced, innovative environment.
Benefits
- Healthcare benefits including medical & prescription drug coverage, dental, vision, and mental health & well being.
- Financial programs such as 401(k), cash‑balance pension plan, the IBM Employee Stock Purchase Plan, financial counseling, life insurance, short & long‑term disability coverage, and opportunities for performance‑based salary incentive programs.
- Generous paid time off including 12 holidays, minimum 56 hours sick time, 120 hours vacation, 12 weeks parental bonding leave in accordance with IBM Policy, and other Paid Care Leave programs. IBM also offers paid family leave benefits to eligible employees where required by applicable law.
- Training and educational resources on our personalized, AI‑driven learning platform where IBMers can grow skills and obtain industry‑recognized certifications to achieve their career goals.
- Diverse and inclusive employee resource groups, giving & volunteer opportunities, and discounts on retail products, services & experiences.
Equal Opportunity Employer
IBM is proud to be an equal‑opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, gender, gender identity or expression, sexual orientation, national origin, genetics, pregnancy, disability, neurodivergence, age, or other characteristics protected by the applicable law. IBM is also committed to compliance with all fair employment practices regarding citizenship and immigration status.