Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Weights & Biases is seeking a Manager / Senior Manager, Solutions Architecture to lead a regional team delivering complex deployments across AWS, GCP, Azure, and on-prem environments. You will stay hands-on with architecture reviews and critical escalations while building scalable processes and reference architectures.
You will hire and ramp the team in-region, establish deployment standards, and partner with Sales Engineering to validate pre-sales technical aspects.
We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren’t a 100% skill or experience matchStrong Linux/Unix command line experienceHands-on expertise with Docker, Kubernetes, and Helm charts, plus networking fundamentals (DNS, TLS, ingress, load balancing, proxies) and cloud-managed services (e.g., MySQL, object stores)3+ years managing a customer-facing technical team (Solutions Architecture, SRE, TAM, professional services, or field engineering) with direct hiring and performance-management responsibility — 5+ years for Senior ManagerExperience supporting both multi-tenant SaaS and customer-managed / on-prem deployments of the same productWorking proficiency in Python and familiarity with ML workflows or toolingApplicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorizationDemonstrated track record of systematically diagnosing and resolving production infrastructure issues in customer-owned environmentsProficiency with at least one cloud platform (AWS, GCP, Azure); working knowledge of more than one is a plus7+ years total hands-on infrastructure, platform, or SRE experience — 10+ for Senior ManagerProduction experience with Infrastructure as Code, preferably TerraformDemonstrated experience delivering technical workshops, demos, or enablement sessions to customer engineering teamsOr (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agencyYou’re an expert in building and running containerized, distributed systemsSaaS, web service, or distributed systems operations experience at scaleExperience building a regional technical team from a small headcount into a durable function with defined coverage, on-call, and rampYou’re curious about how to scale complex ML systems in production environmentsExperience with air-gapped, sovereign, or regulated-industry deploymentsExperience with GPU infrastructure and distributed training workloadsDeep proficiency in Kubernetes design patterns, including Operators and CRDsFamiliarity with data engineering and MLOps tooling (e.g., Ray, Kubeflow, Slurm, Airflow, SageMaker, Vertex AI)You’d rather write the runbook once than answer the same escalation five timesYou’re comfortable being the calm one on a customer bridge call at 2am, and the one who makes sure it doesn’t happen againYou love diving into infrastructure problems and solving them systematically — and you get more satisfaction from your team solving them than from solving them alone