Safety Evaluations Platform Engineer – Secure Infra & Data

B Capital

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited vacation days
Visa sponsorship

Job summary

Reflection is seeking a Research Software Engineer on the Safety team to design, build, and own infrastructure for evaluating open models in high-consequence domains. You will create sandboxed environments, secure data pipelines, and robust tooling to support trusted experimentation.

You will collaborate with safety researchers, legal, and domain experts, translating evaluation needs into scalable, auditable systems with strong access controls and encryption at rest and in transit.

Qualifications

  • Strong software engineering skills, particularly in Python.
  • Experience building sandboxed, isolated, or security-sensitive execution environments.
  • Solid grounding in security engineering fundamentals: least privilege, need-to-know, RBAC, secrets management, encryption, audit logging, defense-in-depth.
  • Experience building data pipelines and handling sensitive or restricted data.
  • Ability to own end-to-end problems in cross-functional settings.
  • Comfort working on sensitive projects requiring discretion and integrity.
  • Thrive in a fast-paced startup environment with a bias toward action.

Responsibilities

  • Design and build secure, sandboxed infrastructure for running sensitive model evaluations, including CBRN and other dangerous-capability domains.
  • Build controlled data pipelines and storage for sensitive evaluation material, applying least-privilege and need-to-know access, RBAC, encryption at rest and in transit, audit logging, and data-minimization safeguards.
  • Partner with safety researchers to translate evaluation designs into reliable, reproducible, and scalable systems.
  • Build eval-orchestration tooling and harnesses for high-throughput evaluations against models in isolated environments.
  • Develop infrastructure for measuring AI capability uplift and integrate results into release pipelines.
  • Implement guardrails, monitoring, and compartmentalization across compute and data.
  • Write production-quality Python and related tooling for data processing and evaluation systems.
  • Improve reliability, security posture, and developer experience of the safety platform.

Job description

Reflection is seeking a Research Software Engineer on the Safety team to design, build, and own infrastructure for evaluating open models in high-consequence domains. You will create sandboxed environments, secure data pipelines, and robust tooling to support trusted experimentation.

You will collaborate with safety researchers, legal, and domain experts, translating evaluation needs into scalable, auditable systems with strong access controls and encryption at rest and in transit.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Research Software Engineer - Safety Evaluations Infrastructure
Member of Technical Staff - Research Software Engineer - Safety Evaluations Infrastructure

B Capital • San Francisco (CA)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+4
Staff Software Engineer, Safeguards Evaluations & Trust
Staff Software Engineer, Safeguards Evaluations & Trust

Menlo Ventures • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Flexible working hours
Generous vacation and parental leave
Office space for collaboration
Staff AI Safety Engineer — Red Team & Guardrails
Staff AI Safety Engineer — Red Team & Guardrails

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Top-tier compensation
Stock options
Health & wellness
+3
Research Program Manager - Model Evals and Safety
Research Program Manager - Model Evals and Safety

Visa Hunt • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Safety Evaluations Analyst, Safeguards
Safety Evaluations Analyst, Safeguards

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 230,000 - 270,000
Competitive compensation
Flexible working hours
Generous vacation
+3
Research Program Manager — Model Safety & Eval
Research Program Manager — Model Safety & Eval

Reflection • New York (NY)

On-site
USD 120,000 - 150,000
Top-tier compensation
Comprehensive health benefits
Fully paid parental leave
+2
Safety Evaluations Specialist, Safeguards
Safety Evaluations Specialist, Safeguards

Anthropic • United States

Remote
USD 120,000 - 180,000
Lead Software Engineer – Safeguards Review Tooling
Lead Software Engineer – Safeguards Review Tooling

United States Digital Space LLC • San Francisco (CA)

On-site
USD 320,000 - 485,000
Safety Evaluations Lead - Safeguards Enforcement
Safety Evaluations Lead - Safeguards Enforcement

United States Digital Space LLC • United States

Hybrid
USD 230,000 - 270,000
Platform Security Engineer
Platform Security Engineer

Thinking Machines Lab • San Francisco (CA)

On-site
USD 200,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1