Member of Technical Staff - Research Software Engineer - Safety Evaluations Infrastructure

B Capital

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited vacation days
Visa sponsorship

Job summary

Reflection is seeking a Research Software Engineer on the Safety team to design, build, and own infrastructure for evaluating open models in high-consequence domains. You will create sandboxed environments, secure data pipelines, and robust tooling to support trusted experimentation.

You will collaborate with safety researchers, legal, and domain experts, translating evaluation needs into scalable, auditable systems with strong access controls and encryption at rest and in transit.

Qualifications

  • Strong software engineering skills, particularly in Python.
  • Experience building sandboxed, isolated, or security-sensitive execution environments.
  • Solid grounding in security engineering fundamentals: least privilege, need-to-know, RBAC, secrets management, encryption, audit logging, defense-in-depth.
  • Experience building data pipelines and handling sensitive or restricted data.
  • Ability to own end-to-end problems in cross-functional settings.
  • Comfort working on sensitive projects requiring discretion and integrity.
  • Thrive in a fast-paced startup environment with a bias toward action.

Responsibilities

  • Design and build secure, sandboxed infrastructure for running sensitive model evaluations, including CBRN and other dangerous-capability domains.
  • Build controlled data pipelines and storage for sensitive evaluation material, applying least-privilege and need-to-know access, RBAC, encryption at rest and in transit, audit logging, and data-minimization safeguards.
  • Partner with safety researchers to translate evaluation designs into reliable, reproducible, and scalable systems.
  • Build eval-orchestration tooling and harnesses for high-throughput evaluations against models in isolated environments.
  • Develop infrastructure for measuring AI capability uplift and integrate results into release pipelines.
  • Implement guardrails, monitoring, and compartmentalization across compute and data.
  • Write production-quality Python and related tooling for data processing and evaluation systems.
  • Improve reliability, security posture, and developer experience of the safety platform.

Job description

Our Mission

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

Our Mission

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

About the Role

As a Research Software Engineer on the Safety team, you will design, build, and own the infrastructure used to run our most sensitive model evaluations — including evaluations in CBRN (chemical, biological, radiological, and nuclear), child safety, and other dangerous-capability domains. These evaluations inform release decisions for our open models, so the systems you build must be secure, isolated, reproducible, and trustworthy under scrutiny.

This is a deeply technical, high-ownership role at the intersection of platform engineering, security, and safety research. You will partner closely with domain experts, legal, and safety researchers to turn their evaluation needs into robust, scalable infrastructure: sandboxed execution environments, controlled data pipelines for sensitive material, access controls, audit logging, and the tooling that lets researchers safely elicit and measure model capabilities in high-consequence areas.

What You'll Do
  • Design and build secure, sandboxed infrastructure for running sensitive model evaluations, including CBRN and other dangerous-capability domains.

  • Build controlled data pipelines and storage for sensitive evaluation material, applying least-privilege and need-to-know access, role-based access control (RBAC), encryption at rest and in transit, audit logging, and data-minimization safeguards.

  • Partner with safety researchers and domain experts to translate evaluation designs into reliable, reproducible, and scalable systems.

  • Build eval-orchestration tooling and harnesses that let researchers run high-throughput evaluations against models and agents in isolated environments.

  • Develop infrastructure for measuring AI capability uplift in high-consequence domains, and integrate results into the pipelines that inform release decisions.

  • Implement guardrails, monitoring, and compartmentalization so sensitive work stays appropriately siloed, applying least-privilege, need-to-know, and defense-in-depth principles across compute, data, and tooling.

  • Write production-quality Python (and related tooling) for high-throughput data processing and evaluation systems.

  • Improve the reliability, security posture, and developer experience of the safety team's evaluation platform over time.

About You
  • Strong software engineering skills, particularly in Python, with a track record of building reliable, scalable infrastructure or platform systems.

  • Experience building sandboxed, isolated, or otherwise security-sensitive execution environments (e.g., containerization, VM isolation, secure compute) for Trust and Safety teams.

  • Solid grounding in security engineering fundamentals: principle of least privilege, need-to-know access, role-based access control (RBAC), secrets management, encryption, audit logging, compartmentalization, and defense-in-depth design.

  • Experience building data pipelines and handling sensitive or restricted data with appropriate safeguards.

  • Ability to own entire problems end-to-end, including ambiguous, cross-functional ones.

  • Comfort working on sensitive projects that require discretion, integrity, and sound judgment.

  • Thrive in a fast-paced, high-agency startup environment with a bias toward action.

Strong candidates may also have
  • Experience building evaluation, benchmarking, or experimentation infrastructure for ML systems.

  • Experience working with LLMs, agents, or ML training/inference pipelines.

  • Familiarity with dangerous-capability or dual-use domains (CBRN, cyber, etc.) and the information-security considerations they involve.

  • Familiarity with compliance frameworks relevant to sensitive data handling.

We encourage you to apply even if you don't meet every qualification. Not all strong candidates will match every item listed.

What We Offer:

We believe that to make intelligence open and accessible to all, you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company, and help define the future of open foundational models.

We want you to do the most impactful work of your career with the confidence that you and the people you care about most are supported.

  • Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.

  • Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.

  • Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.

  • Meals: Lunch and dinner are provided in the office daily.

  • Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.

  • Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.

  • Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.

  • Team building: We have regular off-sites, happy hours, and team celebrations.

Export Control Notice: This position may require access to technology or source code subject to the U.S. Export Administration Regulations. Any offer of employment for this role may be conditioned on the Company's ability to provide the candidate with access to such technology or source code in compliance with applicable U.S. export control laws, which may require the Company to seek government authorization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Safety
Member of Technical Staff - Safety

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Top-tier compensation
Stock options
Health & wellness
+3
Research Program Manager - Model Evals and Safety
Research Program Manager - Model Evals and Safety

Visa Hunt • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Manager of Technology & Security Engineering
Manager of Technology & Security Engineering

Reflection AI • New York (NY)

On-site
USD 250,000 - 350,000
Top-tier compensation
Stock options
Comprehensive health benefits
+4
Member of Technical Staff - Security Engineer
Member of Technical Staff - Security Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 280,000
Top-tier compensation
Stock options
Health & wellness benefits
+3
Manager of Technology & Security Engineering
Manager of Technology & Security Engineering

Visa Hunt • New York (NY)

On-site
USD 260,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
+5
Research Program Manager - Model Evals and Safety
Research Program Manager - Model Evals and Safety

aijoblist • San Francisco (CA)

On-site
USD 130,000 - 180,000
Top-tier compensation
Comprehensive medical, dental, vision insurance
Paid parental leave
+2
Member of Technical Staff - Evaluations
Member of Technical Staff - Evaluations

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Stock options
Health & wellness
Meals provided
+4
Subject Matter Expert - Election Integrity, Violent Extremism, and Fraud
Subject Matter Expert - Election Integrity, Violent Extremism, and Fraud

B Capital • California (MO)

On-site
USD 150,000 - 210,000
Top-tier compensation
Equity
Comprehensive health benefits
+5
Member of Technical Staff - Distributed Systems Engineer
Member of Technical Staff - Distributed Systems Engineer

Visa Hunt • New York (NY)

On-site
USD 140,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+5
Forward Deployed Engineer - LLM Post-training
Forward Deployed Engineer - LLM Post-training

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+5