What open-source tooling does AI safety research need right now?

TL;DR: I'm running a survey to find out what open-source tooling AI safety researchers are wishing for right now. Anything that makes your daily workflow easier or faster, or that allows you to run experiments that you previously couldn't because the setup was so complicated. You can find the survey here. You don't need to read this post to fill it out. In this post I will explain the motivation behind this survey and what answers I'm looking for.

Motivation

I claim that people eager to get into AI safety research would benefit from having a curated list of open-source projects that researchers in this field would be excited about, and that the field as a whole would benefit from this as well.

Based on my own experiences and from reading around social media threads, I think many young professionals and students trying to break into the field of AI safety research right now are having trouble finding their footing. With the field growing, it seems to be getting harder to get into fellowships like MATS or to secure high-quality AI safety mentorship in other ways. In my opinion, however, many people are still very eager to contribute.

I believe that a great way for these people to contribute to the field and simultaneously build valuable career capital is building useful open-source tools or contributing meaningfully to existing ones. However, it can be very hard to know what is useful or needed when you are not deeply involved in a research field. One could argue that the fix to that would be to simply get your hands dirty and do research. I largely agree with that, but I think without getting good mentorship, it can be hard not to get stuck building things that aren't really meaningful and don't provide value to the field. I argue it would be a better use of many people's time to build something they know is valuable, and use that career capital to land roles or fellowships that can expose them to mentorship if that is what they ultimately want to achieve.

This is the reason I'm running this survey. I want to borrow the research-taste of people deeply involved in the field in order to curate a list of open-source projects that could meaningfully help researchers carry out experiments faster or allow them to perform more sophisticated experiments.

I think the contribution of this survey is three-fold:

Breadth: Even the people that are deeply involved in a sub-field of AI safety will have a limited grasp of fields far outside their area. Thus, the results are useful for giving everyone in the field a broad overview of what is needed right now, which could be useful for, e.g., allocating funds.Opportunity: This survey gives talented and ambitious researchers and engineers an opportunity to build career capital, and also the comfort of knowing that what they are building is actually needed by the field.Decreased research friction: If a lot of people build the tools researchers in this field want and need, I expect this to non-trivially increase researcher productivity and decrease the friction they experience in their day-to-day workflows.

Another potential objection to my motivation behind this survey is that if you are not deeply involved in the field, even if you have the idea, you won't be able to build it well. I think that's a fair objection, but I believe that it is arguably the better scenario to have a vetted idea that you don't know how to implement perfectly compared to having an unvetted idea that you also don't know how to implement perfectly. Furthermore, I expect many answers to the survey to be ideas for features to existing tools rather than standalone libraries, in which case quality will be controlled by the maintainers of those libraries.

A similar survey was done in 2022[1], but it focused more on language-model-powered tools than on open-source tooling broadly, and furthermore, it's ~3.5 years old and the field has moved a lot since then.

I hope this post will be of great interest to the LessWrong community: both to the researchers who I hope will take the time to fill out the survey, and to the people trying to break into the field, whom I hope the results will help.

How does the survey work?

The survey itself is a short Google Form that will take ~5 minutes to fill out. I deliberately kept it short to be able to get responses from the very busy people within this field. The form itself asks for the field you work in, how much experience you have working on AI safety research directly, and where you work. Then, it asks you to describe what tools you wish existed that would make your life as a researcher easier. A good example of what I'm looking for here is TransformerLens. It was built because Neel Nanda was frustrated with the state of open-source tooling for mechanistic interpretability research[2]. Now, it is a fundamental tool to the field of mechanistic interpretability. I'm looking for tools that could be equally meaningful, or just PRs that add features to existing tools that would significantly aid you in your research. Some questions to consider that might help you come up with an answer:

What is something in your daily workflow that you keep having to implement, that would be amazing if it would just be a library you import?In a current tool you are using, what is a feature you wish it would have that you don't have the time to implement yourself?Are there experiments or analyses that you currently are not able to carry out because the tools to carry them out do not exist?When will the results be published?

I will let the survey run until the 1st of September 2026. Afterwards, I will analyze the results and publish another blog post here on LessWrong with a list of the open-source projects the field wishes for. I will aggregate the information so it's clear which recommendations come from people working in which field and with what seniority etc., while being careful not to publish the results in a way where they could be deanonymizing (unless a respondent explicitly gave consent for that).

The survey

You can fill out the survey here.

About me

This is my first LessWrong post, so I'm using this opportunity to introduce myself. I'm Fabian Degen. I helped build the TransformerLens 3.0 release on a Manifund grant, and I'm currently doing my MSc at Oxford working on LLM agents and causal reasoning.

Acknowledgements

Thanks to Neel Nanda for quick feedback on the survey.

^^
AI Article