---
title: "My current research and request for collaborators"
description: "The bio I wrote for EAG Bay Area 2024: what I was working on, what I wanted to argue about, and the three kinds of collaborator I was looking for."
published: 2024-01-23
tags: ["Automating alignment research"]
importance: 3
docStatus: "notes"
audio: "https://pub-4ee2f71bc29541a7a6e8d9694f0a1b21.r2.dev/65b01fd53de6b56ce4c0cbf5/audio.mp3"
author: "Jacques Thibodeau"
canonical: "https://jacquesthibodeau.com/update-on-what-ive-been-doing/"
---
**I wrote this as a bio for EAG Bay Area 2024. I'm sharing this here because it gives an overview of what I've been working on and might reach someone who wants to chat or collaborate.**

Hey! I'm Jacques. I'm an independent technical alignment researcher with a background in physics and experience in government (social innovation, strategic foresight, mental health and energy regulation). [Twitter/X](https://x.com/jacquesthibs).

CURRENT WORK

-   Collaborating with [Quintin Pope](https://www.lesswrong.com/users/quintin-pope?from=post_header) on our **Supervising AIs Improving AIs agenda** (*making automated AI science safe and controllable*). The current project involves a new method allowing **unsupervised model behaviour evaluations**. [Our agenda](https://www.lesswrong.com/posts/7e5tyFnpzGCdfT4mR/research-agenda-supervising-ais-improving-ais).
-   I'm a research lead in the AI Safety Camp for a [project on **stable reflectivity**](https://www.lesswrong.com/posts/RHojGPWLgdFLk3PAt/aisc-project-benchmarks-for-stable-reflectivity) (testing models for metacognitive capabilities that impact future training/alignment).
-   [Accelerating Alignment](https://docs.google.com/document/d/1g-p_8d-7c29WHeA_sQih1YZ7QDlpHd2kGL_IskF8Ix4/edit?usp=sharing): **augmenting alignment researchers** using AI systems. A relevant [talk](https://www.youtube.com/watch?v=rDK0XxFyrzQ) I gave. Relevant [survey post](https://www.lesswrong.com/posts/a2io2mcxTWS4mxodF/results-from-a-survey-on-tool-use-and-workflows-in-alignment).
-   Other research that currently interests me: multi-polar AI worlds (and *how that impacts post-deployment model behaviour*), [**understanding-based interpretability**](https://www.lesswrong.com/posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluations), improving evals, designing safer training setups, interpretable architectures, and **limits of current approaches** (what would a new paradigm that addresses these limitations look like?).
-   Used to focus more on [model editing](https://www.lesswrong.com/posts/QL7J9wmS6W2fWpofd/but-is-it-really-in-rome-an-investigation-of-the-rome-model), rethinking interpretability, [causal scrubbing](https://www.lesswrong.com/posts/DFarDnQjMnjsKvW8s/practical-pitfalls-of-causal-scrubbing), etc.

<!--kg-card-begin: html-->
<details>
<summary><strong>TOPICS TO CHAT ABOUT</strong></summary>
<ul>
<li><strong>How do you expect AGI/ASI to actually develop</strong> (so we can align our research accordingly)? <strong>Will scale plateau?</strong> I’d like to get feedback on some of my thoughts on this.</li>
<li>How can we connect the dots between different approaches? For example, connecting the dots between <a href="https://arxiv.org/abs/2308.03296">Influence Functions</a>, Evaluations, Probes (detecting truthful direction), <a href="https://functions.baulab.info/">Function</a>/<a href="https://arxiv.org/abs/2310.15916">Task Vectors</a>, and <a href="https://arxiv.org/abs/2310.01405">Representation Engineering</a> to see if they can work together to give us a better picture than the sum of their parts.</li>
<li>Debate over which agenda actually contributes to solving the core AI x-risk problems.</li>
<li>What if the pendulum swings in the other direction, and we never get the benefits of safe AGI? Is open source really as bad as people make it out to be?</li>
<li>How can we make something like the <a href="https://vitalik.eth.limo/general/2023/11/27/techno_optimism.html#dacc">d/acc vision</a> (by Vitalik Buterin) happen?</li>
<li>How can we design a system that leverages AI to speed up progress on alignment? What would you value the most?</li>
<li>What kinds of orgs are missing in the space?</li>
</ul>
</details>
<!--kg-card-end: html-->

<!--kg-card-begin: html-->
<details>
<summary><strong>POTENTIAL COLLABORATIONS</strong></summary>
<p>Examples of projects I’d be interested in:</p>
<ol>
<li>Extending either the <a href="https://openai.com/research/weak-to-strong-generalization">Weak-to-Strong Generalization</a> paper or the <a href="https://arxiv.org/abs/2401.05566">Sleeper Agents</a> paper</li>
<li>Understanding the impacts of synthetic data on LLM training</li>
<li>Working on ELK-like research for LLMs</li>
<li>Experiments on <a href="https://arxiv.org/abs/2308.03296">influence functions</a> (studying the base model and its SFT, RLHF, iterative training counterparts; I heard that Anthropic is releasing code for this "soon")</li>
<li>Figure out how to resolve <a href="https://www.lesswrong.com/posts/jXjeYYPXipAtA2zmj/jacquesthibs-s-shortform?commentId=oYHoeJRRFAfopz5zX">the problems that arise in RLHF/RLAIF as you increase model size</a></li>
<li>Studying the interpolation/extrapolation distinction in LLMs</li>
<li>People seem overly focused on LLMs rather than what comes after. I want to understand how dangerous levels of agency could arise in the next big advance. I think one true danger that we need to concern ourselves with is autonomously agentic AI agents that can update their own goals (that go outside the bounds of humanity). <em>"Many applications will be much more autonomous, difficult to monitor or even understand, and potentially fully close loop, i.e the agent has a complex enough action space that it can copy itself, buy compute, run itself, etc."</em></li>
</ol>
<p>I’m also interested in talking to grantmakers for feedback on some projects I’d like to get funding for.</p>
<p>I’m slowly working on a guide for practical research productivity for alignment researchers to tackle <a href="https://roamresearch.com/#/app/Jacques-Second-Brain/page/A5eKKnA9t">low-hanging fruits</a> that can quickly improve productivity in the field. I’d like feedback from people with solid track records and productivity coaches.</p>
</details>
<!--kg-card-end: html-->

TYPES OF PEOPLE I'D LIKE TO COLLABORATE WITH

-   Strong math background, can understand [Influence Functions](https://arxiv.org/abs/2308.03296) enough to extend the work.
-   Strong machine learning engineering background. Can run ML experiments and fine-tuning runs with ease. Can effectively create data pipelines.
-   Strong application development background. I have various project ideas that could speed up alignment researchers; I'd be able to execute them much faster if I had someone to help me build my ideas fast.

## Sources

Every external link in this piece that has a captured card, with what that page said
when it was captured. The quoted lines below are not the author of this piece writing:
they are the linked page describing itself, recorded by `bun run link-cards` on the date
given, and kept so that a reader still has them if the original moves or goes away.

- **Research agenda: Supervising AIs improving AIs** — Quintin Pope, lesswrong.com, 2023-04-29
  <https://lesswrong.com/posts/7e5tyFnpzGCdfT4mR/research-agenda-supervising-ais-improving-ais>
  Captured 2026-08-28.

  > [This post summarizes some of the work done by Owen Dudney, Roman Engeler and myself (Quintin Pope) as part of the SERI MATS shard theory stream.] TL;DR Future prosaic AIs will likely shape their own development or that of successor AIs. We're trying to make sure they don't go insane. Summary There are two main ways…

- **AISC Project: Benchmarks for Stable Reflectivity** — jacquesthibs, lesswrong.com, 2023-11-13
  <https://lesswrong.com/posts/RHojGPWLgdFLk3PAt/aisc-project-benchmarks-for-stable-reflectivity>
  Captured 2026-08-28.

  > Apply to work on this project with me at AI Safety Camp 2024 before 1st December 2023. Summary Future prosaic AIs will likely shape their own development or that of successor AIs. We're trying to make sure they don't go insane. There are two main ways AIs can get better: by improving their training algorithms or by…

- **Accelerating Alignment with Language Models (meta-level)** — docs.google.com
  <https://docs.google.com/document/d/1g-p_8d-7c29WHeA_sQih1YZ7QDlpHd2kGL_IskF8Ix4/edit?usp=sharing>
  Captured 2026-08-28.

  > Accelerating Alignment with Language Models (meta-level) Project outcome: An AI Alignment Research Assistant that is actually being used and helpful to alignment researchers. It aims to become the AI system that every alignment researcher uses in their daily workflow to increase their productivit...

- **Accelerating Alignment with Jacques Thibodeau** — EA Software Engineers, youtube.com
  <https://youtube.com/watch?v=rDK0XxFyrzQ>
  Captured 2026-08-28.

- **Results from a survey on tool use and workflows in alignment research** — jacquesthibs, lesswrong.com, 2022-12-19
  <https://lesswrong.com/posts/a2io2mcxTWS4mxodF/results-from-a-survey-on-tool-use-and-workflows-in-alignment>
  Captured 2026-08-28.

  > In March 22nd, 2022, we released a survey with an accompanying post for the purpose of getting more insight into what tools we could build to augment alignment researchers and accelerate alignment research. Since then, we’ve also released a dataset, a manuscript (LW post), and the (relevant) Simulators post was…

- **Towards understanding-based safety evaluations** — evhub, lesswrong.com, 2023-03-15
  <https://lesswrong.com/posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluations>
  Captured 2026-08-28.

  > Thanks to Kate Woolverton, Ethan Perez, Beth Barnes, Holden Karnofsky, and Ansh Radhakrishnan for useful conversations, comments, and feedback. Recently, I have noticed a lot of momentum within AI safety specifically, the broader AI field, and our society more generally, towards the development of standards and…

- **But is it really in Rome? An investigation of the ROME model editing technique** — jacquesthibs, lesswrong.com, 2022-12-30
  <https://lesswrong.com/posts/QL7J9wmS6W2fWpofd/but-is-it-really-in-rome-an-investigation-of-the-rome-model>
  Captured 2026-08-28.

  > Thanks to Andrei Alexandru, Joe Collman, Michael Einhorn, Kyle McDonell, Daniel Paleka, and Neel Nanda for feedback on drafts and/or conversations which led to useful insights for this work. In addition, thank you to both William Saunders and Alex Gray for exceptional mentorship throughout this project. The majority…

- **Practical Pitfalls of Causal Scrubbing** — Jérémy Scheurer, lesswrong.com, 2023-03-27
  <https://lesswrong.com/posts/DFarDnQjMnjsKvW8s/practical-pitfalls-of-causal-scrubbing>
  Captured 2026-08-28.

  > TL;DR: We evaluate Causal Scrubbing (CaSc) on synthetic graphs with known ground truth to determine its reliability in confirming correct hypotheses and rejecting incorrect ones. First, we show that CaSc can accurately identify true hypotheses and quantify the degree to which a hypothesis is wrong. Second, we…

- **Studying Large Language Model Generalization with Influence Functions** — Roger Grosse and 16 others, arxiv.org, 2023-08-07
  <https://arxiv.org/abs/2308.03296>
  Captured 2026-08-28.

  > When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which training examples most contribute to a given behavior? Influence functions aim to answer a counterfactual: how would the model's parameters…

- **Function Vectors in Large Language Models** — functions.baulab.info
  <https://functions.baulab.info/>
  Captured 2026-08-28.

  > LLMs have an embedding space for functions that emerge from in-context learning.

- **In-Context Learning Creates Task Vectors** — Roee Hendel, Mor Geva, Amir Globerson, arxiv.org, 2023-10-24
  <https://arxiv.org/abs/2310.15916>
  Captured 2026-08-28.

  > In-context learning (ICL) in Large Language Models (LLMs) has emerged as a powerful new learning paradigm. However, its underlying mechanism is still not well understood. In particular, it is challenging to map it to the "standard" machine learning framework, where one uses a training set $S$ to find a best-fitting…

- **Representation Engineering: A Top-Down Approach to AI Transparency** — Andy Zou and 20 others, arxiv.org, 2023-10-02
  <https://arxiv.org/abs/2310.01405>
  Captured 2026-08-28.

  > In this paper, we identify and characterize the emerging area of representation engineering (RepE), an approach to enhancing the transparency of AI systems that draws on insights from cognitive neuroscience. RepE places population-level representations, rather than neurons or circuits, at the center of analysis,…

- **My techno-optimism** — vitalik.eth.limo
  <https://vitalik.eth.limo/general/2023/11/27/techno_optimism.html>
  Captured 2026-08-28.

- **Weak-to-strong generalization** — openai.com
  <https://openai.com/research/weak-to-strong-generalization>
  Captured 2026-08-28.

  > We present a new research direction for superalignment, together with promising initial results: can we leverage the generalization properties of deep learning to control strong models with weak supervisors?

- **Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training** — Evan Hubinger and 38 others, arxiv.org, 2024-01-10
  <https://arxiv.org/abs/2401.05566>
  Captured 2026-08-28.

  > Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity. If an AI system learned such a deceptive strategy, could we detect it and remove it using current state-of-the-art…

- **jacquesthibs's Shortform** — jacquesthibs, lesswrong.com, 2022-11-21
  <https://lesswrong.com/posts/jXjeYYPXipAtA2zmj/jacquesthibs-s-shortform?commentId=oYHoeJRRFAfopz5zX>
  Captured 2026-08-28.

- **Roam Research – A note taking tool for networked thought.** — roamresearch.com
  <https://roamresearch.com/>
  Captured 2026-08-28.

  > As easy to use as a word document or bulleted list, and as powerful for finding, collecting, and connecting related ideas as a graph database. Collaborate with others in real time, or store all your data locally.

## Terms used

The author's own definitions for the glossary terms this piece uses. These are his words,
not a standard reference.

- **AGI** — Artificial general intelligence: a system with human level cognitive ability across domains rather than in one narrow task.
- **ASI** — Artificial superintelligence: a system that outperforms the best humans at essentially every cognitive task.
- **EAG** — Effective Altruism Global: the conference series where much of this community meets in person.
  See also: <https://jacquesthibodeau.com/what-does-effective-in-ea-mean-to-you/>
- **extrapolation** — Answering a question that falls outside the region the training data covers, past its edge rather than between its points. How far outside, and in which direction, is what the word leaves unsaid and what any argument using it has to supply.
  Source: Balestriero, Pesenti and LeCun, Learning in High Dimension Always Amounts to Extrapolation <https://arxiv.org/abs/2110.09485>
- **interpolation** — Answering a question that falls inside the region the training data already covers, between the points rather than past their edge. The strict version is geometric: a point interpolates when it lies inside the convex hull of the training set, and in high dimensions almost no point does.
  Source: Balestriero, Pesenti and LeCun, Learning in High Dimension Always Amounts to Extrapolation <https://arxiv.org/abs/2110.09485>
- **LLM** — Large language model: a neural network trained on very large amounts of text to predict what comes next.
- **RLAIF** — Reinforcement learning from AI feedback: the same loop as RLHF with the preference judgements produced by a model instead of a person.
- **RLHF** — Reinforcement learning from human feedback: training a model against a reward signal learned from human preference comparisons.
- **SFT** — Supervised fine tuning: continuing training on curated input and output pairs, usually before any preference learning.
