What we do
We give a model a dose by turning one of its internal features up or down. A feature can be a steering vector, which is a direction in the model's activations, or a single feature of a sparse autoencoder trained on them. Then we measure what changes: how the target behaviour responds to the dose, what else gets worse, how long the effect lasts, how two doses combine, and whether the model reports that anything changed.
How we work
Every study is preregistered. We write down the hypotheses, the doses, the measures and the analysis plan before the full run. If a small pilot runs alongside the plan to check the plumbing, its numbers don't count, and anything it changed is recorded in a dated amendment.
We report every outcome we preregistered, with 95% intervals, and we publish null results on the same terms as positive ones.
We work only with open-weight models, so that anyone can check a result. For now every experiment runs on a CPU, with small models such as GPT-2 small and the public sparse autoencoders trained on them.
The name
Mount Ossa is a peak in British Columbia, next to Mount Pelion. The old phrase about piling Pelion on Ossa describes adding effort on top of effort. We prefer to add one dose at a time and write down what happens.
People and contact
Who runs Ossa Labs and how to reach us will go here.