1 result
for sae
-
- **Linear representation hypothesis**: high-level concepts are represented as *linear directions* in activation space. Empirical support exists from word embeddings up to large language models — but the source states plainly that it "does not hold up universally." Much of the me…field/mechanistic-interpretability · interpretability, mechanistic-interpretability, sae, circuits, llm, safety