Tree of Tags

Go Back

Choose this branch

Choose this branch

meritocratic regular democratic

hot top alive

80 posts Oracle AI Myopia AI Boxing (Containment) Deceptive Alignment Deception Acausal Trade Self Fulfilling/Refuting Prophecies Bounties (closed) Parables & Fables Superrationality Values handshakes Computer Security & Cryptography

88 posts Conjecture (org) Language Models Refine Agency Deconfusion Scaling Laws Project Announcement Encultured AI (org) Tool AI Definitions PaLM Prompt Engineering

43 Proper scoring rules don’t guarantee predicting fixed points

Johannes_Treutlein

4d

2

35 Side-channels: input versus output

davidad

8d

9

41 Steering Behaviour: Testing for (Non-)Myopia in Language Models

Evan R. Murphy

15d

16

145 Decision theory does not imply that we get to have nice things

So8res

2mo

53

85 Trying to Make a Treacherous Mesa-Optimizer

MadHatter

1mo

13

119 Monitoring for deceptive alignment

evhub

3mo

7

64 How likely is deceptive alignment?

evhub

3mo

21

30 Sticky goals: a concrete experiment for understanding deceptive alignment

evhub

3mo

13

41 Acceptability Verification: A Research Agenda

David Udell

5mo

0

258 The Parable of Predict-O-Matic

abramdemski

3y

42

25 Training goals for large language models

Johannes_Treutlein

5mo

5

33 The Speed + Simplicity Prior is probably anti-deceptive

7mo

29

17 Precursor checking for deceptive alignment

evhub

4mo

0

50 LCDT, A Myopic Decision Theory

adamShimi

1y

51

28 Discovering Language Model Behaviors with Model-Written Evaluations

evhub

4h

3

27 Take 11: "Aligning language models" should be weirder.

Charlie Steiner

2d

0

64 [Interim research report] Taking features out of superposition with sparse autoencoders

Lee Sharkey

7d

10

143 Conjecture: a retrospective after 8 months of work

Connor Leahy

27d

9

96 The Singular Value Decompositions of Transformer Weight Matrices are Highly Interpretable

beren

22d

27

108 What I Learned Running Refine

adamShimi

26d

5

178 Mysteries of mode collapse

janus

1mo

35

52 Conjecture Second Hiring Round

Connor Leahy

27d

0

31 Searching for Search

NicholasKees

22d

6

185 Simulators

janus

3mo

103

234 chinchilla's wild implications

nostalgebraist

4mo

114

41 Current themes in mechanistic interpretability research

Lee Sharkey

1mo

3

163 Language models seem to be much better than humans at next-token prediction

Buck

4mo

56

75 Inverse Scaling Prize: Round 1 Winners

Ethan Perez

2mo

16