Go Back
Choose this branch
Choose this branch
meritocratic
regular
democratic
hot
top
alive
80 posts
Oracle AI
Myopia
AI Boxing (Containment)
Deceptive Alignment
Deception
Acausal Trade
Self Fulfilling/Refuting Prophecies
Bounties (closed)
Parables & Fables
Superrationality
Values handshakes
Computer Security & Cryptography
88 posts
Conjecture (org)
Language Models
Refine
Agency
Deconfusion
Scaling Laws
Project Announcement
Encultured AI (org)
Tool AI
Definitions
PaLM
Prompt Engineering
43
Proper scoring rules don’t guarantee predicting fixed points
Johannes_Treutlein
4d
2
35
Side-channels: input versus output
davidad
8d
9
41
Steering Behaviour: Testing for (Non-)Myopia in Language Models
Evan R. Murphy
15d
16
145
Decision theory does not imply that we get to have nice things
So8res
2mo
53
85
Trying to Make a Treacherous Mesa-Optimizer
MadHatter
1mo
13
119
Monitoring for deceptive alignment
evhub
3mo
7
64
How likely is deceptive alignment?
evhub
3mo
21
30
Sticky goals: a concrete experiment for understanding deceptive alignment
evhub
3mo
13
41
Acceptability Verification: A Research Agenda
David Udell
5mo
0
258
The Parable of Predict-O-Matic
abramdemski
3y
42
25
Training goals for large language models
Johannes_Treutlein
5mo
5
33
The Speed + Simplicity Prior is probably anti-deceptive
7mo
29
17
Precursor checking for deceptive alignment
evhub
4mo
0
50
LCDT, A Myopic Decision Theory
adamShimi
1y
51
28
Discovering Language Model Behaviors with Model-Written Evaluations
evhub
4h
3
27
Take 11: "Aligning language models" should be weirder.
Charlie Steiner
2d
0
64
[Interim research report] Taking features out of superposition with sparse autoencoders
Lee Sharkey
7d
10
143
Conjecture: a retrospective after 8 months of work
Connor Leahy
27d
9
96
The Singular Value Decompositions of Transformer Weight Matrices are Highly Interpretable
beren
22d
27
108
What I Learned Running Refine
adamShimi
26d
5
178
Mysteries of mode collapse
janus
1mo
35
52
Conjecture Second Hiring Round
Connor Leahy
27d
0
31
Searching for Search
NicholasKees
22d
6
185
Simulators
janus
3mo
103
234
chinchilla's wild implications
nostalgebraist
4mo
114
41
Current themes in mechanistic interpretability research
Lee Sharkey
1mo
3
163
Language models seem to be much better than humans at next-token prediction
Buck
4mo
56
75
Inverse Scaling Prize: Round 1 Winners
Ethan Perez
2mo
16