Go Back
Choose this branch
Choose this branch
meritocratic
regular
democratic
hot
top
alive
32 posts
Machine Learning (ML)
OpenAI
Lottery Ticket Hypothesis
19 posts
DeepMind
Truth, Semantics, & Meaning
Anthropic
Honesty
Map and Territory
Calibration
47
Reframing inner alignment
davidad
9d
13
226
Common misconceptions about OpenAI
Jacob_Hilton
3mo
138
52
A Data limited future
Donald Hobson
4mo
25
49
Steganography in Chain of Thought Reasoning
A Ray
4mo
13
40
Prosaic AI alignment
paulfchristiano
4y
10
89
Survey of NLP Researchers: NLP is contributing to AGI progress; major catastrophe plausible
Sam Bowman
3mo
6
20
Train first VS prune first in neural networks.
Donald Hobson
5mo
5
14
[MLSN #5]: Prize Compilation
Dan H
2mo
1
14
Grouped Loss may disfavor discontinuous capabilities
Adam Jermyn
5mo
2
125
A Bird's Eye View of the ML Field [Pragmatic AI Safety #2]
Dan H
7mo
5
20
My thoughts on OpenAI's Alignment plan
Donald Hobson
10d
0
26
Discussion on the machine learning approach to AI safety
Vika
4y
3
57
Unsolved ML Safety Problems
jsteinhardt
1y
2
12
Automated Fact Checking: A Look at the Field
Hoagy
1y
0
265
A challenge for AGI organizations, and a challenge for readers
Rob Bensinger
19d
30
102
Clarifying AI X-risk
zac_kenton
1mo
23
80
Paper: Discovering novel algorithms with AlphaTensor [Deepmind]
LawrenceC
2mo
18
364
DeepMind alignment team opinions on AGI ruin arguments
Vika
4mo
34
104
Caution when interpreting Deepmind's In-context RL paper
Sam Marks
1mo
6
28
Paper: In-context Reinforcement Learning with Algorithm Distillation [Deepmind]
LawrenceC
1mo
5
52
Autonomy as taking responsibility for reference maintenance
Ramana Kumar
4mo
3
64
Toy Models of Superposition
evhub
3mo
2
25
Bridging syntax and semantics, empirically
Stuart_Armstrong
4y
4
32
AlphaGo Zero and capability amplification
paulfchristiano
3y
23
21
Knowledge is not just precipitation of action
Alex Flint
1y
6
29
The accumulation of knowledge: literature review
Alex Flint
1y
3
65
Truthful LMs as a warm-up for aligned AGI
Jacob_Hilton
11mo
14
30
Finding the variables
Stuart_Armstrong
3y
1