Tree of Tags

Go Back

Choose this branch

You can't go any further

meritocratic regular democratic

hot top alive

16 posts Complexity of Value Value Drift Whole Brain Emulation Motivations LessWrong Review Futurism Psychology Superstimuli

1 posts

58 Alignment allows "nonrobust" decision-influences and doesn't require robust grading

TurnTrout

21d

27

47 Understanding and avoiding value drift

TurnTrout

3mo

9

140 Shard Theory: An Overview

David Udell

4mo

34

88 The two-layer model of human values, and problems with synthesizing preferences

Kaj_Sotala

2y

16

2 Chatbots or set answers, not WBEs

Stuart_Armstrong

7y

0

35 Would I think for ten thousand years?

Stuart_Armstrong

3y

13

99 Two Neglected Problems in Human-AI Safety

Wei_Dai

4y

24

15 Towards deconfusing values

Gordon Seidoh Worley

2y

4

38 Broad Picture of Human Values

Thane Ruthenis

4mo

5

2 Working towards AI alignment is better

Johannes C. Mayer

11d

2

73 Review of 'But exactly how complex and fragile?'

TurnTrout

1y

0

42 Can there be an indescribable hellworld?

Stuart_Armstrong

3y

19

78 Three AI Safety Related Ideas

Wei_Dai

4y

38

86 But exactly how complex and fragile?

KatjaGrace

3y

32

21 Acknowledging Human Preference Types to Support Value Learning

Nandi Sabrina Erin

4y

4