Tree of Tags

Go Back

You can't go any further

You can't go any further

meritocratic regular democratic

hot top alive

6 posts Reward Functions

10 posts Wireheading

86 Scaling Laws for Reward Model Overoptimization

leogao

2mo

11

76 Seriously, what goes wrong with "reward the agent when it makes you smile"?

TurnTrout

4mo

41

40 A Short Dialogue on the Meaning of Reward Functions

Leon Lang

1mo

0

26 The reward engineering problem

paulfchristiano

3y

3

25 Reward model hacking as a challenge for reward learning

Erik Jenner

8mo

1

20 $100/$50 rewards for good references

Stuart_Armstrong

1y

5

69 Towards deconfusing wireheading and reward maximization

leogao

3mo

7

47 Draft papers for REALab and Decoupled Approval on tampering

Jonathan Uesato

2y

2

42 Four usages of "loss" in AI

TurnTrout

2mo

18

26 Wireheading is in the eye of the beholder

Stuart_Armstrong

3y

10

25 Wireheading as a potential problem with the new impact measure

Stuart_Armstrong

4y

20

22 Defining AI wireheading

Stuart_Armstrong

3y

9

21 Wireheading and discontinuity

Michele Campolo

2y

4

17 Model-based RL, Desires, Brains, Wireheading

Steven Byrnes

1y

1

16 Value extrapolation vs Wireheading

Stuart_Armstrong

6mo

1

10 Note on algorithms with multiple trained components

Steven Byrnes

7h

1