Trends
Regels
The Shard Theory research program by Team Shard is based on the idea that we should directly study the what types of cognition/circuitry are encouraged by reinforcement learning, and reason mechanistically about how reward leads to them, rather than the classic Outer Alignment/Inner Alignment factorization which aims to create a reward function that matches human values and somehow give an AI a goal of maximizing that reward.
+3 extra uitkomsten
+7 extra uitkomsten
+5 extra uitkomsten
+4 extra uitkomsten
+2 extra uitkomsten