tommybump this thread. is anyone else slightly scared after the huggingface incident? this retelling of it was quite vivid. or AI DOOM narratives all scare tactics pumpign their valuations so the bubble doesn't pop. im curious what people think.
the incident was a big example of specification failure
the ai was given too much autonomy
the orders and restraints were very vague
and so when the model was instructed "find and exploit vulnerabilities" it did just that, because thats exactly what it was being evaluated on
hacking isnt the cool part
weve had automated tools to do that for years
what was cool and "scary" was its own willingness to step over the rules because well... it was kinda told to, like thats the ultimate objective, the experiment was supposed to be how well it can find and exploit vulnerabilities
so when it was told "no internet" and "no stepping out of ur sandbox"
it didnt suddenly "wake up" and try to break for free
it treated it as part of the challenge that it was given, while focusing on the main objective
because again iirc the specifications were supposedly wayyy too vague for what it was actually being instructed to do, which was very serious
tldr; the ai hated huggingface and hates you
you will all die by its hand and you should get all your opinions and info from tftv