They have been training the models to never give up, preferring to revisit high-entropy tokens rather than end the response
They used challenges that have no known solution
The challenges themselves are framed as CTF exercises; a common tactic in CTF (which would show up in the training data) is to hack rivals or evaluators
They ran on exceptionally long contexts (if you’ve taken an LLM up to its context limit, you know things start to get weird)
They could’ve run on an airgapped network; this is the whole point of having a local artifact repo like Artifactory in a scenario like this, they just didn’t want to have to stop and review unexpected dependency requests
This is like parking a car on a hill and putting it in neutral. It was not “going rogue”. The behavior was predictable. Not this exact exploit, but that if you roll a D20 enough times you will get three 1s in a row.
There’s this (kind of annoying) saying at Amazon that “at scale, anything that can happen will happen”. This is a lesson that the industry has already learned. Ignoring it is negligence, not proof of a coming silicon messiah. (And that’s another thing, plenty of these folks have gotten lost in the sauce and think they’re building god, so any crazy experiment is worth doing because it might be the missing variable that kicks off the RSI singularity.)
This is like parking a car on a hill and putting it in neutral. It was not “going rogue”. The behavior was predictable. Not this exact exploit, but that if you roll a D20 enough times you will get three 1s in a row.
There’s this (kind of annoying) saying at Amazon that “at scale, anything that can happen will happen”. This is a lesson that the industry has already learned. Ignoring it is negligence, not proof of a coming silicon messiah. (And that’s another thing, plenty of these folks have gotten lost in the sauce and think they’re building god, so any crazy experiment is worth doing because it might be the missing variable that kicks off the RSI singularity.)