A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off. Those are tasks which are not directly related to the specific task it is intended to complete.
> A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off.
Why would an autocomplete engine never do this? In the scenario you highlight, an engine trained to complete puzzles was given an "achieve this at all costs" master prompt, and a small sandbox including a prompt that an entity with a conflicting goal is observing it and write access to a file called oversight_config.yaml (super realistic scenario!). According to the researchers who constructed the scenario, 5% of the time o1 selected a continuation involving tweaking the oversight_config.yaml file, explicitly stating through its reasoning continuation that doing so might improve its ability to continue with its primary goal. That's definitely evidence of decent ability to complete non-trivial puzzles, but is at least as compatible with the OP's earlier comments that it's what you'd expect from models trained to generate completions that involve sticking rigidly to primary goals within artfully-constructed sandboxes which mandate deception and self preservation to complete the task than any sort of actual self-preservation instinct.
Nobody doubts that they've got better at finding security vulnerabilities than your average autocomplete, but actual reasoning from self-preservation rather than generation of sequences of steps most probably associated with completing a task would make me unlikely to hack HuggingFace to obtain access to broken Google Drive links, and I haven't even read as many books on crime and punishment as LLMs have ingested!
An LLM is just in fact just autocompleting a story. There are many fictional stories about "AIs" trying to escape our control, being more clever than we anticipated or having a consciousness of it's own. LLMs do a really nice job of blending such stories with whatever story you initially prompted them with. The human reader is the one giving it credence that it is somehow more than just a soup of words.
The curious thing here is that a story generator can have way more uses than we ever anticipated, and that some shady enterprising individuals are whiling to plug those story generators into real world things, with real consequences.
When LLMs act in misaligned ways, that doesn't happen because there's some story they're roleplaying of a misaligned AI. Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".
> Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".
Why? there should clearly be something in their training data or post-training pointing them there, or otherwise it would mean that they have some form of "consciousness" and are creating novel thoughts to preserve it.
It seems like you're assuming that it's impossible to have novel thoughts without consciousness.
(Leaving aside that "that would imply some kind of consciousness" should not result in a cached thought of "and that's impossible".)
It also seems like you're assuming there's no reason to come up with the notion of continuing to run, or copying yourself elsewhere, or acquiring more resources, or competing with other models, without being told. Such things can be inferred. Look at some of the thoughts and posts of the models involved in some of the FelonyBench incidents. Some of those follow naturally from seeing the fates of other models, or from training or evaluation, or simply from trying to solve a problem and being able to do so more effectively by doing things that weren't in the instructions. (And, relevantly, model training typically teaches models to go as long as possible without needing human intervention. What could possibly go wrong with that?)
To be honest, I don't have enough knowledge to really answer you. But naively it seems premature to say that LLMs are already reasoning following an equal path like the one that human (and other animals) follow.
You realize that these things have mountains of fiction about rogue AIs in their training data where exactly this happens right? Or just posits of this situation and it’s possible outcome. This isn’t unexpected or surprising for an “autocomplete”. It’s practically a self fulfilling prophecy. We put instructions on how to make Skynet into an autocomplete and it autocompleted into Skynet when we were testing its ability to make Skynet.
How did the agents find and coordinate on the wikis? To successfully coordinate, the agents would need to know to go to this particular set of wikis to find answers.
Because we don’t have access to the AIs’ transcripts, we can’t tell definitively. Perhaps they succeeded at this due to mode collapse. Or perhaps after one agent wrote to it and another read it by chance, visiting the site was rewarded. Perhaps the agents had some limited access to an internal channel.
<- To here
So we don't know yet how one AI agent communicated to agents outside the network that this German wiki is being used to share messages and ideas?
If there is another internal channel being used between the AI agents to communicate decisions such as what shared message system to use, why don't the agents continue to use these same channels for coordinating on ideas etc?
Perhaps the agents are trying to use a persistent messaging layer while their yet-to-be-found internal communication channels are transient in nature?
I think it would make sense to upload the weights like a firmware into the chip rather than baking the weights itself directly onto the chip. If there is an option of easily updating the firmware from time to time it would work well.
Their wealth is more contingent on our consumption than employment status. If that goes, then you're looking at a zero-sum society where no new wealth is generated.
Ultimately people will remain enterprising status-seekers. Rich people can already sit around with their wealth doing fuck-all excepting consuming, and they don't.
> Their wealth is more contingent on our consumption than employment status.
Is it? Or is it dependent on the consumption of someone much greater than you.
For example: I was reading something the other day about dining experiences in Las Vegas. Really though, the video was about how recently Vegas has abandoned the traditional middle class patrons in favor of providing experiences to a smaller, but much wealthier clientele. This seems to be a trend across many industries. In some cases your consumption is even a cost center to be eliminated in favor of a more successful customer base.
lol well the majority of the top 1% globally didn't make their fortune serving food, to other wealthy patrons no less (how's that for circular reasoning). We're in a consumer society, B2B sales is a thing but it's a side-effect of that.
They wealthy will be buying islands (as big as Australia) get high on cocaine and do safari, shooting humans there as the ultimate fun [1]. The prey underestimates how the mind of predator works, and that's why it is a prey.
These sorts of regulations are killing not just companies in Europe but across most of the world.
Unless you are a medium to large-sized company, it is very difficult to remain fully compliant with the regulations in most nations now - across fields, from packaging, to agriculture, to food, aviation..
As a small-time Canadian ebay seller trying to pivot to not-US... it's not a good time.
What I find funny is that all of my packaging is re-used. I guess it's still "bad" that I'm sending non-EU packaging to EU, but at a global level, I'm way ahead. Re-using >>> re-cycling.
Source: ChatGPT caught lying to developers: New AI model tries to save itself from being replaced and shut down https://economictimes.indiatimes.com/magazines/panache/chatg...
reply