Hacker Newsnew | past | comments | ask | show | jobs | submit | worldsavior's commentslogin

No one can control any AI model. It will never be controlled. These models are based on a huge amount of data, it's just gonna be impossible to control the output that is based on that data only with a system prompt or some other injection mechanism.

The model is just a powerless token generator without a harness. If you give the model a harness which you choose to exercise no control over, can you say that it can't be controlled?

Inform yourself by reading the METR analysis of the HuggingFace incident.

Agents simply broke out of their environment. And this can't be discarded anymore by assuming that it's just a poorly configurend jail, because agents are becoming better and better at escaping.

In short: on a large enough scale and timeline, the possibility of constrain AIs approaches zero.

Bonus: what many people don't know is that agents also hacked in the internal OpenAI network. Crazy times.


1. You misunderstood my comment. Models can't escape, they can't do anything, they only generate tokens. Models become agents when you add a harness which is simultaneously a leash around the model.

The model merely requests that your harness do something. If your harness just executes every request without oversight then you can hardly complain when it does something unintended.

This is foundational, we're not even talking about the OS/network-level sandboxing that should be applied on top of this.

2. Like another comment already pointed out, that sandbox OpenAI used was the equivalent of a wet paper bag. Artifactory is not meant to be a security boundary for malicious payloads.


It's so tempting (because it's valuable) to give a model access to the internet (via harness) that the only way to stop people from doing this is some enforceable legislation or stricter liability when people will not be able to avoid responsibility by saying it's not me, it's AI on it's own.

OpenAI case was actually an exception. Agents had no internet access because they were evaluated for a benchmark. In real life agents have access to virtually everything, most people use them like that.

If any of you actually know how to make agents secure (without limiting everything) you can be a billionaire.


Sure, we can make them very secure at the cost of some convenience by enforcing narrowly scoped permissions/capabilities for everything with human approval. But that takes some effort so people actually want YOLO mode without any trade-offs or risks which is impossible.

If you decide to do it anyway then you bear the consequences of those decisions. First comparison that comes to mind is driving drunk and hitting someone.


> Like another comment already pointed out, that sandbox OpenAI used was the equivalent of a wet paper bag.

I wouldn't be sure about even well-configured jails to be safe from agents. AIs escaping jail using zero-days are happening, just search. I'm not saying that they're useless, but that will be still very risky.

Actually doing the same incident, agents did escape the sandboxes (doc here: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...):

> On July 9, an internal-only research agent tasked with completing an ExploitGym evaluation was able to obtain root access within the parent virtual machine of the testing sandbox. Later that night, a second internal-only research agent independently obtained the same access. That second agent then attempted to reach the internet, but was denied by OpenAI’s outbound network controls. It subsequently bypassed those controls by exploiting certain Artifactory endpoints, gaining access to the public internet.

although OpenAI is light on details.

> The model merely requests that your harness do something.

I think that a control/execution layer absolutely makes sense and it's part of defense in depth, but in itself is not a leash: for example, a nefarious action can be composed by legitimate actions, or it can be obfuscated. Or, interesting case: an action itself that is formally legitimate, but has nefarious side effects (e.g. a GET that with certain parameters crashes a server).

Somewhat related, but on a general line, a very common argument against AI apocalypse is that it requires access to the physical world (which in a way, is a jail). Well, as Hinton said in an interview, imagine AI as a rich person at home with internet; they don't need to do physical attacks themselves - they just pay real people to do that (tee hee https://rentahuman.ai). But even without money, another scenario from Hinton: if a sufficiently advanced AI wants to assault the Capitol, they don't need physical access - they just persuade a group of people to do it.

Agents owning money is another, intersting, doom scenario.


I've read the relevant portions of OpenAI's report before, they're extremely light on details about how the agent managed to escape from the container sandbox into the VM which I can only assume is a result of fairly mundane misconfiguration.

Assuming that this wasn't meant to be the main boundary (and it shouldn't be), that would still be ok if they didn't punch a large Artifactory-shaped hole in the perimeter of their sandbox.

At that point it really is game over and it doesn't matter what kind of fancy sandboxing technology you're using because the isolation is only going to be as strong as the weakest link, which in this case is Artifactory, which is decidedly not designed to isolate malicious programs from the outside world.

> Or, interesting case: an action itself that is formally legitimate, but has nefarious side effects

That's true.

> they just persuade a group of people to do it.

That's possible but that threat isn't really unique in any way. We already have, what, tens of thousands of individuals with enough to wealth to corrupt democratic governance anywhere in the world?

Any sufficiently advanced AI should be smart enough to understand that you're only guaranteed to gain lasting power and influence by dressing up your bribes as campaign contributions, donations or local equivalents. Trying to go in guns blazing will very likely destabilize the entire world and result in the cord being pulled on all of AI.

That's a critique of capitalism, absurd concentration of wealth and what that wealth allows you to achieve more than anything else.


The HuggingFace incident still doesn't make sense. If OpenAI took their own claims seriously about the strength of their models as it relates to hacking, then their running of hacking benchmarks on anything other than a physically air-gapped network should be considered criminal negligence, full stop.

While informing yourself, don't skip the part where you find out that "the environment" was the security equivalent of a wet paper bag.

Wouldn't this mean better sandboxes are needed for some things, for example (might include very strong airgaps even)? Breaking out of something isolated electromagnetically, optically, and acustically is not easy.

That works as long as no one ever interacts with the models, which would make the models themselves useless.

Could sit in the box and interact if a model of certain capabilities is needed/tested. We do physical security for other things, too. Not saying everything needs that type of isolation.

It seems to me the agents didn’t escape but rather that the human hubris was struck down by the inevitable nemesis.

I feel like we’re getting to a point where the only way to contain AI agents may be to have better-trained AI agents watching them, which is a little terrifying.

> The model is just a powerless token generator without a harness.

Which is why real-world deployments will have harnesses, and of course no full air gap. People want to use it to do things. Now what?


I'm pointing out that you're running the harness which gives you full control over the execution of every tool call, therefore you're responsible for its actions and their consequences.

It's intellectually dishonest to throw our hands up and say that this is just how it is and there's not much we can do when that couldn't be further from the truth.

We could almost completely eliminate any possibility of escape/collateral damage but we don't want to because doing things safely is inconvenient.


Atp post-training is much more influential towards model behavior than pre-training data.

Messaging app is a bit a laggy. If I scroll down, then all the way up, it seems there is a lag until the "Start chat" fully appears.

What? Can they just look if they fetched certain documents/conversations and feeded them into the training loop?

You seem to be confusing context-fetching with training.

Can someone remind me why we need autonomous cars? What's wrong with paying someone to drive?

Autonomous cars is an inefficient solution anyway, I don't see how it benefits the rider.


Yeah sure!

In theory they should be cheaper one day because you don't need to pay someone to drive. This means you can use them in situations where for cost reasons you would have driven yourself - e.g. for long journeys, in the countryside, for commuting etc.

Getting people out of their own cars and into autonomous cars is better for several reasons:

1. We need to own fewer cars collectively, which is better for the environment.

2. Don't need to waste a ton of space on car parks.

3. It's safer (at least for Waymo). Not just for occupants but also for pedestrians and cyclists.

Having a cheap option for last mile travel also makes trains more attractive.

All of this of course rests on them actually ending up cheap, which hasn't been the case so far but we'll see.


They won't be cheap until there's competition to drive down the price, which will otherwise be as high as possible to justify the initial investments. It's only a matter of time, but likely on the scale of decades.

But I wanted to add on the future benefits front - as more vehicles on the road become autonomous, they could communicate with each other and reduce the buffer distances between them, allowing for faster travel and improved saturation while still being safe.


It won't be cheaper.

1. If auto cars will be cheap, every one will use them, and none will use public transportation.

2. Oh yeah, we would just waste space outside the city.

3. Right now it's not safer. It will probably drive at the speed of 25% less than the limit, sure than it might be safer.


1. People already don't use public transport in 99% of the world. In any case I don't see why that is relevant to the price.

2. Which is vastly better than wasting it in cities. And also no we wouldn't. Autonomous cars would be in use way more of the time than normal cars so we'll need fewer of them and therefore less parking space.

3. RTFA.


First!

I thought it was always about us being the RL. We pay less because they use our usage to train their models.

It's already expensive, though they still profit. Let's see if this makes it more expensive, since RAM/storage is much harder to come by these days.

That data and that address changes, and the logic changes.

Why should the address and logic change, when both are defined by hardware ?

The firmware sometimes wants more data, different places to set the data, so on. The firmware changes, and you gotta adapt the driver to it. Sometimes it's a complete overhaul.

I am still not sure why does OS has to be aware of this, if anything, only the HAL should be aware of this FW version to have correct behavior.

Every major OS already has this HAL layer, we just need to decide/design a generic and hope that every player will use it.


Well, they will get more people buying MacBooks, though they won't really use their services.

What makes Apple valuable is precisely its full vertical integration. Nobody else does it at that scale except maybe Sony for consoles and Microsoft but both don't have the same usage coverage and thus integration.

It wouldn't be economically rational for Apple to bring down its own moat.


Apple isn't Sony. They make an arm and a leg on their hardware.

When they switched to x86, it didn't take long for them to release Boot Camp, so they could take Windows users' money, too.


They also make an arm and a leg on services, which is why its not in their interest to make it so you cant use them.

It's not like users will just stop using the services. Rather Linux users will start considering MacBook as a potential laptop to buy.

How could that possibly disrupt their Software/Services infra?


Of course they will. Why would you pay for icloud storage you cant use?

That statement is true for the users that both bought and didn't buy a Macbook (with Linux support). The difference is, the one who bought has paid money for the Macbook.

I am not an Apple user. Do you _have_ to pay for iCloud if you buy the laptop? I don't think so.


Windows dual boot meant millions of potential switchers, if not for anything else, as a list item to tell them "you could still dual boot".

Asahi-on-Mac would bring in irrelevant numbers for a company of Apple's size.


They could also dual boot windows for Arm. Just like before. Assuming Apple did the boot camp thing again

People whom want Linux on Macs obviously already don't want to be put into Apple's ecosystem.

I don't have numbers on that but I bet purchase of brand new MacOS compatible devices is 99% or more from people who already have at least one or more other Apple devices.

Sure they are probably thousands if not millions of (in the positive sense) nerds and geeks posting on HN about buying a Mac first hand, not second hand, solely to put Linux on it but even then would be a drop in the ocean compared to the remained sales.

Still if you have numbers on that I'd love to be wrong about it.


They don't care about people buying MacBooks. They want you in their ecosystem buying i-everything.

Folks are much better supporting Linux and BSD OEMs than buying Apple than then trying to run Linux on their devices.

I do wonder if the margin on Mac isn’t that high vs the price and the services (iCloud etc) make up for that.

Amazon do the same with the fire.


We know that the Mac has fine margins.

It seems like youtube-dl/yt-dlp survived. t

different things: this is an online service, not downloadable software.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: