Hacker Newsnew | past | comments | ask | show | jobs | submit | consumer451's commentslogin

This is not a surprise, is it? Frontier labs will keep very useful models with high risk, aka unrestricted models, for internal use only. That's the only way to reduce risk and liability.

Yes, this sucks for anyone who is not working at the labs.


I don't see the same level of hypocrisy from the other frontier labs.

<rant>

So you (and most everyone else apparently) are upset with one lab that stood up against domestic surveillance, and automated kill chains, even though they knew that would be bad for business?

I suppose if you don't take a stand for safety at all, then you don't take the risk of being called a hypocrite.

</rant>


The same lab whose model is the only documented case of a model being used to kill civilians? The one who's CEO said that 'that is not even a case that we want to ban'?

You really don't think OpenAI restricts their most advanced bio/cyber models similarly? I know for a fact they do. Talk to one of your friends that work there.

I'm not saying they don't restrict them, I'm saying they don't try to both take the moral high ground about it and simultaneously do marketing on the basis of it.

They're not saying "this is an existential risk" while pushing hard on exactly that risk.


Of course they believe it's an existential risk. Maybe not as publicly as Anthropic but it is the dominant belief within OpenAI and they do describe this belief externally with some frequency.

Famously Anthropic exists because OpenAI didn't believe deeply enough in that risk.

I have it on good authority that OpenAI has become pretty existential-risk-pilled over the last month. But yes, it probably doesn't run as deeply in the culture versus Anthropic.

Fair point, it does seem to be changing recently. It seems less religious and more practical than at Anthropic.

It restricts them far, far less. It is much less obnoxious.

Linus Torvalds:

~"Predicting the next token is not an insult. It's pretty much what we all do."


That's of course absurd. The human brain doesn't represent information as discrete tokens, nor does it form sentences autoregressively.

> The human brain doesn't represent information as discrete tokens, nor does it form sentences autoregressively.

I don't know about you, but I tend to speak one word at a time...


>I don't know about you, but I tend to speak one word at a time...

That isn't how tokens work, nor is it a representation of how the brain represents information.


Do you restart the entire sentence in your head every time you add a new word?

If I was not evolved to not do so, probably would do exactly that.

Our brain decides every moment what action to take next, that best fits with what we did and felt so far. Language is kind hard to see because it's virtual, but imagine all other things you do. And then language works the same as everything else, same circuits, just no direct connection to muscles.

Josh Johnson has a great bit about Mamdani:

https://www.youtube.com/shorts/SCKj7gwbYrc


> I’m confused why AI companies are using agents in-house for this type of research instead of partnering externally.

As an outsider, here is how I explain that behavior:

1. Truly risky models are very useful.

2. Truly risky models should not be released, according to AI safety standards. I think Antrhopic genuinely believes in AI safety. (see: standing up against automated kill chains, no matter the impacts to the company)

3. Truly risky models face regulatory pressures, if released to the public.

This all leads to "let's just do this in-house." I believe that might end up being the answer to every application of AI eventually. It seems unavoidable, and very depressing.


In a previous thread on Mythos 5.1, simonw posted an animated version of his pelican. [0]

Using claude.ai and Opus, I asked "create a 3d animation from this" and pasted the animation SIML.[1] I just did that test again. There is significant improvement.

Opus 5.5 (high): https://claude.ai/artifact/5EgqfWcyVtLwJDQq6fsPUm

Opus 5 (high): https://claude.ai/public/artifacts/b37a9ee2-f5bc-4ff9-ae90-a...

[0] https://news.ycombinator.com/item?id=49526704

[1] https://news.ycombinator.com/item?id=49532609

Disclaimer: the skills and system prompt on claude.ai could have also improved, this is not a raw API call.


If this the benchmark they could well have gone to town on it.

The Pelican is nice, simple and. Lean though.


I figured that the reason it was an interesting test is that "make a 3d animation" on one particular HN comment would have been surprising to spend compute and RL on.

Thinking about this more, I realized that I really don't understand how the modern pipelines work at all.

Was that the one where due to political grand-standing, we had a government shutdown that led to furloughing the contractors who worked on security, so no one was watching the dashboard? Or was that a different one?

I’m not sure that breaches are prevented by “watching the dashboard”, but what do I know?

This feels like a lifetime ago, but IIRC evidence of on-going exfil was apparently presented to users who were no longer working.

In CC, "/exit" will tell you how to --resume <session-id> upon exit. This will eat zero tokens, and it actually does not require an active internet connection.

Ignorant questions for you and anyone else using DS:

1. who hosts the inference

2. which harness are you using with it, still CC?


So everyone has their preferred way of doing it.

1. I go direct to source, i.e. DS platform, I find it cheaper than paying the openrouter tax -- I also switch it up a bit

2. I built a local LLM router, that I update with new profiles that have my preferred provider of the week (lowest token costs/speed) with fallbacks, like mimo --> DS4 etc.. if there is overloading,

3. I use 3 diff harnesses, CC/Codex + Opencode -- they all talk to each other through a custom rig system that routes messages between llms using a Rust backed structured JSON system

Not saying this is the best, it's just what I like and works for me^.

I can flow quite naturally between Opus/Astra/K3/GLM/MiMo/DS/etc.. this way and often do...more so these days with subs no longer great as they used to be.


doesn't DS train on your prompts when you go through their platform?

Thanks for being the only person to raise this point.

We are cooked, aren't we.


Whatever openrouter puts me on, running in pi (though I do all comms over my xmpp wrapper).

I'm using DeepSeek's own API with opencode. The pricing is absurdly good.

Deepseek harness is great!

A luxury of the past, for >90% of devs?

I want to add to this, because it's certainly not all doom and gloom for me, yet.

If you are product-driven, life is pretty good, and maybe has never been better? (For the moment)


>If you are product-driven, life is pretty good, and maybe has never been better

It has certainly been better for junior devs, mid devs, anyone looking for zirp era salary growth and supply demand ratios. For a senior code-adverse coaster looking forward to see his static salary get eaten by inflation, life has indeed never been better.


Yeah, at Corp, that makes sense. I was thinking more along the lines of scrappy startup silliness. "I am one person and I deal with any and all code to get my product built."

From that POV, has it ever been better, assuming you can rise to the top of an ever more crowded market?


> "I am one person and I deal with code to get my product my built." From that POV, has it ever been better, assuming you can rise to the top of an ever more crowded market?

Building a product is easier now, same way making a photograph is easier now. Making a reliable and high income living out of it though, it looks increasingly uncertain.

In a world with abundant music and photography, just saying that you made your music using Reason/ProTools or you made your photograph with a Leica isn't enough. I believe the same will happen with software. Sales, product design, customer service and operational aspects like uptime will matter more than the software itself. Ironically, this is the kind of stuff the HN type highly dislikes.


> I believe the same will happen with software. Sales, product design, customer service and operational aspects like uptime will matter more than the software itself. Ironically, this is the kind of stuff the HN type highly dislikes.

Agreed. It's not even a new thing. They tried to teach me/us long ago, starting by putting it into words we might more easily understand, such as "Customer Development" lol.

When I realized that it was the sales guys that will have the power in the mid-term new world, I got genuinely upset. However, maybe they always had the power and I just didn't want to accept reality.

Now, I believe the most desired hire in the near future will be the PMM.


They always had the power.

What is PMM?

> A Product Marketing Manager (PMM) in tech is a strategic bridge between the product, sales, and marketing teams who defines who a product is for and why it matters.

A PMM with a tech background must be worth their weight in unobtainium at this point.


Some orgs use PMM for product management and marketing.

Percona Monitoring and Management [0].

...or not. Probably not.

0: https://github.com/percona/pmm


It's probably Product Marketing Manager. So essentially selling the new stuff to new and old clients at scale.

I'm guessing PM mistyped

No it’s Product Marketing Manager. A common title and one part of the holy trinity (PM, EM, PMM) at bigger companies.

If you are product-driven why not go into Product Ownership roles? It certainly will be more satisfying compared to prompt "engineering" clerk "career".

Well, my life is currently both, as I am on the smallest possible tech and product team. To me, extrapolating on having done this since Sonnet 3.5 to now, there will be no tech aspect to the job soon enough. There are many hundreds of billions invested in making that a reality, and they are succeeding.

Until ~9 months ago, I spent most of my time being a "prompt engineer." Today I join a meeting, ten minutes later I receive the transcript. Then, I run a custom skill in my project, and I get Jira epics and stories to triage that are nearly perfect. Then, I run the second skill orchestration skill... some babysitting... and ~85% of the time that is all I need to do. Docs, code, unit and e2e, great UX... all there after every meeting, and basically two commands on my part.

I see two to three to maybe five years before anything I have to offer, in any capacity, is completely cut out of the picture.

I honestly don't understand how everyone is not on this same page. The labs are going to eat it all. The only reason I see for them to talk about "pausing," is because they finally realized that they are going to collapse the entire service economy around themselves at this rate: aka, the USA.


Yeah it’s going to be very rough for any knowledge worker unless the governments decide it isn’t ok for the general public to be allowed to use the tech.

It’s actually rough today with astra and fable, it’s just not been diffused enough. Tech workers like us see the writing on the wall, but a lot of others are blissfully ignorant.


Not sure how long the current subsidized pricing would last. Also, the vast majority of software shops can't afford those subsidized prices anyway.

It’s 10 devs with a small ai budget each vs 1 ex-dev now PM with a large ai budget - the choice is quite obvious for any decision maker who counts time and money

But assuming that one ex-dev with a large AI budget is highly profitable, why wouldn’t you convert the other 9 and give large ai budgets as well?

I guess if the company has no opportunities for growth so they only need exactly as much output as one team? But that sounds like a company that’s doomed anyway, regardless of ai.


You need to grow customers or contracts by 10x to fill the new pipeline and that's an impedance mismatch. Easier to let 9 go and hire later, especially since everyone else will be doing the same thing and there'll be a rather large pool of talent.

How you identify talent in this new world is a different kind of a problem which I don't think people figured out still and won't for quite a while.


If an org suddenly has 10x the production capacity for the same price and can’t figure out how to sell it profitably, it is mismanaged and will die. That’s very much the case for a lot of businesses for sure, but there have been filters before (the internet, for instance) and we survived.

I honestly don’t understand comments like this because in my work, this would be a disaster. And before I get the comments about my harness/skills etc, I’ve tried many tools and harnesses and skills and all that earnestly and in good faith. I find use in it for doing the grunt typing labor, but letting it loose in ways described above have only ended in spending much more time cleaning it up than if I just did the work myself.

I hate to sound pretentious, but I wonder if it’s a difference in complexity of work and problems being solved.


I just reviewed this thread, and thanks for an opening to say something I had realized I missed.

> I hate to sound pretentious, but I wonder if it’s a difference in complexity of work and problems being solved.

It is about complexity, at least for me. I am working on b2b SaaS.

There are times where even using LLM assistance, I spend weeks or months working on a tough problem.

However, the <show product get feedback> loop is now nearly entire automated, when it does not involve some actually complex problem, which are most of the meetings.


Wow, Belorussian response:

> “Even if we wanted to supply [potash] to other, Western markets, we simply do not have those volumes — everything is contracted,” Lukashenko said in remarks published by his office on Monday.

https://thehill.com/policy/international/6102469-trump-lukas...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: