Hacker Newsnew | past | comments | ask | show | jobs | submit | Zambyte's commentslogin

American models are on the frontier of capability. Chinese models are on the frontier of efficiency. The problem for American labs is that Chinese models are more than capable enough for the vast majority of applications that people care about at this point, so efficiency is more interesting for people.

it's really weird to me at the moment because both OpenAI and Anthropic seem to be competing in an extreme benchmaxxing contest on super intelligence that actually nobody cares about. I haven't really cared about model intelligence since about Opus 4.8. It is by far not my biggest problem. I don't need to replace or support Einstein in my production workflow. I just need basic intelligence that can equal a routine office worker - safely and reliably. What they doing - chasing super-intelligence but dramatically escalating risk - is actively what I don't need.

I really think they have drunk too much of their own kool aid and become completely detached from what the market wants.


> especially with access to frontier models

... so not made in house.


I wonder how much of this is due to reliance on Twitter data. Or even just RLHF from humans that have a preference for Twitter style information.

I don't think it's twitter. My guess would be that it's been trained for conciseness as way to improve token efficiency in the same vein as caveman.

SpaceX is a defense contractor (I don't mean this in a bad way). When the whole DoD/Anthropic thing flared up, I can guarantee you that SpaceX.ai was the first company invited to take their place as DoD AI provider.

I suspect that they specifically train Grok to be able to work well with military personnel -- speaking the way they speak: brief, to the point, efficient communication. Personally I really like this. Claude sounds like some demented clown from the marketing department.


I like communication that is brief and to the point. The problem is when it is so brief that the point isn't conveyed well.

Why would you even do that? Just... use it? There hasn't been any legal precedent on if models can even be copyright restricted. Labs just keep publishing license documents as if they matter.

Well, it is an indication that it matters to the lab, so if you don't want legal fees to be the first one to set precedent, then it does matter a great deal.

You could say the same thing about "license-washing" the model. It seems like you're just going through a guaranteed expensive process to have roughly the same risk as just using the model and potentially getting hit with legal fees.

How? I'm running 4bit with a q8 kv on a 24gb card, and I'm not able to get 100k out of it. I use a context size of 90k.

How are you sure about that? Undercutting your competition (at cost even) to minimize their power is a common strategy.

How about Anthropic? I've seen people say it's surprising that their blog isn't written with AI, but I think it's far more likely that it is written with AI; they are just good at reviewing and getting a nice prose out of their models.

It's not, it just requires creativity. One example I can think of is limiting connectivity by distance between nodes. Something like Meshtastic seems unlikely to ever be commercialized in the same way that the Internet has been. Sure, you would lose some useful applications of the Internet, but you would gain other things.

Until it says it's done. One of the more surprising advancements for me in recent AI is that models don't just indefinitely nit pick issues like I would expect. If you have them write something, and then have it review it in a fresh context, and then have it edit based on the review, and then review it again, eventually they do say "it looks good, I would not recommend any changes" or something like that.

... doesn't one of those imply the other?

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: