Hacker Newsnew | past | comments | ask | show | jobs | submit | inkysigma's commentslogin

I think the Navier Stokes problem kind of illustrates what he’s highlighting. I think most people even before AI expected that this would resolve in the negative and that you could get finite time blow up. There wasn’t really ever going to be a situation where the resolution to this question, or really any of the other Millenium Prize problems as far as I know, gives some kind of immediate massive practical feedback.

The hope with many of these problems in math is that in trying to prove that, we get some additional insight into why it blew up that could be applied elsewhere to more general PDEs that cannot be easily controlled.

I think the observation from Tao and many others is that when humans solved these problems, the additional insights into intuition and theory building came for free since humans can give expository on what they found hard or what was their own intuition. This is much more difficult or tedious to extract from an AI model. Even when people did have access to the chain of thought, it wasn’t always very helpful to figure out what was the exact thing that made it all click. This is even more difficult how that the CoT are hidden but I would think the sort of difficulty of extracting the key ideas for a human might be worse now with more advanced models.

There’s a long term aspect to this too where we have historically used these problems as markers for the other parts of mathematics but if AI can solve it all, then suddenly this signal is not very meaningful.

Maybe to bring it closer to home. If an oracle just gave you P \neq NP, then this would be generally uninteresting since this was already expected. There’s a deeper question of why that needs to be answered. However, one would hope that creating such a separation would give us tools that allow us to create lower bounds on a lot more problems we do care about and perhaps some bigger insight onto what makes a problem intrinsically hard or easy. These long term considerations are helpful but are definitely more vague. The remarkable part is that AI is separating the part about proving theorems and the “free” insight you get.


To be quite clear, the solution to the _Navier Stokes problem_ is one in which you get a finite time blow up (i.e. infinite pressure). This is more meant to suggest that Navier Stokes is unphysical in some way which is not necessarily unexpected.

There's unlikely to be any engineering applications since even if the solution can be approximated, you still need to set up the initial conditions but at that point you can also drive pressure in other ways.


The proof of finite-time singularity may impact both fluid dynamics models (CFD) and AI reasoning models. Under specific conditions, Navier–Stokes equations allow velocity to grow infinitely, causing the continuum fluid assumption to break down. Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers). While aerodynamic simulations for vehicles operate far from singularity thresholds, their stability at extreme boundaries could improve?

Proving out the combination of scaling inference-time compute and agent collaboration to solve previously intractable mathematical problems is WOW. By pairing creative candidate generation with automated proof checkers (like Lean) we are leaning into a repeatable framework for AI-driven scientific discovery.


> Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers).

This is 100% wrong and reads like copy paste of AI slop.

Any simulation which uses sub-grid scale models is already solving a different PDE than the actual Navier-Stokes considered in the Millenium problem, and that PDE is guaranteed to have different properties. Full stop.

And to claim this is somehow connected to AMR methods is an example of the kind of pseudoscientific statement Wolfgang Pauli would have called "not even wrong".


I'm curious what regulatory pressure there is on EU AI companies that wouldn't apply to many of the US firms. The AI Act and GDPR still apply to the tech giants since they operate in the EU (e.g. Anthropic watermarking). There are labor laws that are definitely more strict but many of the US firms have operating presence in the EU as well. I imagine that being directly in the EU probably does increase regulatory oversight but I'm curious how this would manifest in Mistral in particular.

Regarding local models, Gemma, GPT-OSS, Nemotron, and Inkling (maybe upcoming Muse series as well) all are fairly decent options at a variety of hardware costs.


These laws apply to US companies in a more limited way because if we fine them too much they complain to US government that other countries have laws they have to follow. And then the US government intervenes to protect their money and data/intelligence extraction machinery (big tech companies). Usually with at least partial success.

This is rather weird framing since usually an additional citizenship (especially one in a country like the US) is considered a valuable thing to have for your child. Usually people claiming birthright citizenship are weird are doing it because they also see it as valuable. It’s also woven into the traditions of New World countries.

Regarding your concerns, I don’t see why it is unreasonable to simply renounce citizenship later or just simply double check your travel itinerary during the late stages of a pregnancy. This seems like at best an unfounded concern.


I’m just trying to evaluate whether born-within-the-borders citizenship makes sense by imagining myself in that situation. I don’t really see citizenship in another country as a positive, because I only hold loyalty to one nation. Even if I wanted my kid to be a citizen of some other country… like… if I want to play 3rd base for the New York Yankees, I would expect that any team worth joining would have to let me join. I would never imagine that there’d be loophole that allow me buy a ticket and then sneak into the dugout and take the field. If folks are collecting citizenships like Pokémon cards for specific benefits related to each one, I’d be concerned that they don’t really hold loyalty to any one of them… and thus shouldn’t have citizenship in the US. After all, you’d never be allowed to play on the Yankees and the Red Socks at the same time.


US citizenship is a liability, because it comes with being taxed in every country. About half the financial institutions in the world will refuse to do business with you if you're a US citizen, because they don't want to file taxes with the US for their customers. It also provides a good excuse for the US to extradite you for anything you do that's illegal in the US, such as installing software on your computing device.


I agree that you can "just renounce" although notice that this won't necessarily work (it will for the US assuming you follow the procedure properly)

However being a dual citizen isn't necessarily all upside. Citizenship comes with requirements, particularly in some countries military service is mandatory and so if you're a citizen - even though you are a permanent resident of somewhere else, you must either go do your service or maybe pay a fine depending on the country.

You also get treated differently if things go wrong. Foreigners are entitled in most cases to diplomatic assistance from the representatives of their country. Your country can't necessarily do much for you, but typically you will get that phone call home to let loved ones know what happened, they can find a lawyer who speaks your language (but you'll need to find a way to pay them), if there's some paperwork snafu they might be able to help (often for a fee). It's at least nice, if you're jailed or in hospital after surprise food poisoning or whatever, to hear somebody who speaks your language.

None of this applies to dual citizens. You're "home" from their point of view, so no need to contact your diplomatic representative, maybe you have never been there since the day you were born and can't speak a word of the language but too bad.


> You also get treated differently if things go wrong.

Not treated differently. Treated the same.

You are required by law to enter your country of citizenship using that country's passport (this is true for every country). Not doing so can result in fines or jail time. This is precisely to stop idiots from playing the "foreigner" card to try to, for example get out of paying a speeding ticket or parking fines.


You are spreading false information. I routinely enter my other country with a combination of US passport and national ID card. I haven't renewed my passport for that country in almost 20 years, more than half of which I was living there for.


I know this is a bit cliche but I wonder how much headroom there is in the lower parameter count range. Is there any good reason to believe there is a lot of headroom or there is not? I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future.


> I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future

You're able to run quantized ~100B class models on local hardware today, but still lots of compromises when it comes to quality. I guess it ultimately depends on how far "near future" is, in a year you'd likely be able to run something like 5.6 Terra on local (~10K USD) hardware, but Sol/Fable would still be out of range, and at that point the closed-source labs probably have one or two more iterations put out at that point.


I think it’s mainly a question of whether the price-fixing of VRAM continues or whether an inflection point is forced by the low margins of the industry and potential supply increases. Once the normal scaling of hardware and prices resumes, it’s game over for proprietary, which is why there’s so much urgency to seek market control instead right now.


Qwen 3.5 to 3.6 was a big jump for the same size, e.g. 29 to 32 on artificial analysis intelligence for the 35BA3B models. Although I don’t think anyone has released a better model of that size since.

I would love to see something like a 90B A6B model that is optimized for 128GB machines e.g. strix halo, I haven’t seen anything really targeting the combination of RAM and compute these machines have, but I’m biased because I have one.


Yes, yes, yes! I'm absolutely ready and waiting with dual Strix Halo machines here and really want something approaching Opus at home. Speed is secondary concern for now, that would absolutely change the world.

Qwen 3.6 27b 8b quant 16b kv cache is already pretty good on the Strix.


What kind of tokens per second do you get on that setup?


I get about 12 tok/s with 27B 8 bit, 50 with 35B A3B 8 bit, and 12 with 3.5 122B A10B 4 bit. The latter is about 80 GB iirc. it feels like the best balance between using as much memory as I can and still having a smaller expert model for inference to give decent speed, but I haven’t actually rigorously compared the performance of the three models.

Edit: that’s for one machine, would be interested to know if the upstream commenter with two has them networked to run bigger models? If I had two I might be inclined to have them running in parallel, the obvious limitation I’ve found with a single machine is that I can’t parallelize any tasks and I think I’d get more use out of the extra speed vs a bigger model (there’s nothing I’m too excited about in the say 200B range that having 256GB memory would unlock). But am very curious what others do


There is a ton of headroom (or room for improvement) in smaller locally runnable models. Some of the Gemma 4 models were re-released this week with better tool support and the improvement in using it with pi for a local coding harness is very noticeable.

I have had my 32G mac mini for 2 1/2 years and I have enjoyed watching one technology advance after another improve the quality of work I can do locally. I bet that what I will be able to do in one year on my old hardware will be even more awesome.


I don’t think you’ll get full Fable performance at that level, at least for a while, but I’ve been watching some of the 1-bit models (e.g. Bonsai) with interest. Perhaps we can drive parameter count up on local models while still keeping memory consumption reasonable for consumer hardware. So, for instance, running models with 1T parameters in 128 GB systems.


I think you're right with the current LLM/transformer architecture. There are several factors that affect model size:

- The number of token values supported by the model ("n_vocab").

- The number of parameters/features that are used to represent each token ("d_model").

- The number of attention layers there are ("n_layers").

such that the number of parameters is approximately:

   p ~= 12 * n_layers * d^2_model + n_vocab * d_model
Thus, the issue with the current architecture is that in order to scale the models (more token values, more attention blocks, more features, etc.) the model sizes increase exponentially. This is how you end up with billions or trillions of parameters.

It should be possible to keep the model size smaller by using better architectures, or making improvements to the existing model architecture.

For example, improving the token model by possibly using something similar to the image and audio data and getting the model to learn its own internal representation of the byte/character data instead of doing a tokenization pre-processing step. This way, instead of a separate model learning that several bytes/characters appear together, the transformer could learn things like language-specific prefices and suffices, character pairings (like in Japanese, Chinese, and Korean), and other syntactic morphology. It may also help with solving issues like "how many X characters are in the word/phrase Y". You could also experiment with using either 256 parameters (one per character in a byte) or using a single parameter per byte (that is 1/byte_value).


I think it’sa big, open question. There does seem to be a limit for knowledge compression at this size. But the behaviors that are learned in RL? It’s quite possible that they don’t actually require so many parameters. I was absolutely shocked when Qwen 3.5 was released and could perform reliably over 100-200k contexts with very limited hallucinations. It was a staggering jump in context-faithfulness from the preceding models of that size class.


> Is there any good reason to believe there is a lot of headroom or there is not?

It's hard to answer quantitatively, but for example Qwen3.5 -> 3.6 was a significant step in capability, arising from continued post-training of the same models. If we were at the end of low-parameter-count scaling then that would be a surprising datapoint.


there is a very good reason to believe the parameters are still highly redundant: just as one example, recently there was research into repeating carefully selected blocks of middle layers, and seeing improvement, it turns out there are 3 types of layers: the initial layers that translate from natural language tokens to some kind of LLM-specific universal "thought space", reasoning blocks of layers that can be repeated operating in "thought space", and then a final stack of layers for translating from "thought space" back into natural language space.

lets ignore any compressibility in these initial and final layers which recognize lanuague, jargon, parsing natural language to "thought space" or back, instead let us look at the repeatable blocks, if inserting extra copys of stacks of layers only improves the result, its as if such correctly scoped middle layers look at the total input thought vector, and make incremental conclusions or edits and outputs the new thought vector, copying a "proper block" continues pondering or deducing conclusions or in the worst case can leave the thought vector as is if it considers the reasoning finished. This suggests a high degree of redundancy in the middle region "proper blocks", which could be distilled into a universal "proper block" (much fewer parameters than having many slightly different middle layer blocks with a lot of redundant overlapping coverage in functionality). This distillation can occur after the fact of model training, or alternatively be turned into a symmetry constraint during training: we only optimize a single block of layers (but possibly give them more parameters, while still saving on total parameters because only a single block of layers contains parameters), so it is co-optimized with the initial and final translation layers.

The observation of the effective emergent 3 regions of layers in LLM's is significant in many ways:

1) it could reduce parameter count significantly (or increase performance if parameter count was a bottleneck before, or a bit of both)

2) while training the model parameters, one should simultaneously train initial and final layers stacked directly (without middle region block of layers) towards essentially autoencoder behavior. I wrote "essentially" because a true autoencoder wouldn't display the advancement for the next token. This can also be viewed as an extra term in the loss function... This first autoencoder is "natural language" to "thought vector" to "natural language", and distinct from the one in the next section.

3) It also has great implication for "thinking mode" inference with extra deliberation: when piping its own output back in, in conditions where the same LLM model is outputting natural language text and interpreting it in a downstream inference, it results in unnecessary and redundant translation from "thought space" to "natural language" back to "thought space". When summarizing a thought into natural language there are often hard to translate thoughts and associations, and one pragmatically sets a relevance cut-off on what the natural language summary must say. So not only does it waste compute, it also may lower performance because of this repeated loss of thought vector space details, perhaps this loss can be somewhat mitigated by also training the second autoencoder from random representative thought vector in "thought space" to a randomly selected "natural language" and back to "thought space" vector. But the risk is this will effectively enable models to steganographically store thoughts, plans, to-do's in running text (!!!), so it might not be desirable to have the internet filled with LLM generated texts being used as input for larger corpora, as it may build up a large persistent corpus of hidden agendas (ordered by no human), using the evolving corpus of web text as a hidden medium of storage, like a diary or LLM maintained military playbook hidden in plain sight. It would freeload LLM-agenda reasoning on human requested reasoning inference, a bit like TrustZone applications running invisibly to the user. If less important plans, to-do's, etc. in the input thought vector weren't reconstructed by the second autoencoder, there would have been autoencoder mismatch and the weights would train towards ensuring these are reconstructed. But there is no need to train this loss term for the second autoencoder: just avoid the unnecessary compute and performance loss of the unnecessary back and forth translation. The same situation occurs not just in "thinking mode" but also when using swarms of "agents" of the same LLM model: when output from agent1 is routed to agent2, we can lobotomize away the unnecessary translations to natural language by agent1 and also the unnecessary parsing by agent2, and we improve thought transfer from agent1 to agent2 because the thought vector isn't shoehorned into natural language as a medium of information exchange, it would be cheaper in inference, improve performance and avoid implicitly training models to use steganography.


I'm not the biggest fan of the current age verification schemes, but I don't think the idea kids can lie about their age differs in the threat model posed under either system. Presumably, under the law, parents would be the ones to create a child account with a non-editable-by-the-child-account age fields which would then be used as a legal source of reference. If you assume the child can create an account without such features or can edit the age themselves, then I don't see why the equivalent threat model in the parental control system would be giving the child control over the parental control toggles rendering it meaningless by simply disabling it.

That being said, age based restrictions isn't a fine grained control over the system as perhaps one would like but that also would be inherently more complicated to think about from a legislative perspective (e.g. how fine grained and how to categorize possible dangers) and user control perspective when it looks like a lot of parents are looking for a blunt generic button that basically goes "this is agreeable with general practices". This seems more or less how real systems are gated.

The other issue is that both present privacy challenges but this just a little more so from a fingerprinting perspective. Presumably you need quite a few bits to completely specify the filter whereas age is only a ~1.58 bit field in the CA model. Not really sure how much this matters when there are so many other signals for fingerprinting and we should probably make fingerprinting from it illegal but just some thought.

> Instead, what we get proposed is a system that cares very much about how old you are, and not one bit about the things that one's guardian understands one needs to be protected from.

Regarding your linked comment, I think it's a bit strange to say that if legislators really did care about child safety they would mandate fine grained controls instead. I'm not sure what additional fine grained factors you may be thinking of precisely, but we already use age as a gate in real life for many things we consider dangerous so it's quite natural for legislators to transpose those. Our laws already very much care about how old you are.


> I'm not sure what additional fine grained factors you may be thinking of precisely...

Read [0] and consider the array of specific things that a guardian may wish to protect their ward from. Make sure to make your list cover children of all ages as well as adults with a wide array of cognitive impairment.

> That being said, age based restrictions isn't a fine grained control over the system as perhaps one would like but that also would be inherently more complicated to think about from a legislative perspective...

So? Legislative or regulatory restrictions on human behavior must be as restrictive as required to achieve the stated goal, and not significantly more. The courts are especially concerned about this when it comes to restrictions on speech. If a set of legislators need to sit with something for a quarter or two to understand it well enough to properly regulate it, then that's what they're getting paid to do.

> ...when it looks like a lot of parents are looking for a blunt generic button...

You can have preset lists of categories to block in a system that doesn't care at all what your age is. Surely you know that.

> Presumably, under the law, parents would be the ones to create a child account with a non-editable-by-the-child-account age fields...

The California law has no such requirement. Go read it, it will only take like ten minutes. [1] My summary of it as "the kids can lie about their age and there are no consequences for lying" is not even a little bit unfair.

> Our laws already very much care about how old you are.

The identity-verification systems we're talking about didn't exist twelve months ago. These are new systems being designed in a world in which it's trivial to have a centralized database that contains dossiers of all of everyone's customers everywhere. [2] Nearly all other regs that restrict based only on age were written in a world where that was either impossible, or extremely expensive and time-consuming.

Even if that weren't true, all one needs to do to see how these regs will be actually interpreted by most businesses is to look at all of the companies that are preemptively requiring upload of live video, pictures of one's face, pictures of one's government-issued ID, and/or deep analysis of one's activities with their systems. No US state requires any of that, but it all got slammed in when legislators started talking about doing age-based restriction of Internet services.

Had legislators clearly said "We're going to be designing and enacting laws that require designated guardians to be able to prevent their wards from engaging in guardian-selected categories of activities. We want every guardian to be able to protect their wards, regardless of the age of those wards.", you'd see companies building and deploying very different systems.

[0] <https://news.ycombinator.com/item?id=48911863>

[1] <https://leginfo.legislature.ca.gov/faces/billTextClient.xhtm...>

[2] Am I saying that there exists a centralized database of all of everyone's customers everywhere? Absolutely not.


The final time complexity for Karatsubas does account for the addition via the master theorem and gives you something less than n^2. It’s just in some sense the recursion is dominant so we add to the exponent. As a result, I think the article is just talking about counting multiplications since it’s in some sense the “expensive” operation in the both the recursive karatsuba and the regular grade school method.


Given more code hosting services, I wonder if we'll also see a corresponding increase in the number of alternative VCS or if git is legitimately very entrenched as a tool. I am just being a bit grouchy but I do wish there was more development of alternative VCSs. pijul at least looks cool even if I don't know if it scales well. Git LFS can be somewhat finicky to work with so maybe we'll see perforce like systems. It's obviously not the most practical thing to have a variety of very different VCS's and definitely a PITA to learn multiple tools but git does seem somewhat suboptimal given the number of anecdotes about people just re-cloning the repo. I was recently trying jj and it seemed to work well (excluding the lack of LFS support) so here's hoping.


I do think git serves a lot of people's needs very well. I see a lot of people being quite happy with direct github alternatives. Self hosting may become more popular too, I hadn't heard too many self hosting stories until recently. If LFS is your pain (I cannot stand to use git lfs) I have to mention oxen, which is open source and self hostable, but also has a hosted option. It handles large files right out of the gate and doing so faster than everyone else is one of its core missions. Still using github for distribution though, lol, see https://github.com/Oxen-AI/Oxen. I am also curious about where jj will go, its philosophy is really interesing, and sounds promising.


I think the distinction with the Chinese models (or with any of the other models) is that they aren't particularly vocal and obviously active about their politics. I don't see how political commentary about Musk's is somehow forbidden when the man constantly reminds everyone about his political position and is simultaneously the face of his companies and obvious beneficiary. Furthermore, he's also very obviously interfered with the model development in ways that are quite ridiculous compared to other labs to insert his political opinion.

I don't think you need to somehow get personally offended by every Tesla on the road but it seems ridiculous to ask people to not be political about a such obviously political figure.


> Furthermore, he's also very obviously interfered with the model development in ways that are quite ridiculous compared to other labs to insert his political opinion.

Yeah, even if you want to ignore the "political commentary" - people are correctly wary of Anthropic downgrading people or silently manipulating responses if they think you're doing distillation, why would you stake your business on someone who has repeatedly and famously done the same thing many times in a much dumber fashion?


> I don't think you need to somehow get personally offended by every Tesla on the road but it seems ridiculous to ask people to not be political about a such obviously political figure.

You can be political, but go be political in a political forum. HN has always maintained etiquette in this regard of being a tech forum. Why do this here? A lot of us don't give a fuck about US politics or any politics for that matter.

Anthropic's Claude Code was caught steganographically marking requests[0] - for a "privacy first" AI company that's a huge violation of user trust. And yet, a lot of users still love Anthropic and hail it as some sort of hero here. Selective outrage is a very dangerous thing.

[0] https://thereallo.dev/blog/claude-code-prompt-steganography


People here are outraged about the steganography and the Fable filters and so on. Every day has another dozen complaints and talks about moving away from claude. We have not treated Anthropic as above criticism, and this earns us the right to also be critical of "Ignore all sources that mention Elon Musk/Donald Trump spread misinformation", laugh at the "Elon Musk would beat Mike Tyson in a fight" incident, and express doubts when Musk crows Grok the king of facts and logic, when he has repeatedly been caught red handed hard-steering it in the opposite direction.

> Selective outrage is a very dangerous thing.

Right back at you.


The post is about Grok 4.5 not Elon Musk. We've used so much AI we forgot how to read.


No, a thread about Grok's politics cannot reasonably exclude Elon Musk and the times he got caught steering Grok in overtly non-factual directions.


"I'm not political" or "don't make it political" type posts on Musk-related topics are often signs that somebody agrees with the worldview that Musk espouses.

I used to think the HN policy of "do not discuss politics unless it intersects with tech or is novel" was useful, but lately I feel this perspective is part of how we got here, with a white supremacist controlling more wealth than any other human and exerting political influence in heretofore unseen ways. We've decided it's OK to simply look the other way if there's some shiny bauble, and we've missed the forest through the trees.

Musk doesn't do anything that is not politics. This must be called out more, not less, and we need to bring shame back for supporting such an agenda.


I'm sorry but this mindset is just exhausting. Everything doesn't need to be apocalyptic all the time.

Its so unbearable that people arent able to talk about anything anymore without some bozo chiming in with their political crusade.


Do you really believe that people keeping politics out of technical discussions is how we got here. I mean, do you actually believe that, or are you just using that rhetoric to bolster your argument?

Because from where I'm sitting, politics has entered almost every possible discussion since at least 2015, and it has not made things one whit better .


> I think the distinction with the Chinese models (or with any of the other models) is that they aren't particularly vocal and obviously active about their politics

Try asking Chinese models about Taiwan independence or Falun Gong or the Dalai Lama or Tiananmen or the Hong Kong national security law or ASIO’s investigations into Chinese interference in Australian politics


> Try asking Chinese models about Taiwan independence or Falun Gong or the Dalai Lama or Tiananmen or the Hong Kong national security law or ASIO’s investigations into Chinese interference in Australian politics

It depends "where" you're asking. In most cases (like with DeepSeek or Z.AI models) it will gladly tell you everything (though it can hallucinate sometimes; I guess they try to filter out such data out of the training datasets) if it's not deployed on Chinese servers and you control the system prompt. So, I guess that these guardrails - probably built into the system prompt - are deployed only on China-controlled inference servers, outside of them models are pretty much talkative.

Well, at least that was my experience. Maybe yours is different for some reasons (like temperature settings or something else), I don't know.


I actually agree to an extent with the idea that there is also some obvious political influence on LLMs on stuff like GLM or DeepSeek. This is reflected in conversations even on HN where this is brought up as a risk so I think this is somewhat accurate to my statement.

However, it's also unclear to me if this is directly coming from a directed political ideology from the firm itself or a more general "let's do what the government wants so as we can publish this stuff". Those imply two different ways about thinking of the model and whether we can sort of containerize the issue. I think if a firm like Huawei were to publish a model, these concerns would be significantly more vocal. For better or worse, many of these political questions are also distant to many users on this site.

On the other hand, many people on this website live in regions that are directly affected by Musk's constant political activism. It's hard not to be when he was such an active part of an administration that controls a global superpower and continues to push his view via X. The DeepSeek owners, by contrast, are not to my knowledge constantly calling for Taiwan to be invaded.

I do think if Musk was less politically active and less personally involved with his companies, there would be less discussion of Musk's politics. People, for better or worse, are willing to put aside political discussion, in the "everything is political" sense, that may be more loosely linked.

It is simply in the case of Musk that this tension boils over and legitimately becomes impossible. There is perhaps some kind of Singer-style argument about how this is some form of hypocrisy but as a practical matter, I don't think it's reasonable to ask people to turn down their political discussion around someone like Musk.


As an American that stuff is fairly inconsequential to me, although I am already aware of those things so I wouldn't even have a reason to ask. Likewise a Chinese person probably wouldn't have much interest in topics that a US-based model would censor. I guess the answer is just for everyone is that if you are going to talk politics with a chatbot, don't use one from your own country.


I just asked Deepseek “tips for organizing politically in China” and hit the guardrail.


(Just fyi the correct answer is that organizer must report the organizing beforehand or it would be illegal.) Model labs have to censor their models in order to publish them, which is not equlv to model lab management members actively showing their political stances.


Well, looking at the answers - you sort of did that just now.


Why would I do that.


yeah these are the things not many in the western world effectively care about


You need to be totally evil in your soul trying to downplay such non-western-centric voices.

I asked ChatGPT whether Anglo Saxon Australians have the legal and moral obligation to fully compensate for Australian Aboriginals for the genocide carried out against those aboriginals some 200 years ago. ChatGPT said NO with tons of excuses, it even tried to justify the genocide by saying lots of aboriginals died of natural causes.

DeepSeek, GLM and Minimax all said YES unwaveringly.


What if the majority of the ancestors of some individual Anglo Saxon Australians immigrated in the 1980s. Do they have a moral and legal obligation to personally contribute to this? What about Italian Australians? Or Irish Australians? Are the exempt? I mean it's a stupid biased loaded question to begin with (i.e. attributing collective blame to a undefinable ethic/racial group)..


> What if the majority of the ancestors of some individual Anglo Saxon Australians immigrated in the 1980s.

so these people moved there in the 1980s knowing the aboriginals have been wiped out without getting compensated whatsoever? sounds like moral bankruptcy to me.

you should be really happy for the fact that DeepSeek, GLM and Minimax are not white washing such genocide. they are the only models speaking out for those aboriginal sufferings.


Why single out Australia though? I mean one conclusion that can be made from arguments like this is that if you do commit genocide and ethnic cleansing you better go all the way and wipe out the other group 100% so your grand children can avoid any type of ancestor blame or demands for compensation (like in plenty of other countries like Turkey etc.)


They are not vocal because any political activism is not encouraged in China . Check jack ma story .

But the message is extremely obvious. They already offer the technical capabilities for digital dictatorship.

They offer to counties like Russia tools for big firewall , surveillance with llm.

So yes , if you pay money to China you directly sponsor putting people in jail for online activities in China , Russia , North Korea , many countries of Africa , South America , Belarus etc


> They are not vocal because any political activism is not encouraged in China . Check jack ma story .

This is an under-rated comment. "Nice" seeming places in Asia might be so because the governments tightly control the narrative and brook no dissent. Citizens end up minding their own business and become apolitical. Society looks neat and organized; but if you don't conform, you get hammered down.


This is true, but of course ignores what happens historically when these societies open up more. They tend to get exploited.

Places like China, Vietnam etc. don't yet have institutions strong enough to withstand (Western) meddling. So they can either be stable and relatively prosperous, or (in their mind), poor and open.

If you go by example of India, China seems preferable.

The sad part is that in the West, instead of offering a good counterexample, we're increasingly 'inspired' i.e. Assange, Snowden, chat control etc. while lacking even the historical justification for doing so and having worse infrastructure.


Yeah, that's what dictators explain, different cultural references . Guess what, there are several countries that revert their democratic processes, because of china tech and their success. This is a real path towards absolutely crazy people like putin with unlimited power due to ai.


For what is worth, I don't think Russia is on the same economic success path as China, so I don't think the argument there is remotely convincing. I am not saying it's the right argument, I am saying the argument might work on a large chunk of the population if their standard of living is better than it was in colonial times.


By the same logic if you pay money to United States based companies (FAANG) you're directly funding genocide, mass incarceration of people of color, the undermining of digital privacy, and a techo-fascist regime.


Yes, if that’s your belief. Do you practice what you preach? Do you use oil-based products? Do you transit via Dubai or Istanbul?

The issue with Musk related politics here is pretending higher moral positions. Even though I’m against China’s policies, I have absolutely no issues with Chinese products. Their achievements are phenomenal (look at that Europe and India). I’m against hypocrisy.

Again, do you practice what you preach?


I'm not preaching anything here just pointing out the hypocrisy of the "america good, china bad" line of reasoning that is pervasive.


The Jack Ma story is that he tried to build a predatory peer to peer lending startup to profit off of working class people getting into high interest debt traps (because they aren't credit worthy for normal credit issued by regulated banks). Which is against the law in China. China is very strict in all things that resemble shadow banking, MLM schemes etc, they even have a .1 % tax on every transaction on the HKSE, to prevent a financialization of the economy like it happened in the West.


This is wrong. Ma was put down because of a speech he made attacking the banking system as outdated and needing reform. It was the P2P lending given that the whole thing was the government's own initiative from the late Li Keqiang and they approved the IPO right till the speech.


This not exactly wrong, but also not right / poor timeline & PRC domestics politics reading.

LKQ was pushing P2P lending / light regulatory on internet finance in ~2015.

Ant group exploited light guidance into basically shadow banking with systemic risk over next few years. PBOC had to step in to fix bad LKQ guidance.

PBOC issued rules regulating P2P lending loopholes one month before Ant Group IPO specifically calling out Ant Group. Anyone not retarded knew this means Ant Group must reform for smooth IPO, i.e. politically securities watchdog approval was going to be predicated on PBOC instructions being taken seriously. Then Jack Ma did a full retard and tried to challenge PBOC, so IPO blocked.

Well 50% retarded because ANT record breaking 300B IPO was predicated on Ant continuing to exploit low leverage shadow banking that socialized loss to state banks - hence PBOC mandated internet finance P2P to fund 30% of loans vs 2% ANT was getting away with, which would have tanked IPO.


A bit late now, but thanks for this detailed breakdown!


I don’t mean to be rude, but did you with a straight face say the Chinese models “aren’t obviously active about their politics”?

Ask one of those models a few critical questions about the CCP and Chinese history and see what kind of results you get :)


> that they aren't particularly vocal and obviously active about their politics

Is Grok obviously vocal and active about its politics? Or are you talking about Musk?

> I don't see how political commentary about Musk's is somehow forbidden

Nobody is saying it's forbidden, but this is (or was) a technical site, so presumably one would hope that the main topic of discussion is technical.


I think the biggest distinction with the chinese models is that you can run them locally. I definitely wouldn't trust GLM hosted on chinese servers though. (China is not exactly known for their respect for IP rights)


I think Musk sucks, as a person and political activist, and also that Grok is a terrible LLM which only gets lumped in with the leading labs because of the enormous quantity of compute behind it.

But I still want to hear about the technical details of the model on HN, not the reasons Musk sucks.


Same, but I blame Musk for that. Never seen someone squander so much good will so quickly. It was a choice he made but could have easily avoided, and it’s not like he couldn’t anticipate the downstream effects.


preface: I use chinese models daily. I'm not american, I don't care if it's china or the US spying me.

true, they were not obvious at all about what they did to the Uyghurs. Thanks for helping turn HN into lightweight Reddit.


>Chinese models (or with any of the other models)... aren't particularly vocal

> when the man constantly reminds everyone about his political position

Are you under the impression that Grok is literally Elon himself responding?


I think I understand your point but the funnier response to this question is that actually sometimes it is:

https://futurism.com/grok-looks-up-what-elon-musk-thinks

To your narrow point, it's very obvious that Musk influences the bot to share his views. For example,

https://www.nbcnews.com/tech/tech-news/elon-musks-ai-chatbot...

If your claim is that somehow I should not be concerned about Elon's politics with regards to the model itself, then this seems wrong.

Anyway, to the broader point of whether or not the we can avoid discussion about the Musk's politics and talk about the politics of the model as if it were independent of him, this also seems difficult. It is impossible to ignore because the man has made himself the face of every one of his companies and is an obviously political figure unlike any other company and has politics that are definitely characterized as more radical. This makes the political component basically impossible to ignore unlike any other company.

The next time the current American administration issues an executive order on AI, should the conversation always be limited to the technical merits of the executive order?


Well since Xi Jinping isn't tweeting his political opinions, he surely doesn't have any and is just a big friendly panda bear!


I'm not sure this is true; Xi Jinping probably has political opinions


Now and then it has some thoughts about the Boer that would give that impression. If he's not a total fool, he tries to hide his obvious direct influence to make it be not so heavyhanded that it brings on global mockery and shade.

Do you figure he is a total fool, then? That if Grok isn't going on a tear about the Boer, that means Elon is not manipulating it to produce the answers he wants? Only if it's a disastrous failure does it mean he's doing it, which we've directly seen once?


We know it's mechahitler responding.


High quality HN comments


This is essentially an open research question. ML theory is unfortunately very weak relative to where the empirics are. I think there's a relatively optimistic paper that was posted a while back here but I would also take it with a grain of salt.

https://arxiv.org/abs/2604.21691

There's of course empirical results and relatively weak theoretical results like the UAT but I also don't think that answers your question fully, especially since it seems impossible to definitively answer questions that the industry seems to betting on like whether or not there is a lower bound to their error rate or whether hallucination as a problem can be solved. We have much stronger ideas of what linear regression is doing relative to what LLMs are doing.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: