Hacker Newsnew | past | comments | ask | show | jobs | submit | rsfern's commentslogin

For numerical code I like einops.reduce more than numpy/pytorch sum reductions because you can reduce over named dimensions. It’s much more readable than having to reason through axis indexing again every time you come back to the code

That’s the prevailing narrative, but I think this controversy calls it into question to some extent. If the OpenAI result wouldn’t have been possible without experts seeding the training data with feedback on promising solution routes, there’s less reason to believe this, IMO. More information and transparency is needed

That's the pickle isn't it? AI solved it with human help. But it's standing on their shoulders. Neither AI nor the humans got it alone.

Humanity is better having solved this issue. But which humans were credited and benefited is the issue.


I agree (and so does Buckmaster based on his written statement) that we are better having solved this.

But I disagree that which humans were credited is the heart of the issue in this particular controversy. The question is what do you need to bring to the table for a result like this. A pre-release frontier model trained on the open literature and $15 million of inference? Or all that plus a year of the experts finding the path to the solution for the model to run with?

I think it makes a huge difference in terms of what we think the future of mathematical research will be like, and whether we should still encourage students to go into this field, which was the original topic of this thread


I think you’re missing an important distinction. “Major damage” to the talent pipeline because models become capable of original end-to-end mathematics is what the community has been discussing. But if the models rely on sniping nearly complete work then this damage is antisocial without a lot of upside, it would be destroying a talent pipeline that would still necessary for continued progress.

Which is it? I don’t think OpenAI is being transparent enough for us to really understand whether these results would have been possible without relying on unpublished information from the solution strategies of the experts


The session data could be cryptographically signed. Probably easier in an open harness?

Without seeing the full correspondence it’s hard to evaluate for sure, but parent linked to a tweet from the OpenAI employee at the center of the controversy, that’s a primary source you can read and evaluate yourself

Personally I don’t find the tweet a satisfactory explanation of their behavior, it seems like a lot of deflection without directly responding to the specific claims of front-running, and the screen cap of the correspondence doesn’t include all the relevant context. If there was really no bad behavior, why not post the whole thing?


> If there was really no bad behavior, why not post the whole thing?

what is whole thing?..


The screen cap of their full correspondence in the Twitter thread, which was the subject of the preceding sentence. The Twitter post has an obviously incomplete fragment of the conversation that doesn’t resolve what the author presents it as resolving, which is the dispute over how the discussion of authorship of Alpöge actually went down

I read Tristan's allegation from https://cims.nyu.edu/~tristanb/statement.pdf, and he didn't make it clear if threats to him were made over email, texts, calls or during personal meetings, so it is not clear if the whole thing was in writing.

He mentioned that they had meetings:

I asked to speak the following week. On Friday, September 4th, I was asked whether I could meet that day; I again said the following week. At 12:45 on Sunday, September 6th, I was asked whether I could meet “at any point today.” Sebastien Bubeck joined. The three of us spoke twice that afternoon. Levent was not on the calls.


Regardless of what you think of the priority dispute issue discussed on sibling threads, I’m highly skeptical of the closing quote that this Navier Stokes result means that the same approach of casually spending a few million on agentic computation is going to solve end to end materials design or drug development.

Those problems can’t be formally verified with an automated theorem prover. We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations because otherwise they’d be too computationally expensive, or we just don’t have the right data to parameterize them beyond describing qualitative behavior. Agents are helping accelerate research in these fields but I think it’s mostly a different class of problem that’s a lot harder to specify and verify


Yeah, I think you can't just throw money randomly at problems and expect results unless you know a line of attack that can get you all the way. OpenAI chose the line of attack only after it became known to them via rumors. They "front-ran" the researchers.

Yes. What the headlines hailed as an AGI discovery the facts show more to be someone spending years mining for gold, rumor gets to OpenAI that there might be gold in this specific place, they mine there and instantly discover gold, then tell the world they’ve developed the worlds best gold finding/mining machine.

Separate from all the allegations of more nefarious actions and ethical issues, that’s the most charitable version of what happened here.


they threw it on all the millenial math problems (I think there are 6 at this point unsolved, well, 5 now).

And according to them at some point they saw that one was close to being solved, so they pointed all the agents at it.

The same thing happens to humans - at this time there are no simple problems left, so solving the hard ones requires using prior knowledge and attempts at solving things.


Yeah but the one they decided the AI was close to solving may have been so because the researchers' progress on this problem became part of the training data for that AI...

Almost but not quite I think. You can throw money at parts of problems. I think it's helpful to think it kind of like supercomputer MD/MC or electronic structure calculations. A tool that can get you valuable answers but not necessarily aid understanding. Simulations can be used to aid understanding also, and are integral to theory development. In the same way the approach to this result is.

Worth noting they claim they did not choose the line of attack. Of course we don’t know whether that is true.

Plausible deniability - The line of attack is in their sessions/prompts data. Just make the prompt pointed enough that the search space is tractable and use your ginormous compute.

> "Of course we don’t know whether that is true"

Yep. Who is verifying these claims? We all know how trustworthy Altman & Co are.


Yeah, they didn't choose the line of attack, the person they copied it from did..

but the researchers were also largely relying on AI

“relying on” is misleading here relative to what the researchers have said.

If I write a book and pass it through a spelling and polish checker, I still wrote the book and its core IP. I didn’t “rely on” the tool to create the IP.


it’s much more like you come up with the premise and someone else writes the book. the released prompts for other foundational problems (like unit distance) prove that.

The tools the researchers used though was much more than an spellchecker, because spellcheckers don't come up with chains of reasoning for the arguments in the book. The LLMs did in the case of the Navier-Stokes problem.

If this were the case the problem would not have remained unsolved for this long. A new spell checker is not what cracked the problem.

In the same way you rely on a keyboard or touchscreen to type this comment. It doesn't mean the tool is the brain behind the work.

that's not how AI was used in this case. It's more like a professor with assistants.

Professor says the assistants - why don't you dig in this direction, I have a hunch it might produce something valuable. And AI assistant does just that, proving or disproving a hunch. This would take the professor a lot of time if doing by themselves.


Yes, and keyboards also save a lot of time over handwriting. Numerical methods and proof engines save even more time. LLMs are just another tool in the kit.

Keyboards don't suggest chains of reasoning or words to type. When I press the K key, I know exactly what will happen. It's just a translation layer that gives an output known ahead of time and thus does not impinge upon the creativity of putting words together.

A better example would be playing chess against a player slightly stronger than me and using a chess computer to suggest some good moves. I could win, but it certianly wouldn't be just my brain that wins. It would be an amalgamation of my brain with a machine that suggests good moves.

One cannot simply reason by analogy.


This is straight up misleading. When you press your K key on a touchscreen, your keyboard program may decide you meant to press the neighboring L key (by dynamically inflating the collision geometry on it) because it was statistically far more likely that you meant to press the L key next.

This likely doesn't happen exactly on an analog keyboard, but then many text-processing environments that do the same thing in post. My keyboard just edited 'yuor' to 'your' even though I successfully input the prior string.


> Keyboards don't suggest chains of reasoning or words to type

My iPhone keyboard does


I was talking about regular old keyboards. But you're just being pedantic.

frankly don’t know how to reply to these sorts of comments anymore

That's usually a good sign you are on shaky ground!

Obvious false analogy in your earlier argument.

No so obvious to this guy.

it is truly not obvious to you why keyboard isn’t a good analogy for LLM?

The researchers were driving prompts and trying to actually do math.

The OpenAI effort was a pure brute force attempt. I'm not even sure an LLM was actually involved. I think they just used their hardware to run the matrix multiplies required by the search for a counter example. Perhaps some clever approach guided the search but that seems to be about it.


You can’t do anything novel with these models from scratch and let it fly. I’ve observed something over the past few months

Work on something novel -> llm is kinda useless and low value-add -> Keep at it and in the process feed it more information -> keep doing this periodically -> a few months go by and you realise the model outputs are almost like-for-like regurgitations of what was inputted in some prior period.

Once it’s accumulated new info can it produce something automated that is somewhat useful? Sure.

But by itself - absolutely not.

I clearly see humans will be needed - the best ones that is. For ‘rote work’ and stuff that is not IP sensitive firms will be ok with employees putting that as inputs into models.

But I’d wary about trusting the labs. They will push the letter of the law to the max.

Personally I’ve stopped doing anything novel with these models. If I do use a model on something adjacent but not totally novel I have to craft the inputs in a strategic way not to give much away.

I’d wager firms will soon realise this and that growth rate of revenues of the frontier labs will become questionable. The economic cost that firms have brought out thus far is only financial. There’s a whole bunch of other costs people aren’t talking about.


Agree all. And as the revenues become questionable, the frontier labs practices will necessarily become (more) questionable. Vicious cycle.

To avert that dynamic, the frontier labs must deflect and otherwise act to prevent this controversy from breaking through. Both to the general public, but also more specifically to the firms' decision-makers. All of whom are generally aware of the IP issues, and some of whom are aware of what happened with Cursor and Figma, but with few exceptions have not yet themselves acted to protect their property.


This doesn't follow for me. There are what, Dozens or Erdos tier problems that got solved with no progress for decades? How does that factor in to your view?

IIUC the argument is that while unsolved there was much work done on them that shows up in the training data. The idea being that the LLM is limited to a small amount of inference over externally supplied data.

And yet the best humans could not use that same available data to solve the problems.

I think you stopped at the wrong time with the wrong perspective. Why can't that info accumulation part also be made more self-contained?

I guess I'm having trouble unraveling your experience and personal usage vs. what you're concluding about the labs.


I’m pretty sure that OpenAI has some of the best mathematicians prompting the models and analysing the results. While they are marketing as if the model solves problems themselves.

Prompting them yes, suggesting potentially fruitful research directions and so on, but the actual research was conducted by hundreds of agents swapping millions of messages and using billions of output tokens over 88 hours. The result being a huge Lean proof: https://github.com/openai/NavierStokesAndEuler. It's not just possible for humans to manually guide such a process in a meaningful way. They can set the direction and attempt to understand the result, but they solution itself must emerge (or not) from the agent swarm.

So yes, the models do seem to be "solving" the problems themselves, but not necessarily in the way we think of mathematical discoveries happening. Academic mathematics has historically been resource constrained: There are a limited number of top-level mathematicians, and they only have so much time and brain power to spend. So when approaching a problem, they are essentially forced to be as efficient as possible, not just searching for a solution, but for one that can be achieved within their cognitive budget. This induces them to develop novel techniques and abstractions, and it is actually those techniques and abstractions that tend to be the valuable part for further research, not the proof itself.

An agentic swarm is like getting a single skilled mathematician, cloning them a hundred times, then locking them in a room with the single objective of solving a problem. No longer constrained by time or brain power, they can approach it differently, using pre-existing techniques to gradually build their way to a solution. This process might not require a single intuitive leap or new discovery, and the solution will not be simple or elegant, but they will probably get there. It is more like a process of intelligently guided search than invention.


The OpenAI team didn't make a Lean proof. They brute forced a counter example. The "other" team was doing what you described but they haven't "finished" their work yet. Also, their Lean proof was for a simpler version of the problem, not the full NS.

Also, OpenAI wanted the actual mathematician taken off the resulting paper. I'm not sure I would describe what OpenAI did as research. What the other team was doing does seem to be more like research but the hardware was still in those cases mostly brute forcing things and then doing something like a genetic algorithm to compose an actual proof based upon the results of a large set of brute force attempts.


Didn't OpenAI make a Lean proof? https://github.com/openai/NavierStokesAndEuler/tree/main/Nav... "This repository contains Lean 4 formalizations of the results presented in “Finite time blowup for Navier–Stokes” and “Finite time blowup for the Euler equation” by OpenAI."

Theres a reason those same mathematicians did not solve the problem on their own. Minimizing the impact the model made here seems unjustified.

For any practical application, numerical solvers for Navier-Stokes already exist and do a good job.

This proof is just checking the boxes for mathematicians.


The efficacy of applied NS was never in doubt. "Checking the box" is downplaying the magnitude of the discovery quite a bit as it has been unsolved for almost 100 years. Yes, this particular problem with NS no real-world applications, but that's true for 99.9% of math research.

There is no NS proof here. Its just a counter example. There is another team working on a proof but they aren't associated with OpenAI.

Agreed, but i think this underscores my point. We have numerical simulations in materials science too, but that doesn’t mean formally verified theorems about the underlying equations automatically translate to formal (or even informal) verification of simulation results. That’s not to say you can’t make progress with agents, but I think it’s less well defined how you write the goal and progress assessment for an agent

you're as sure of what you say as wrong about it.

Note that your reply has exactly 0 value for anyone who doesn’t already know where and how the parent poster is wrong.

fair enough

which part is wrong?

> For any practical application, numerical solvers for Navier-Stokes already exist and do a good job.

or

> This proof is just checking the boxes for mathematicians.


If that were true you could explain it. There are lots of solvers for navier stokes simulations and they do a good job.

I was referring mostly to

> This proof is just checking the boxes for mathematicians

There are already quite a lot of summaries of the story that lead to the solution, and the impact that the intermediate results have had.


True, this is just a much bigger one with a vastly larger amounts of hardware.

The same could be said of your post.

OpenAI (claim to) show the existence of *a* finite time singularity. It could stimulate more research in PDE solving, and maybe physics, but it has zero impact on practical applications, that I can see. The Millenium problems were chosen based on hardness not practical relevance.


I was referring to the "it's just mathematicians checking boxes" claim

> same approach of casually spending a few million on agentic computation is going to solve end to end materials design or drug development."

you're not actually spending that money. it's sunk cost, as you already bought the hardware. at least for the big pharmaceutical companies for drug development. then you run your own local model, trained on special data, with special etc, etc... to the end of buying GPUs for what, 3.5-6.5M/rack or so (GB300 NVL72, Google AI summary pricing quote) becomes a bargain (vs the double digit billions you need to spend on a new drug R&D).


Solve logically? Sure.

Solve for how to implement and synthesize physically? Not likely.

Humans solved for launching rockets to the Moon on paper decades before it happened.

Pareto type thing; the logical work is the easy 80%. The last 20% is fighting physics.

There is no beating physics but there is still plenty of room for us to improve our understanding of it.

Which we weren't focused on at all sitting millions primates at well understood physical computers searching for Shakespeare Python and Ruby code yet merely getting same old contemporary software outputs.


> Those problems can’t be formally verified with an automated theorem prover.

It certainly seems like any problem that is amenable to reinforcement learning will be solved.


It does, yes. So designing objections functions and making sure you can afford the training rollouts becomes really important in defining which problems are tractable. It will be really interesting to see how that shapes the kinds of problems people choose to work on

>We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations

Do you think it is possible that better math will lead to better physics models?


It might but the math results from GenAI so far have been limited to finding counterexamples to known conjectures, not building new mathematics.

> results from GenAI so far have been limited to finding counterexamples

Not all.

Ehrhart’s volume conjecture

Quantum parallel repetition for general two-player quantum games

Erdős Problem #183 on multicolor Ramsey numbers

Erdős–Sárközy Problem #12(i)/(ii)

Erdős Problem #125

Log-concavity of codimension-3, type-2 pure O-sequences

Optimal O(1/t) last-iterate convergence for Anchored Gradient Descent-Ascent


True, these results are from last month. My info was a little out of date.

Prior updated :)


Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods

Working with AI on science (not LLMs though), couldn't agree more.

The TL;DR is still "AI helpful, but not end of the line". The live discussion about these matters is always ridiculously inflated by hyperbole.

I think a lot of people don't understand, with respect to a mathematical theory, the relation between a carefully stated conjecture requiring formal proof vs. using the objects in the theory effectively. Could better understanding of NS lead to better practical tools? Almost certainly, even if only to give us bounds on performance. Has its unresolved status stopped us from using NS? No. Almost no one using it cares. Resolving it is valuable, especially if it comes with mathematical and/or physical insight leading to greater understanding. But it is not this is grand result that like instantly unlocks 100+ day weather forecasts.

It would be like saying proving ergodicity more generally for physical systems would unlock condensed matter physics, ignoring how well stat mech has served us regardless.

I am not anti-AI and I don't think we should stop throwing them at conjectures. I'm against this fundamentally misleading type framing that's become prominent. Millennium prize problems are important. Treating this specific aspect of NS as the one missing piece is just harmful. If we just throw compute at formal conjectures voila cancer and fusion.

I think the better example of "AI" usefulness toward solving problems is AlphaFold, and immensely powerful tool. But also suffering from a false framing/marketing problem as "solving protein folding". It feels like the right use of compute. Considering many factors that we can't hold in our head at once. "Solving" something that was already "solved" via computation (simulation) but now much more efficiently. The output is a valuable tool itself, it was not about "solving the protein folding problem", which it didn't do. It is a tool to solve problems requiring a sequence->ground state calculation. Which is a very broad set.

Formal verification of a conjecture we set up as a benchmark we set to test human understanding is not valuable in the same way.

I'm failing to make multiple points and gotta run, but i think that final point is important. The millennium prizes are not about technological/practical value, at least not intentionally. They're about shit that seems fundamental to us, things that feel[1] to us based on our understanding are important AND feel like they should be solvable in a human-comprehensible way. So formally resolving them with pure compute is not really the point. It seems closer to that story about one of those prime conjectures where some guy just ran brute force enumerations to find a counterexample. Valuable for sure, time-saving. And knowing the answer makes it a lot easier to solve a problem.

TL:DR science and math are more than formally resolving conjectures, they're about building up understanding and tooling that you can then build more on. AI should be an increasingly big part of it, but declaring "AI will solve fusion because it's smart" is like the rest of the fucking owl meme. I have no doubt it will help, most likely via simulations/quicker testing/calculations and verification. Maybe partly via reactor designs. Maybe partly being fed conjectures about bounds/limits that would be useful as inputs for the next iteration. And maybe even in the form of resolving some formally stated conjectures (I don't know enough plasma physics to name any).

[1] obviously to the mathematicians it's more than a feeling..


Interesting project idea! The link seems to be 404, is the repo still private?

Yes, I forgot to open-source it. Now, it’s publicly available on GitHub!

Thanks! This seems really cool. If I’ve got it right, your UI builds and displays these diffs (and implements undo/redo) by parsing the edit tool calls?

If so that seems really nice, one of the things I don’t like with coding agents is having to defensively commit changes to roll back if the agent goes off the rails. I’m not sure if that’s a problem with my workflow, but your tool seems great for exploratory stuff

How does it change the way you personally use pi? How git aware is it, do you have to commit at the end of a branching session, and does it handle manual edits?


Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model

Indefensible, not impossible. As you say it is quite technically feasible.

The bit you quoted doesn’t capture why the reporters are upset:

> White House journalists are outraged that a threat credible enough to force Donald Trump to escape from Air Force One using an airport catering truck wasn’t relayed to them.


Point, however...

Assume you're the Secret Service's Grand Poobah of Presidential Protection. And actually trying to do your job. How, exactly, might you tell a gaggle of reporters that the President slipped out of the baddie's crosshairs, without a huge risk of them giving the game away? This ain't 1989, when their being inside Air Force One gave you total control over their communications with the outside world. And "President is in that slow, unarmored catering truck over there" may be all the targeting info that the baddies need.

Note too that "everybody else's lives are worth less than the POTUS's life" is right in that Grand Poobah's job description. And legal duties.

My guess: What actually has the journalists pissed is that they were reporters, sitting figurative feet away from a huge secret news story...and nobody felt they were worth letting in on the secret.


Exactly right? Why would anyone expect a competent security outfit to tell the reporters anything?


Let’s distinguish a bit. There are political appointees (Trump’s government employees as you say) who are mostly upper management, and there are career civil servants (all the government scientists are under this category) who have a strong culture of apolitical dedication to the mission of their agency and to the American people and Constitution, regardless of who the current president is. And in the DOE labs in particular most (not all) of the scientists are actually employed as government contractors, but they have a similar non-partisan ethos.

That doesn’t necessarily mean there’s no need to be concerned with potential impact of policy and priority changes from the administration, but it does temper the threat model because the government employees you’re considering trusting have given oaths of office to protect and defend the Constitution.


> and there are career civil servants (all the government scientists are under this category)

US national lab scientists are not even civil servants. The labs themselves are run by a corporation under contract to the DOE and the scientists work for that corp. The managing corporation changes from time to time and the scientists transparently start working for whatever assumes the replacement. The land, the hardware, the buildings and any physical products are owned by the US gov't. To a very large extent, the intellectual output is set free to the world in the form of papers, presentations and to some small extent (eg compared to CERN) in the form of software.


Right, I did specifically say that most of the DOE scientists are contractors, but I concede the phrase “government scientist” is a bit ambiguous. I appreciate the extra detail you added. I think the distinction between political appointee and scientist/researcher stands.

As an added complication, some of the DOE labs do have civil servant scientists, for example National Energy Technology Lab and National Renewable Energy Lab are like 50/50 civil servants and contractors. And most of the funding arm of DOE are career civil servants. LANL, Sandia, Livermore, Argonne are all staffed by contractors


> I think the distinction between political appointee and scientist/researcher stands.

I agree with this.

I'm a civil servant and I know (personally; my work is nowhere near the labs) a number of people who work or have worked in that weird contracting DOE/DOD structure that includes the labs.

The difference I've seen is far more from the nature of work rather than the employment details.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: