> to the point that they can solve major outstanding problems in many fields of mathematics
I will die on this hill, humans solve math problems not machines, there's no automatic math prover out there. There are humans attempting to solve this problems either by leveraging these tools or not
Imo Eventually we will circle back to nonfiction quality books, for one just to decrease screen time :) for second we actually need them even for chatgpt
Lately whenever I enter a book store in Italy, I'm overwhelmed by the amount of fiction romance for girls and womens. I get 3/4 of readers are womens but all I'm left with is a gigantic pile of nonfiction shit I'm not interested about, I know for sure most of it will leave nothing to me. Increasingly way too many people attempted to sell a bunch of nonfiction books, and it shows there amounts of unsold copies. Most of them being a marketing political campaign. There gems for sure but it's becoming increasingly harder to find them. Also the only fiction books I'm interested about are sci-fi, and the dedicated space is soo small that you feel a weirdo with strange hobbies. But I swear to you, cause I read a fantasy novel from my wife's selection, that a saga like The Expanse is 100 times better written than all that new shit. Finally rest of the book store is now dedicated to manga and graphic novels. Not bad I love them but the ratio is overwhelming as well
"Smarter" is a vague term. If a bot can beat you at chess then in some sense it's "smarter" than you about chess. After repeating this feat in enough narrow domains, if you say "but it's not really smarter," this objection might technically be true in some sense, but it starts sounding increasingly hollow.
Playing chess, writing code, finding security bugs, and proving mathematical theorems all seem fairly similar to thinking and don't seem much like being tall.
> writing code, finding security bugs, and proving mathematical theorems
While the second can be automated as in "this is the repo go on and look for security issues", the first one and especially the last one do not ever happen alone. Terence Tao did the math proof not chatGPT that was just used as a tool, a tool can do smart things but it's not smart
The crucial distinction here is that "seem" does not at all mean the same thing as "is".
Thunder seems like the anger of the gods but it isn't. We've had chess playing programs for a long time now and despite it seeming like thinking is required for them, it isn't.
The principle you're using here isn't a scientific one but magical [1]. Abandoning empiricism and rationality is not a good way to make progress.
You're insisting on a particular definition of a vague term.
It makes sense now to say that temperature is what a thermometer measures. However, before there were good thermometers, people often thought that heat and cold were different things. The meanings of the words we use were influenced by scientific progress.
For thinking, we don't have a good thermometer. There are IQ tests, but they aren't aren't necessarily all that useful for comparing what people do to what machines do. And that's why there are a zillion AI benchmarks - none are entirely satisfactory.
So what does "smart" mean to you? How do you define it in practical sense? What definition should scientists settle on?
Without a proper definition, how do you tell the difference between "seems smart" and "is smart?"
People in the past being wrong doesn't mean you have to repeat the same mistakes, and it definitely doesn't mean you should throw your hands up and declare that all similar things must be identical.
Quibbling over commonly-understood definitions is not a strong argument. If you're genuinely struggling to understand that El Ajedrecista [1] did not meet any definition of thought, then the solution is not to demand that people redefine all terms to accomodate you, but that you consult a dictionary.
You clearly have a working definition, or you wouldn't have been able to declare that tallness isn't thinking; please engage honestly.
I have a vague understanding, enough to know that tallness isn't thinking. I think I can usually use the word correctly. (I don't think El Ajedrecista qualifies, but it was starting to play chess, so it's closer to "smart" than a rock is.)
That doesn't mean I know whether "smart" should be applied to what AI's do, and I suspect nobody else knows either. This is the sort of thing philosophers debate about, not common sense.
Turing invented the imitation game because he didn't really know either.
It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.
The amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data.
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
reply