There is so obviously competition in this market it’s astounding to me that people attribute all this malicious behavior to the model companies.
They’re growing over 10x a year. They want users and revenue. In order to get users and revenue, they want to provide the smartest models at affordable prices. If they unnecessarily burn tokens, users will get less value and switch.
This thread is filled with competing comments about their monopolistic power and how when one model provider was no longer doing a good job people switched to a different one.
In my experience starting with Gemini 2.5 Pro, moving to 3 and 3.1, 3.5 Flash, 3.6 Flash, and finally 3.7 Flash, 3.7 Flash is just as good if not better than 3 especially on high resolution mode (same token count per page as 3.1).
I run complicated, messy PDFs through these models. 2.5 Pro required a lot of kludgy hacks to get it to fully "see," but from 3.1 pro on I've removed many of them and haven't spotted problems.
3.7 Flash scores better than 3.1 pro on most benchmarks, leading me to believe that even if your OCR requires reasoning to interpret text or data, 3.7 Flash is probably going to be better.
Yes. I have Claude route to codex all the time as part of a dev process where fable is manager and it oversees the work of sidekicks and subagents. It’s happy to comply.
I believe that there are studies that show that merely getting very sick increases your chance of dementia - essentially it ages you faster or brings chronic disease forward. If that's the case, vaccines for things like the flu - a disease you're likely to get - are probably good overall.
They also activate the immune system in generaly, which could probably go either way in terms of longevity.
In general I don't think vaccines are preventing so much as delaying dementia, but if they stop chronic infection they might be.
When the vaccine came out in the UK, they included a hard age cutoff: Above a certain age, you weren't eligible. Below that age, you were eligible.
They looked at the probability of a dementia diagnosis over the 7 years after the vaccine was introduced.
People who were born in the "can get the vaccine" group have markedly lower rates of dementia. People in the "too old" group have higher rates. It's cut and dry. The researchers didn't separate out the people who actually got the vaccine.
It's one of those studies where you don't even need to look at the p-value to see the difference between the cohorts.
It's not even dementia, it's shingles itself. Two friends of mine had shingles in the last year or two, one of whom has a pain threshold at a level where I'm not entirely sure he's human ("I pulled the flesh back and you could see the bare bone, it was interesting, I've never seen my insides like that before"). After hearing their description of the pain levels involved I got a shot the next week.
My older neighbour also had it a few years ago, but in her case wasn't aware it was shingles, just some rash on her face. Doctors said if they hadn't stopped it within the next few days she'd have lost her sight.
Get the shot. You probably won't need it, but if you do then you'll really need it.
So much of neurodegeneration has turned out to be the effects of latent viruses in the body. HSV-1 is correlated with developing dementia along with many respiratory viruses.
Why on earth wouldn't you need to look at the P value? You give good arguments for the soundness is the methodology, but you still need to look at the results and their statistical significance.
Their P value is 0.02, which is good but certainly not definitive. Also the effect is kind of small, 3.5% reduction in diagnoses.
I think I'm failing to understand your subtext. Obviously older people have higher rates of dementia. A study reporting that doesn't tell us anything about the effect of the vaccine.
Sorry - the older group has higher rates of dementia than the younger group when they reach the same age - so when they reach 75, they're more likely to have dementia.
I see what you mean, there's a clear gap in the fitted (regression) lines, suggesting the trend of dementia with age is different in the two groups.
But I wonder if that's just a statistical artifact. The overall trend looks the same in both groups if you ignore a couple points (ages) on either side of introducing the vaccine.
A single line appears to fit all the points well, except two points on either side of the divide.
The normalization of the age variable still looks problematic, as it requires that there is a linear dependency between age and incidence.
Both the scatter plot and also the probability graph seem to be curving upwards w.r.t age, i.e. the incidence of dementia cases seems to accelerate with age. Which at least makes a lot of intuitive sense, and would forbid just plotting in a regression line.
Having a more or less circumstantial point to arbitrarily cut your regression line in half also just begs for introducing Simpson's Paradox.
There was a hard age cutoff in the UK study. Above a certain age, you weren't eligible. Below it, you were. People who were born in the "can get the vaccine" group have markedly lower rates of dementia. People in the "too old" group have higher rates. It's one of those studies where you don't even need to look at the p-value to see the difference.
Thanks very much for linking that substack, very informative.
But in your summary:
> People who were born in the "can get the vaccine" group have markedly lower rates of dementia. People in the "too old" group have higher rates.
I would change "people" to "women". I thought it was very interesting that the benefit of the Shingles vaccine eligibility for Alzheimer's was largely confined to women - men showed no such benefit per the graph in that substack article.
The video explains why the analysis is wrong and gives a very clear graph of how when you look at the proper data, there’s no benefit from the shingles vaccine for dementia. The signal is clear as day.
FWIW Eric Topol is an extremely unreliable source. He has “fame” but most of his stories end up being wrong because of his poor analysis like the review above. I subscribed to him during the pandemic when he migrated to substack but ended my subscription after countless bad articles.
Right. And RCTs not showing things than can be seen in larger observational studies seems plausibly related to the problems with RCTs, for instance statistical power being generally lower because n sizes are lower, attrition is an issue, etc… I want to believe the research though, especially since I like most smart people are going to get the Shingrix shot anyway.
So you have a 30 person RCT, half the treatment bails midway through, you include the assigned but not treated (ITT) in the analysis. That’s going to give you the power to make a definitive statement while a large scale longitudinal with n=50,000 will not?
Computer use is a great idea. It gets the job done when nothing else will.
If you're a person trying to get their job done at a big company, but half your job is in 1-2 proprietary tools or is stuck behind an API you can't program against, computer use can allow you, a non-techie, to do your job more efficiently.
I think it's an awesome way to circumvent gate keepers and the IT department to let people accomplish their goals.
It does. I used to be an ahk "script kiddie" and know it front and back. It's sort of burnt into my brains. As a result, I can prompt really really well, notice issues at a glance, and I have a sheer volume of scripts locally for all sorts of tasks some from as far back as 2014. From tiling window managers to OCR all the way to simple hotkeys/hotstrings. I let it grep in that folder and build out whatever I want using those primitives. This gives actually 1-shot immediately usable 100% working scripts even with GPT3.5 level models, as opposed to the iterations needed for typical development.
Example: adding copyright text box to bottom of every slide
Yeah, it's not that computer use is the most theoretically optimal paradigm, but there's a reasonable case that given the constraints of modern software systems and how they're built, that it's the most realistically optimal paradigm.
I think there's a sweet spot- a lot of the time you're probably better off with "reverse engineer this web page and build me an API or personalized chrome extension to meet my needs".
I have an agent doing price checks for me for an item on a certain website. Instead of blasting through a zillion tokens processing the DOM over and over, it loaded the page once and figured out how to download a json with the price.
It curls the page. I think the approach it took won't actually wouldn't work in my local browser- its getting the value from some conversion reporting code that I'm guessing my ublock extension would hide.
Computer use is very useful for developing GUI applications since claude code can build and test the entire app end-to-end (accessibility APIs exist but depending on the UI framework of your choosing you can run into walls very fast).
I run it in a VM using a headless wayland compositor, I'd never trust even fable with access to my real system.
How are folks using “computer use” to click things on intranet portals that are behind an SSO?
Even this OP example shows visitors a url and enter this search term… that is port of useless.
How can I automate things behind an SSO wall? Even if it means I manually authorize it once and watch it do things on its own..
I've never used Gemini computer use, but I assume it's the same:
Claude computer use takes control of your whole computer inputs (mouse and keyboard) plus screenshots. You just log in, tell Claude you're logged in, and let it get to work. It'll use the browser you're logged in with.
The chrome extension is a little better because it only takes control of its own chrome tabs (again: you just log in.)
For anyone trying to figure out how to build a society where no one wants to be a criminal, I highly recommend When Brute Force Fails: How to Have Less Crime and Less Punishment by Mark Kleiman.
There are evidence-backed ways of reducing criminality.
One counterintuitive way of reducing crime is to increase the likelihood of being caught, to have small-but-increasing consequences for committing crimes, and to increase the swiftness of sentencing.
For example, if you are caught drinking and driving, you immediately spend 1-2 days in jail.
Long sentences are not very productive at reducing crime or at least are a very inefficient way to do so.
An intuitive way to understand it is imagining that there was a system where if you stole something, you 100% of the time got hit with a charge to your account of the item value + $10. No one would steal again even if the penalty for getting caught was relatively nothing. Because the feedback loop is so short and guaranteed.
No ones life would be ruined over a dumb choice and yet they would change their behavior very fast.
It’s still the same. If you steal something and have no money you lose the item and get some small penalty, perhaps a day in jail. If there is absolutely no chance you’d get away with it, everyone quickly realises there is no point.
On the machine I have (ResMed AirSense 10) the data gets written to an SD card I can pull out and import into OSCAR to see what the machine detected about my sleep apnea events.
I also (not always, but when I'm testing changes) use a continuous pulse oximeter to test my SpO2 blood oxygen % over the entire night. The data from this can either be combined with the data from the CPAP machine, or used on its own to test things like how I do with the collar and no CPAP machine, since the SpO2 drops are the dangerous thing I'm trying to avoid by eliminating the apnea events.
I use the Wellue SleepU for SpO2 monitoring but there are a lot of similar devices (the kind with finger tip sensors are generally more accurate than the wristwatch based solutions).
They’re growing over 10x a year. They want users and revenue. In order to get users and revenue, they want to provide the smartest models at affordable prices. If they unnecessarily burn tokens, users will get less value and switch.
This thread is filled with competing comments about their monopolistic power and how when one model provider was no longer doing a good job people switched to a different one.
The competition in this market is ferocious!