I suspect there'll be a lot of differing opinions on this, and I'm looking forward to seeing the discussion. What would be helpful would be if people could say how much experience they have in hard statistics, and how much what they say is driven a priori from the data.
I know that hackers, in particular, have real problems with "Argument from Authority", but stats is one place where it's really, really easy to go wrong, so knowing how much formal training someone has can be an important indicator.
Statistics, like engineering, is done with numbers. When you attempt to do it without numbers, that's called "opinion".
I have not tried to look at the numbers. Here is my opinion.
My experience of school was that a small minority of teachers are truly excellent, and a small minority were horrible. The truly excellent ones are not distinguished so much by what happened in their class as by what happened in the following classes, and what classes they left people excited about taking. The terrible ones, by contrast, showed up as poor performance all around.
The described data set has a clear correlation. Whether or not the correlation is meaningful in individual cases is much less clear. My initial approach would be to try to fit a hierarchical Bayesian model to a data set, then use the resulting model to come back with predictions about individual teachers. Teachers whom, after several measurements, are overwhelmingly identified as terrible should be removed. If you find a population of teachers who show up as superstars, they should be subjected to further study to see if we can predict the quality of incoming teachers, and to see if we can learn lessons from them that improve other teachers.
However this is a well-studied problem. I'm sure someone has tried something like this. I am sure that there are a lot of vested interests. I have not attempted to evaluate work in this area, and I have no opinion on how good it is.
I'll bite. I won't discuss my experience in my posts since I don't want to argue from authority but will do so below. But first, I should point out that you're committing a similar sin to that which is alleged in the article. What you really care about is the correctness of an argument. You hypothesize that formal training is an important indicator of correctness. Presumably there's also noise in that scatterplot but we don't even have it. Since it's the best thing available, you are choosing to rely on it.
As for my formal stats training: None beyond High School. That said, I've worked with and interviewed many with far more training. One thing, I learned is that most people don't have great intuition for statistical problems and it's not terribly highly correlated with years of schooling. I've met PhD's in economics and statistics who have made basic conceptual errors as well as those with less training who were more reliable.
So while training may be correlated with accuracy, you should still demand a well-reasoned argument and think critically about it.
I agree almost entirely. I'm specifically not asking for people to say "I'm a PhD in statistics, and here's the answer. Accept it because I know better than you." What I'm asking for a a complete and reasoned response, along with some evidence as to how much I should listen to you in the first place.
If you have no such evidence then the onus will be on you to make your argument more complete, more coherent, and more comprehensive. If you have evidence (note: evidence, not proof) then you can be a little less rigorous in what you say, and rely on people giving you the benefit of the doubt while they work through the argument.
What I see a lot of is long, apparently good arguments, that then turn out not to be as complete or coherent. they sometimes just don't hang together.
Significant amounts of formal study in a subject is evidence that someone might just have a better understanding. After spending a lot of time on the internet I'm tired of having to wade through every single argument in detail looking for all the possible chinks.
Maybe that's just impossible. Maybe every person has to redo every analysis for every argument. Seems like a complete waste of almost everyone's time. What about "Don't Repeat Yourself" or "Don't re-invent the wheel." I guess we are doomed to reinvent the wheel in every single discussion.
Reinventing the wheel would be a major problem if our goal is to solve the education problem with this discussion. No one here has done even the basic work I would expect of someone trying to understand teacher evaluation as a solution and compare it to other alternatives.
I would argue that the whole point of HN is to think through arguments in other domains and build intuition by reasoning through problems and arguments. Otherwise, what's the point? No one is going to arrive at this thread and scan the top-rated comments for the solution to his school district's problems.
In cases where actual decisions are being made where the analyses are much more thorough and fully validating much more expensive, other techniques are available. First, one generally builds an awareness of the strengths of each team member which suggests where errors may be more likely. Additionally, one can check a random set of the most likely problem areas. Perhaps, most importantly, while everyone won't re-do every analysis, it's highly unlikely that an any important analysis will only be done once. So one can expect that the high-level results are generally consistent.
How much formal training does the author of the blog post have? From the short bio, it doesn't sound like much.
That said, my training -- read Friedman's intro text in high school (just freetime reading on my own -- so not formal per se). Two quarters of stats as an undergrad. Two quarters in grad school. All of which at least 15 years ago -- so mostly forgotten anyways. :-)
Or you could ask people to put forth coherent mathematical arguments, since research has shown, for example, that most professional PhD-holding published medical research is statistically incorrect.
I know that hackers, in particular, have real problems with "Argument from Authority", but stats is one place where it's really, really easy to go wrong, so knowing how much formal training someone has can be an important indicator.