Nice, thanks for sharing. I used the same variables with the exception of dbn; my code is here https://gist.github.com/4652968 (I later started applying filters such as 4th-grade teachers only, etc., which is reflected in the gist.)
I was pretty curious how the author specifically munged his data, since then we could put to rest the speculation about the degree of correlation.
https://github.com/tmoertel/nyva-cursory
Anyway, here are the variables I used to identify comparable observations:
So I used subject, grade, school (identified by dbn) and the teacher's first and last names.