David Ravenwood, click
here to update your pages on AuthorsDen.
Blogs by David Ravenwood
How Politics Affects What our Kids Learn 2/15/2015 1:34:02 PM
Excerpts from a good article written by Jason Stanford
Mute the Messenger
Share on Reddit
When Dr. Walter Stroup showed that Texas’ standardized testing regime is flawed, the testing company struck back.
by Jason Stanford Published on Wednesday, September 3, 2014, at 8:00 CST
Rebellions sometimes begin slowly, and Walter Stroup had to wait almost seven hours to start his. The setting was a legislative hearing at the Texas Capitol in the summer of 2012 at which the growing opposition to high-stakes standardized testing in Texas public schools was about to come to a head. Stroup, a University of Texas professor, was there to testify, but there was a long line of witnesses ahead of him. For hours he waited patiently, listening to everyone else struggle to explain why 15 years of standardized testing hadn’t improved schools. Stroup believed he had the answer.
Using standardized testing as the yardstick to measure our children’s educational growth wasn’t new in Texas. But in the summer of 2012 people had discovered a brand-new reason to be pissed off about it. “Rigor” was the new watchword in education policy. Testing advocates believed that more rigorous curricula and tests would boost student achievement—the “rising tide lifts all boats” theory. But that’s not how it worked out. In fact, more than a few sank. More than one-third of the statewide high school class of 2015 has already failed at least one of the newly implemented STAAR tests, disqualifying them from graduation without a successful re-test. As often happens, moms got mad. As happens less often, they got organized, and they got results.
Texas Education Commissioner Robert Scott, long an advocate of using tests to hold schools accountable, broke from orthodoxy when he called the STAAR test a “perversion of its original intent.” Almost every school board in Texas passed resolutions against over-testing, prompting Bill Hammond, a business lobbyist and leading testing advocate, to accuse school officials of “scaring” mothers. State legislators could barely step outside without hearing demands for testing relief. So in June 2012, the Texas House Public Education Committee did what elected officials do when they don’t know what to say. They held a hearing. To his credit, Committee Chair Rob Eissler began the hearing by posing a question that someone should have asked a generation ago: What exactly are we getting from these tests? And for six hours and 45 minutes, his committee couldn’t get a straight answer. Witness after witness attacked the latest standardized-testing regime that the Legislature had imposed. Everyone knew the system was broken, but no one knew exactly why.
Except for one person. Stroup, a bookishly handsome associate professor in the University of Texas College of Education, sat patiently until it was his turn to testify. Then Stroup sat down at the witness table and offered the scientific basis behind the widely held suspicion that what the tests measured was not what students have learned but how well students take tests. Every other witness got three minutes; it is a rough measure of the size of the rock that Stroup dropped into this pond that he was allowed to talk and answer lawmakers’ questions for 20 minutes.
A tenured professor at UT with a doctorate in education from Harvard University, Stroup isn’t frequently let out of the lab to address politicians in front of cameras. He talks with no evident concern that he might upset the powerful, and he speaks so quickly that his sentences have to hurry to keep up as he darts down tangents without warning. He taught in classrooms for almost a decade—it must have been a nightmare for his students.
Everyone knew the system was broken, but no one knew exactly why. Except for one person. But his testimony to the committee broke through the usual assumption that equated standardized testing with high standards. He reframed the debate over accountability by questioning whether the tests were the right tool for the job. The question wasn’t whether to test or not to test, but whether the tests measured what we thought they did.
Stroup argued that the tests were working exactly as designed, but that the politicians who mandated that schools use them didn’t understand this. In effect, Stroup had caught the government using a bathroom scale to measure a student’s height. The scale wasn’t broken or badly made. The scale was working exactly as designed. It was just the wrong tool for the job. The tests, Stroup said, simply couldn’t measure how much students learned in school.
Stroup testified that for $468 million the Legislature had bought a pile of stress and wasted time from Pearson Education, the biggest player in the standardized-testing industry. Lest anyone miss that Stroup’s message threatened Pearson’s hegemony in the accountability industry, Rep. Jimmie Don Aycock (R-Killeen) brought Stroup’s testimony to a close with a joke that made it perfectly clear. “I’d like to have you and someone from Pearson have a little debate,” Aycock said. “Would you be willing to come back?”
br> “Sure,” Stroup said. “I’ll come back and mud wrestle.”
But that never happened. Stroup had picked a fight with a special interest in front of politicians. The winner wouldn’t be determined by reason and science but by politics and power. Pearson’s real counterattack took place largely out of public view, where the company attempted to discredit Stroup’s research. Instead of a public debate, Pearson used its money and influence to engage in the time-honored academic tradition of trashing its rival’s work and career behind his back.
Stroup knew from his experience teaching impoverished students in inner-city Boston, Mexico City and North Texas that students could improve their mastery of a subject by more than 15 percent in a school year, but the tests couldn’t measure that change. Stroup came to believe that the biggest portion of the test scores that hardly changed—that 72 percent—simply measured test-taking ability. For almost $100 million a year, Texas taxpayers were sold these tests as a gauge of whether schools are doing a good job. Lawmakers were using the wrong tool.
The tests, Stroup said, simply couldn’t measure how much students had learned in school. The paradox of Texas’ grand experiment with standardized testing is that the tests are working exactly as designed from a psychometric (the term for the science of testing) perspective, but their results don’t show what policymakers think they show. Stroup concluded that the tests were 72 percent “insensitive to instruction,” a graduate- school way of saying that the tests don’t measure what students learn in the classroom.
This claim earned Stroup a rebuke from the TEA, which stated that his findings betrayed “fundamental misunderstandings” about the way tests were constructed. The idea that most of a student’s test score carries over almost automatically, with little variance, year to year, was new, but it shouldn’t have been. After three years, STAAR scores have not budged much at all, and the TEA’s own recent report on the STAAR test results largely agrees with Stroup’s finding: The state agency declared that about 58 percent of middle school test scores showed little change from year to year.
Recently, the American Statistical Association condemned the use of student test scores to rate teacher performance. In a statement last April, the association cautioned that most studies find that “teachers account for about 1% to 14% of the variability in test scores,” largely confirming Stroup’s apparently controversial conclusion.
If it’s true that the test measured primarily students’ ability to take a test, then, Stroup reasoned to the House Public Education Committee in June 2012, “it is rational game theory strategy to target the 72 percent.” That means more Pearson worksheets and fewer field trips, more multiple-choice literary analysis and fewer book reports, and weeks devoted to practice tests and less classroom time devoted to learning new things. In other words, logic explained exactly what was going on in Texas’ public schools.
When business lobbyists and legislators desired tests that measure whether a student was “college and career ready,” they didn’t dramatically reform the curriculum. They needed harder questions based on the same curriculum, a trick Pearson managed by incorporating logic puzzles into questions about knowledge.
“My son came home from the third grade, and he said, ‘You know daddy, someone is out there trying to trick me, and all I have to do is figure out how they’re tricking me,’” Stroup told the legislators. “I’m not sure if it translates all that well to society if we teach kids gaming. All right, we end up with adults and professionals spending most of their time gaming the system.”
Pearson’s real counterattack took place largely out of public view. In a cover letter to the dean, Bomer explained the change from “unsatisfactory” to “does not meet expectations” as the result of a misunderstanding about the rating definitions. No mention was made of the first draft’s supposed mistakes. In fact, Bomer wrote that the review committee did not even consider the omissions Stroup pointed out because “these things were not on his [curriculum] vita,” even though they were on the forms provided by the university.
Bomer’s cover letter indicates that the subject of Pearson came up with the review committee when Stroup requested a meeting after receiving the original unsatisfactory rating. “At that meeting, Dr. Stroup asked whether any member of the review committee or I had any relationship to Pearson publishing,” Bomer wrote. “None of us has any such relationship.”
Maybe Stroup’s “emperor has no clothes” rebellion against UT’s generous benefactor has nothing to do with his post-tenure review. For its part, Pearson Education said through a spokesperson that the company had no contact with UT about Stroup.
Maybe Stroup and his cloud computing and networked calculators don’t fit neatly into an academic world, so his colleagues think he’s slacking off. Maybe there’s another explanation for why the UT College of Education is seemingly trying to get rid of a tenured professor.
But if Pearson were trying to strike back against a researcher who told legislators that they were paying $100 million a year for tests that mostly measure test-taking ability, it would look an awful lot like what is happening to Walter Stroup.