Showing posts with label high-stakes testing. Show all posts
Showing posts with label high-stakes testing. Show all posts

Monday, June 30, 2014

Arne Duncan Outdoes Himself

In my antepenultimate (to this) post, I described Secretary of Education Arne Duncan as “the worst cabinet member of the millennium (and yes, Curmie includes the likes of Alberto Gonzales and Donald Rumsfeld in that analysis).” It wasn’t always that way—I even praised him for his confrontation with the NCAA over graduation rates for athletes. But a). virtually anyone looks good by comparison to the NCAA, b). that was over four years ago, and c). give enough monkeys enough typewriters…

Since that good start, moreover, Duncan has managed to espouse positions which represent the worst of both political perspectives. An arrogant buffoon who has never actually taught a day in his life, Secretary Duncan manages to blend the union-busting, anti-teacher, corporatist Machiavellianism of the GOP with the top-heavy bureaucracies, nanny-state sensibilities, and documentation fetishes of the Democrats. He has become a self-styled Tsar, and President Obama has not only let him get away with it, he’s encouraged it. Obama’s education policy is probably no worse than Bush’s, but it’s no better, either, and that’s a rather scathing condemnation when you get right down to it.

Arne Duncan Attempts to Be Worst Cabinet Secretary Ever


But now comes a statement from Arne the Idiot that boggles the mind in its inanity—even by Duncan’s standards. In announcing a “major shift” in the way the government evaluates federally-funded special education programs, he declared that whereas most states are indeed in compliance with federal standards, including an “individualized education plan” for each student, “it is not enough for a state to be compliant if students can’t read or do math.” And it is certainly true that the dropout rate for students with disabilities is twice that for those without, and that two-thirds of students in special education programs perform below grade level in reading and math. Um… that’s why they’re in those programs, Ace.

Here’s the response of teacher and blogger Peter Greene, in a post aptly entitled “Quite Possibly the Stupidest Thing To Come Out of the US DOE”:
Arne Duncan announced that, shockingly, students with disabilities do poorly in school. They perform below level in both English and math. No, there aren’t any qualifiers attached to that. Arne is bothered that students with very low IQs, students with low function, students who have processing problems, students who have any number of impairments—these students are performing below grade level….

But who knows. Maybe Arne is on to something. Maybe blind students can’t see because nobody expects them to. Maybe the student a colleague had in class years ago, who was literally rolled into the room and propped up in a corner so that he could be “exposed” to band—maybe that child’s problems were just low expectations. Maybe IEPs are actually assigned randomly, for no reason at all….

We don't need IEPs—we need expectations and demands. We don’t need student support and special education programs—we need more testing. We don’t need consideration for the individual child’s needs—we just need to demand that the child get up to speed, learn things, and most of all TAKE THE DAMN TESTS. Because then, and only then, will we be able to make all student disabilities simply disappear.

This is just so stunningly, awesomely dumb, it’s hard to take in. Do they imagine that disabled students are just all faking, or that the specialists who diagnose these various problems are just making shit up for giggles?
If what we were discussing here was only that group of students with ADHD, dyslexia, or similar conditions, it might make a little sense to expect to see progress roughly equivalent to norms for students without those conditions. But no, we’re also talking about kids with developmental disorders so severe they can’t sit, talk, or hold a pencil to take one of Duncan’s precious high-stakes tests.

And now we get the capper, an utterance so mind-meltingly idiotic that it would embarrass Michele Bachmann: “We know that when students with disabilities are held to high expectations and have access to a robust curriculum, they excel.” Really, Arne, and where is the evidence for that assertion? Any evidence for that? You’re dealing with educators here, dude. You can’t just make shit up and think you can get away with it.

Despite the cringe-worthiness of Duncan's absurd assertion, the Secretary did manage not to be the stupidest person on the conference call. That dubious distinction went to Tennessee’s education commissioner, Kevin Huffman, who put forth the proposition that it is lack of testing, of those magical words “strong assessments,” that’s the real problem. Because mandated testing cures everything from Down Syndrome to celebral palsy, apparently.

Seriously, it’s difficult to imagine what it must be like in the universe these guys inhabit. Unfortunately, the fact that what Duncan, Huffman, and their fellow charlatans propose is utter nonsense doesn’t change the fact that there are serious implications associated with their delusional ravings.

First, tens of thousands of good and effective teachers will have their hard work demeaned by Duncan’s transcendent silliness. Second, schools, already facing budget crises across the country, will have to re-direct resources to accommodate this boondoggle. That means less money to pay teachers, to support libraries and technology centers, to underwrite gifted and talented programs, in short to, well, be a school. Third, since Duncan seems pathologically incapable of doing anything without attaching a threat to it (do it my way or lose your funding), he further alienates anyone who actually knows anything about education from both his own inanities and the DOE in general, and enhances the impression of Chicago-style politics run amok in the Obama administration.

Finally, whereas high-stakes testing of the regular student population is unnecessarily stressful, often incompetently administered, and frequently used as “evidence” of utter falsehoods, at least we can understand the impulse. As a university professor, I do often despair at how remarkably underprepared many of my students are when they arrive in my freshman classes. If testing actually worked (it generally doesn’t), at least we’d have some means of determining what they know and what they don’t—and, as I’ve said before, I do look at a prospective student’s ACT or SAT scores as part of my decision of how to vote on a scholarship application. (I’d never use those scores to evaluate a teacher or a school in any way, however.)

Here, though, the proposal makes no sense at all. There’s no possible way that testing disabled students could do any good at all, could provide any useful information, could in fact accomplish anything remotely positive. The only way this makes sense is if it’s some sort of elaborate ruse to get people like Curmie to say “testing of the regular student population isn’t so bad, because see how much worse it could be.” (Note: ain’t gonna happen Arne—regular high-stakes testing is still awful, even if this is worse.)

Either that, or Arne Duncan is off his meds.

Sunday, June 10, 2012

A Few Thoughts on High-Stakes Testing

One of the Facebook pages I “like,” “Wear Red for Ed,” posted a link to this article by Dan DeWitt, “FCAT pressure helps students and teachers to achieve,” published on the website (at least) of the Tampa Bay Times. WRFE asks, quasi-rhetorically, “Anyone care to comment on this article?” Well, yes; yes, I would.

Not that I haven’t done so, in effect, on numerous occasions in the past (in one way or another, I’ve written about one of the various forms of standardized testing here, here, here, here, here, here, and here), but on the day when my alma mater ruined the good news of giving an honorary degree to one of my heroes, Johnny Clegg, by giving one also to Teach for America pseudo-educator Wendy Kopp, another chance to stand up for real education cannot be allowed to pass by.

OK, I’m no authority on Florida’s specific version of high-stakes testing, but I’ve seen enough in other states to have a reasonable understanding of the issues. First, let’s clarify some terms. There are two kinds of high-stakes testing: one kind—the ACT, the AP and the SAT and all their respective variations on the theme—serves a useful purpose, even if the instrument is flawed. As I have said repeatedly in the past, I have a voice in how our departmental scholarship money is spent, and having some idea of how to compare students of fundamentally different backgrounds is very helpful. True, I need to be aware of the fact that scores skew towards students who are affluent, male, white, and go to good schools. I need to know, too, that cheating is far more pervasive than the testing agencies admit. But as a part of an evaluation of a prospective student’s potential, these tests are probably a net positive.

The other kind of high-stakes testing is more insidious. Students prepare all year for a single test, and their ability to pass on to the next grade level is contingent upon a satisfactory score. Worse yet, teachers and school districts are held accountable, not for whether their students actually learned anything this year, but whether they did well on that exam.

Even apart from the expense, the possibility of corruption, the vindictiveness of those who would see public education perish altogether, the tests (wait for it) are at best a partial indicator of a student’s achievement. Discounting an entire year’s work based on a single test score is like criticizing Michael Jordan for missing a single jump-shot.

More importantly, whereas scores have “improved” under FCAT, this is a meaningless statistic in all sorts of ways, not least that, in the words of a Miami Herald headline last month, “After FCAT scores plunge, state quickly lowers the passing grade.” Yeah, see, more people are passing because not enough were, so we decided that the test was too hard. Now more students are passing, so we must be doing great.

Moreover, it only stands to reason, DeWitt’s protestations to the contrary notwithstanding, that students will do better on a test when they know what is going to be on it, and what strategy to employ. I recall an incident in my own past: I’m pretty sure I’ve written about this before, so please forgive me if you’ve heard this. I was pretty good at math when I was in high school: good enough that I was one of a handful of students chosen to take some exam distributed by the Mathematical Association of America or some such organization.

I should mention here that there are three fundamental ways of scoring a multiple choice exam: the standard model (your score reflects only how many questions you get right, thereby encouraging guessing), the SAT model (your score is lowered slightly for incorrect answers, meaning that random guessing is discouraged, but if you can narrow the field a little, it’s to your advantage to guess), and the Jeopardy model (you lose as many points for a wrong answer as you get for a right one). So, hypothetically, a student who takes a 50-question test, gets 30 right, 10 wrong, and leaves 10 blank, would get 30, 27, and 20 points, respectively. Of course, this student would be compared to other students who tests were scored the same way, so a 25 on the Jeopardy model might be “better” than a 30 on the standard model.

This test was on the Jeopardy model, up to and including have more points at stake for harder questions. I took the test as a junior. I got a negative score. Yes, a negative score. I had more incorrect answers than correct ones, or at least I had more points worth of wrong responses than of right ones. A year later, having endured about ¾ of a school year with the worst (for me) math teacher I ever had and ¼ of a year with the next-worst, I took the test again. I got the highest score in the history of my school’s participation in the program. Did I know any more math as a senior than I did as a junior? Maybe a little. But what I really learned was how to approach the test: not to guess, to skim over the test to see if there were high-point questions I was confident in my ability to handle, to completely ignore low-point questions I wasn’t absolutely sure about, and so on.

This is the lesson of high-stakes testing, and I consciously employ it in reverse in my freshman-level classes at my university. I know students have been taught to look for keywords like “always” or “never” on exams, figuring that they’re keys to test-taking. Few things are “always” or “never” true, so “usually” or “sometimes” are, in Bubbletestland, nearly always (see?) the correct answer. I therefore am careful to include a question or two for which the right response is indeed “always” or “never.” I also like to go 10 or 15 questions in a row with no C’s, or with 6 B’s in a row, or whatever. It generally doesn’t take too long to convince students that understanding the material is actually a superior strategy to trying to out-think the test.

This is a struggle, however, for students who, increasingly, want to be told the “answer,” not a means of arriving at the truth. I periodically get a course evaluation from an irate student who objects to my asking questions “backward” (not “who was Marlowe?”, but “who was the most important English pre-Shakespearean playwright?”) or that I’d actually expect both a definition and an example of a term.

Corollary to this is the simple fact that test scores measure only test scores, not skills. It is more than a little troubling that when a year’s worth of classes and a day’s worth of test-taking yield different results, we seem to concede the superiority of the test-providers’ commercial product, especially given that the tests are often both written and graded by people who couldn’t get jobs as teachers. So the fact that test scores are going up, even when it’s true (unlike in Florida), means precisely that: test scores are going up. Let’s not pretend it means anything else.

Moreover, be it noted that Advanced Placement tests are an entirely different phenomenon. How many students take them, or how well they do on them, is completely unrelated to basic skills tests. The same do-well-on-the-exam mentality may prevail, but the students taking AP exams aren’t the ones who are fretting about summer school if they don’t pass the Big Scary Test at the end of the year. It’s a good thing if more students are prospering on the AP, but no one who understands education or educational testing sees much of a correlation between that process on the one hand and the FCAT and similar projects on the other, although of course the same person might embrace both strategies.

Finally, there’s the inane argument that “for these tests to have any meaning at all, good scores must be rewarded and poor ones punished. High stakes are the whole point.” Boy, does that sound better in theory than in practice. Yes, it is reasonable to have some sort of standardized testing, and it is reasonable to have scores on those tests count in some appreciable way. Scores could factor into a student’s class rank, could be used internally to compare teachers, etc. But there’s a huge caveat here: there are enormous variations in student competencies even among what would normally be described as the same population, meaning that measuring teachers or districts by student performance on these tests is even more fraught with peril than measuring students by that imperfect yardstick is.

I have occasionally taught two sections of the same class in the same semester. Not infrequently, one class is great and the other horrible, at least in relative terms. Like other universities, we are subject to our own form of legislative meddling, namely “assessment.” I need to compile statistics about how well how many of my students do on certain prescribed tasks and assignments. Compare the results of the assessments from my 2011 and 2010 Play Analysis classes without providing for context, and you’ll think I suddenly learned how to teach (after 30+ years in the classroom). Those scores sure looked a lot better. Of course, as a group, the 2011 classes were comprised of better students: more intelligent, more self-motivated, more engaged… and they came to class more often. Funny thing, they did better. There were, of course, some good students in 2010 and some lesser ones in 2011, but as groups there was no comparison.

I know this to be utterly commonplace. I know this, of course, because, unlike Jeb Bush or Dan DeWitt, I am an educator. I know that any assessment of my abilities in the classroom that doesn’t take into account what I’ve got to work with is, by definition, useless at best and counter-productive at worst. I know that a student who performs adequately but only adequately on some assessment tool might be a credit to my pedagogical skills or an indictment of them.

Similarly, a public school teacher who gives a kid a D is probably doing a better job if the kid fails the standardized test than if s/he passes. That teacher has failed to motivate the student whose skill level exceeds his/her grade; conversely, the teacher has accurately measured the commitment, intelligence, focus, etc. of the student who goes down in flames on the standardized test. Yet I know of no organized attempt—I’m sure some competent principals do this on a ad hoc basis—to compare student test scores to how they did in the actual classroom. That’s insane… which means it sounds great to the average state legislator.

There is a veneer of truth to DeWitt’s observations. But it’s only that. If you really want students who are college-ready, take it from someone who has seen three decades' worth of college freshmen: that test-driven model may seem turbo-charged, but when it comes down to the ability to actually think, trust me, there’s a Honda engine in that Porsche… and I’m not sure it isn’t from a lawnmower.



Sunday, April 22, 2012

A Defense of Pineapples and Hares

There is copious brouhaha of late about a series of reading comprehension questions on a standardized test administered to 8th graders in New York State. I suspect that Daniel Pinkwater, the author of the story from which the exercise was adapted (but not of the piece itself), was not responsible for the headline in the New York Daily News that appears over his name, decrying “the world’s dumbest test question,” but he certainly gives the impression that he endorses the sentiment. An editorial in the same paper proclaims the selection “flunks the test of common sense,” refers to a “dumb question,” a “now-infamous question,” and a “lollapalooza of a damaging embarrassment.”

The New York Times is a little subtler: “When Pineapple Races Hare, Students Lose, Critics of Standardized Tests Say.” NPR pronounces two questions in particular as “bizarre.” New York magazine suggests the exam “gets trippy.” You get the idea.

The situation is confused even further by the fact that there are two different versions of the test question floating around. Both appear in the NPR article, linked above. NPR also points out that the Daily News changed the version that appears on its website, apparently without comment. Here’s a link to, presumably, the story exactly as it appears in the exam. I mean, if you can’t trust the authenticity of a pdf file of a document that explicitly forbids reproduction, what (and whom) can you trust?

OK, so where does that leave us? Well, there are problems, to be sure. A variation on the theme of this series of questions has apparently been making the rounds on Pearson-devised tests in other states for some time, and there has been outcry every time. It would be reasonable to suggest, therefore, that Pearson is more than a little arrogant (who knew?) in re-cycling a question that has been challenged previously. And the edits of Pinkwater’s original story show why he writes books for children and young adults, and the editor… erm… works for Pearson. Giving authorship “credit” to Pinkwater after rendering his tale less interesting and less coherent isn’t necessarily doing him any favors.

All that said, the furor is way overblown, especially if the “real” reading selection was what I presume it to be (since the other one has punctuation errors, I’m guessing it’s the fake). I say this as a staunch opponent of high-stakes testing and the accompanying teach-to-the-test mentality. The story is fine. Five of the six questions are fine, and the sixth… well, I got it right without much thought, but a). I’ve got a little more education than the average 8th-grader, b). interpreting text is what I do, and c). I wasn’t 100% confident of my answer.

One question that spawned some debate was #7: “The animals ate the pineapple most likely because they were: a). hungry b). excited c). annoyed d). amused.” There’s no mention of hunger, excitement, or amusement in the story. But the animals had tricked themselves into cheering for the pineapple in its race against the hare. One might reasonably surmise that, two hours into a race in which their chosen hero had not yet budged would lead to a little annoyance. C: annoyed. Final answer.

The question that has generated the biggest kerfuffle, however, is #8: “Which animal spoke the wisest words? a). the hare b). the moose c). the crow d). the owl.” I can see where someone reading the first version of the question to be distributed to the public (almost certainly not the one actually used on the test) would be confused, as there’s no owl at all in that story. But the no doubt real version is actually fairly clear. True, what the hare says isn’t dumb, but neither is it very profound. The moose and crow speak out of a paranoid suspicion of the Other, and are proved wrong in their beliefs. The leaves the owl, whose observation that “Pineapples don’t have sleeves” is true and both the literal and symbolic level, and is the stated moral of the story. The owl has the wisest words to say.

This is not to discount the uproar, but rather to re-contextualize it. The problem with this question isn’t that it’s “dumb,” or “bizarre,” or “trippy.” It may not be a great question, but it’s a good one. The problem is that it tests precisely what ought to be tested: not rote memorization or even mere vocabulary, but actual comprehension. Moreover, it commits the cardinal sin of being told in a rather flippant manner, with talking animals and headstrong flora. It is, in other words, even in the Pearson-ized version, kind of fun. You know, like the works of Aesop, and Aristophanes, and Jonathan Swift, and Alexander Pope, and Molière, and Lewis Carroll, and P.G. Wodehouse, and James Barrie, and….

You may color me unsurprised, Gentle Reader, that both sides of the debate over high-stakes testing, or educational accountability, or whatever you want to call it, are on the wrong side of this one. Both sides, you see, are more interested in winning political points than in being right. So those who would corporatize the educational system are using this phony issue to argue that the educationists need to be removed from power: look at the crazy stuff they want our kids to have to answer! Meanwhile, the teachers unions and similar folk revel in the same opportunity to claim that standardized tests themselves are the problem. I need hardly mention that I’m a lot closer to the latter mentality, and have been for a long time (here’s a link to a blog piece I wrote nearly seven years ago).

But the consternation over wise owls and annoyed animals isn’t justified. What has happened is in fact an indictment of the status quo, but not for the reasons being advanced by anyone I’ve seen weigh in on the issue. The problem is that anything outside the norm, anything that treats learning as a process rather than a collection of data, anything, in short, that has the slightest bit to do with actual education: this isn’t what students or teachers expect to see on tests. The T^4 (Teach To The Test) mentality cannot function in a world in which creativity, reasoning, and openness are prioritized over memorization, formula application, and test-taking strategies.

There is, moreover, nothing wrong with asking a question most appropriate to a grade level slightly above that of the students taking the test. How else are we to distinguish the truly accomplished from the merely proficient? And if identifying excellence isn’t a goal of these tests, it bloody well should be.

The crux of the situation is that, however much they bellow to the contrary, virtually everyone—teachers, students, education critics, journalists, everyone—really thinks education works best if it turns out Jeopardy champions instead of novelists, physicists, and sociologists. It’s neater, cleaner, less work, and far easier to measure. No one knows how to react when the paradigm is a little different than what’s been experienced already. “That wasn’t on the practice test” is somehow regarded as an inherently legitimate complaint.

The question about which animal speaks the wisest words is mildly problematic because of its failures as a determinant of student success: the second-best answer is a little too appealing. But ultimately, that’s not what the ululation is about. The real problem with the question in the minds of the critics is that it seeks to find out precisely what it should be seeking to find out. It demonstrates by contrast everything that is wrong with the country’s fetishistic craving for standardization in education. And that heresy cannot be tolerated.