Showing posts with label standardized testing. Show all posts
Showing posts with label standardized testing. Show all posts

Tuesday, January 6, 2026

Perhaps the Stupidest F*cking Idea in the History of Stupid F*cking Ideas

Thirty-something years ago, Curmie gave a paper at a conference called The Core and the Canon: A National Debate, held at the University of North Texas.  (A revised version of that work was subsequently published in what amounted to a Proceedings volume.)  As a direct result of his participation in the conference, Curmie was invited to join the rather newly founded National Association of Scholars (NAS), only to receive a snotty and condescending rejection letter because at the time he didn’t have a PhD (he was chairing a department at an accredited college, but that was insufficient, apparently).  Today, membership is open to anyone whose credit card payment goes through.  Curmie, needless to say, is not a member.

The NAS webpage claims that the organization “upholds the standards of a liberal arts education that fosters intellectual freedom, searches for the truth, and promotes virtuous citizenship.”  As Curmie is wont to say, “if you have to tell me, it ain’t so.”  The NAS is predictably right-wing, even de facto advocating signing on to that absurd Compact for Academic Excellence in Higher Education.  But if the Compact was a really bad idea, one of the NAS’s more recent forays into the world of educational policy is a classic of pseudo-intellectuality and downright daftness.  Indeed, this is a contender for the coveted title of Stupidest Fucking Idea in the History of Stupid Fucking Ideas.  The NAS was apparently primarily responsible for drafting this nonsense.

The proposal in question is something called the Faculty Merit Act.  It would require:

… all parts of a state university system to publish every higher-education standardized test score (SAT, ACT, CRT, GRE, LSAT, MCAT, etc.) of every faculty member, as well as the standardized test score of every applicant for the faculty member’s position, of every applicant selected for a first interview, and every applicant selected for a final interview. The Act also requires the university to post the average standardized test score of the faculty in every department. It finally requires everyone in the hiring process, both applicant and administrators, to affirm under penalty of perjury that they have provided every standardized test score.

Predictably, the National Review proclaimed this “a very good idea.”  It is not.  It is not in the same universe as a good idea.

This inanity is based on a series of, shall we say, dubious speculations masquerading as facts: that universities “draft job advertisements with specializations that will ensure only radicals need apply,” that “few close observers believe that the average professor of ethnic studies is as acute as the average professor of physics,” and above all that a standardized test score, albeit that it “is only a rough proxy for academic merit,” nonetheless will “provide some measure of general intelligence.”  All of this suggests a quantification fetish, completely oblivious to the fact that such a measure is virtually meaningless… or, perhaps, not so much “oblivious” as “fully conscious of the fact but seeking to avoid admitting it.”

First off, even the NAS admits that “Some professors will have a greater ability to teach and do research than appears on a SAT score.”  Please substitute the words “virtually all” for “some” and add “or lesser” after “greater” in the previous sentence, Gentle Reader; then it will be accurate.  The idea that a 40-year-old PhD should be judged at all by how they did on a single day when they were 17 is beyond laughable.  Standardized test scores are determined by a lot of variables.  Native intelligence is one, but so are the quality of teachers a student has had up to that point, socio-economic status, whether the test-taker has taken this kind of exam before, whether they’re running a fever or just heard some bad news about a dear friend or family member… 

Curmie has written about his own experience with standardized tests several times.  He won’t link them all here; you can use the word search feature on the blog page as well as he can, Gentle Reader.  But it might be worth mentioning that Curmie did better on the GRE than on the SAT, and better on the SAT than on the PSAT.  Did he learn something between taking those exams?  Sure.  But so, presumably, did everyone else in his age group, and the competition was presumably getting tougher: in Curmie’s day, at least, everyone took the PSAT; you took the SAT if you were part of the smaller percentage of students intending to be college-bound, and the GRE only if you were looking at grad school.  Curmie did better because he’d learned how to take that kind of test, not because he’d grown appreciably in intellect.

More to the point, those scores tell us literally nothing about someone’s skillset, only about a very rough approximation of their aptitude.  (That’s the “A” in “SAT,” after all.).  Curmie actually got a perfect score on the GRE in math.  But he’s never been more qualified to teach a college-level math course (except perhaps what is euphemistically called “College Algebra”) than his colleague in the Math Department is to teach Theatre History or Acting.  Even at the most introductory collegiate level, specific disciplinary knowledge and teaching ability are both vastly more important than intelligence, even if those standardized tests really did measure the latter.

But then we get to what the NAS considers the principal benefit of their proposal: “Perhaps most importantly, this information will provide a mass of statistical information that can be used for lawsuits…. The Faculty Merit Act will provide a mass of information that can be used by plaintiffs against discriminatory colleges and universities” (emphasis added).  Were Curmie of a cynical disposition, he might suggest that the NAS’s real goal is to dismantle public post-secondary education: add to the administrivia, costing a mountain of time and resources; limit applications because there will be a lot of prospective candidates who decide it’s none of anyone’s damned business how they did on a standardized test decades ago; and, above all, open the university up to frivolous lawsuits just because some rejected candidate who would put coffee to sleep got good board scores.  Curmie hasn’t received a single offer to teach math despite his GRE score; he’s a white male and therefore a victim of woke ideology, and dammit, he’s going to sue somebody.  😉

Curmie also notes that David Randall, the director of research at the NAS who seems to be running point for this operation, does indeed have a PhD and some scholarly publications.  What he doesn’t appear to have is any experience whatsoever as a university faculty member.  Imagine Curmie’s surprise.

This proposal is particularly diabolical because some of its foundation is indeed true: there have been plenty of DEI hires that didn’t exactly work out to the benefit of the university or its students.  Are job applicants individuals or representatives of a group?  The answer, of course, is “yes.”  To the extent that they’re the former as well as the latter, the impetus for this proposition is understandable.  That doesn’t make it anything other than moronic.

Curmie first learned of this proposed legislation from Peter Greene at Curmudgucation.  Unsurprisingly, we’re in agreement.  But you might want to check out his take, too.  Let’s give him the last word, shall we? “The Faculty Merit Act is just dumb. It's a dumb idea that wants to turn dumb policy into a dumb law and some National Review editor should feel dumb for giving it any space. If this dumb bill shows its face in your state, do be sure to call out its dumbness and note that whoever attached their name to it is just not a serious person.” 

Saturday, April 24, 2021

A Response to the Cruz Amendment

Over at Ethics Alarms, Jack Marshall excoriates Senate Democrats for not supporting an amendment by Senator Ted Cruz that would deny federal funding to any college or university which “discriminates against Asian Americans in recruitment, applicant review, or admission.” Jack closes his piece with this: “There is no persuasive argument to be made in defense of those 48 Democrats. I challenge any reader here to present one.” 

The following started out as a comment to his post, but sort of kept going long past comment length.  I won’t clutter his page with this missive, but perhaps the following may be at least a partial response to his commentary...

I’ll take up that challenge, Jack. Sort of. 

That is, I would suggest a couple of reasons to be skeptical of this proposed legislation. They require a little more abstract reasoning and a little more knowledge of the way admissions operations work than the average politician is likely to be able to muster. I hasten to add that I am not defending the rationale that diversity is such an inherent good that it overrides all other considerations, but (for example) discussions of The Death of the Last Black Man in the Whole Entire World or Fires in the Mirror in my Advanced Play Analysis class have always benefited enormously from the presence of students with different demographic profiles. 

But to the matter at hand… First off, it’s difficult to assert with any authority that discrimination is actually taking place, since no one can agree on what elements of a student’s record ought to be weighted more heavily: grades? class rank? standardized test scores? extra- and co-curricular activities? recommendations? interviews with admissions personnel and/or alumni? Should legacies count? Should the ability to pay tuition without institution-funded financial aid? Should the inability to do so count? All of these elements are part of the admissions evaluation process at universities, especially those at the elite level. (At my university, a student with a certain class rank and/or board scores is automatically offered admission. But we’d ultimately end up accepting probably 95% of the students rejected by Harvard or Yale.) 

Grades are “objective”; except of course, they aren’t. Students take classes with different degrees of difficulty. Someone who gets a 90 in calculus and a 75 in phys ed is a better student than one who gets a 95 in phys ed and an 80 in algebra, but will have a lower average. Moreover, it almost goes without saying that the #20 student in a class of 400 at School X might be better or worse than the #1 student in a class of 40 at School Y. Using sports as an example: I happen to live in the small city where one of the top half dozen soccer players in US history went to high school. He was, of course, all-conference, etc. But an all-conference player in that league today might be lucky to get a scholarship to a Division II university, let alone even hope to become a star for the national team. So how do we compare, if all we know is “all conference”? 

Of course, pols of both parties seem enamored of standardized testing as the ultimate arbiter of student (or teacher) success. The closest thing to an objective means of comparing students is to give them exactly the same test under exactly the same circumstances. But even such an endeavor is problematic, even apart from such considerations as recognizing that Student X is fighting off a cold, or student Y’s mother just got diagnosed with cancer. Someone, a mortal, designed that test, and the questions are almost by definition going to favor someone who knows precisely those concepts, terms, or formulas over someone who may have mastered different, but equally important, material. I remember a question on one of those tests that asked me to figure out the Earned Run Average of a baseball pitcher. The question provided the formula, but because I knew baseball and statistics, I could glance at the multiple choice responses and determine the right answer in perhaps two seconds. Someone fully as adept at math but who didn’t know baseball would have spent a lot more time, and perhaps been rushed at the end of the exam, omitting questions or not having time to really think them through. 

So it’s perfectly possible that a student taking a standardized test would score better or worse, perhaps even appreciably so, were they to take the exam a week earlier or later. But this isn’t even the biggest problem. Those who argue for significant weight for SAT, ACT, GRE, LSAT, MCAT (etc.) scores do so based on a fanciful belief that performance on standardized tests measures skill or intelligence. It does not. In fact, they measure test-taking ability. To provide just two examples from my own personal experience as a standardized test-taker: 

I was pretty good at math when I was in high school, so in my junior year I was among the 15 or 20 students who took a standardized test under the auspices of the American Mathematical Association, or some organization like that. The test was difficult; no one could be expected to finish it in the appointed time. And it was scored like Jeopardy: harder questions were worth more points, and you literally lost points if you answered a question and got it wrong. I got a negative score! Still, I was asked back the following year, after several months of taking a class with the worst math teacher I ever had, and a couple of months with the next-worst. I adopted a different strategy… and got the highest score in the history of the school. Was I as bad as my junior year score? Nope. Was I the best math student in however long my high school had been administering that test? Not even close. 

Skip ahead a few years. I took the GRE some five or six years after the second of the two freshman-level math courses (Calculus I and Intro to Finite Math) I took in college. It was 8:00 a.m. on a Saturday. I’d had about four hours of sleep, and I was more than a little hungover. Perfect conditions, right? I got a perfect score in math. Is it plausible that I deserved such a score? I suppose so. The test I took, unlike the subject exam in math, didn’t really test anything past algebra and a little introductory probability. But no rational person would think that my math skills exceeded my verbal skills (my verbal score was good, but not perfect). 

But, as they say on the late-night infomercials, wait, there’s more. There are studies that show there is little correlation between board scores at one level and performance at the next. More to the point in the current discussion: I remember reading about (as opposed to reading) studies which concluded that the average African-American student with a certain board score would outperform the average white student with that same score. I forget the exact figures, and I’m not sure I could find them now, but the following provides the general idea: on average, a black student with a 1080 SAT score would get better grades in college than a white student with an 1100. I don’t recall that Asian-American students were included in the findings, but I would suspect that the same principles would apply. So should we be comparing the scores themselves, or their predictive value? 

Of course, multiple studies have shown that these tests tend to privilege affluent white males: precisely the stereotype of the GOP pol. Were I of cynical disposition (perish the thought!), I might be tempted to suspect that the Republicans are using Asian-American rights as a front for their attempts to codify white preference. Wait, Republicans being as virtue-signaling and contemptuous of actual reality as the Dems? Surely not! 

Be it noted: for all my distaste for standardized tests to measure, well, anything, I still use them in considering whether a particular student should be admitted into our program or receive some of our sparse scholarship money. This is different, however, from believing that a higher score on the SAT or ACT makes for an inherently better student. Taken as part of an overall assessment, a good score will do a student more good than a bad score will, but by itself it’s merely a statistic. Another sports analogy: the best shooter in the NBA will never have the highest field goal percentage, because he’s shooting a lot from 3-point range, where 40% is excellent, but there are centers and power forwards shooting well over 50% because they never take a shot from more than 5 feet from the basket. 

So if the question is whether discrimination is a good idea, the answer is of course not. But the phrasing of the amendment opens up a lot of questions without answering any of them, especially about what, exactly, constitutes discrimination. Catch phrases and gotcha politics on either side of the aisle aren’t going to solve any problems. An admissions policy which privileges class rank over SAT scores would probably benefit African-American students over Asian-Americans, but it is not inherently discriminatory, any more than a strategy that reverses those priorities would be. And certainly a procedure which attempts to predict success at the collegiate level rather than judging success at the high school level doesn’t seem to me to be problematic, although it could be called discriminatory, I suppose. 

I don’t want some federal judge to decide how my alma maters, my employer, or indeed any other college or university runs its admissions process, absent clear and convincing evidence of actual discrimination as opposed to simply that which could be construed that way. I’m not suggesting that we ought to break existing laws to advance a cause. There are no doubt cases in which the disparity in qualifications between accepted applicants of Group X and rejected applicants of Group Y is so profound that racial discrimination (against Asian-Americans, perhaps against whites) is the only plausible explanation. In such cases, the legal system does indeed need to be employed. 

What I am suggesting, however, is that the Cruz amendment smacks more of politics—“look, we got the Dems to vote against equal rights!”—than of ethics. It seems too broad a statement to ensure just administration without further clarification. In something like collegiate admissions, one person’s discrimination is simply another person’s (ethically legitimate and comprehensible) difference in priorities. 

As a career educator at the college level, I’d probably have voted against the Cruz amendment, too… although I’d be fully aware I was walking into a trap.

Thursday, July 9, 2015

When Perfection Isn't Good Enough

Curmie is way behind on his writing, which accounts for not getting to this story quicker. We all know that the GOP and the corporatist wing of the Democratic Party have joined forces to declare war on public education in this country. It’s all about “accountability” and similar catch-phrases, which are achieved only by high-stakes standardized testing.

The problems with this philosophy occur at two strata. First, the tests and the curricula with which they survive in perverse symbiosis are designed almost exclusively by non-educators with a single purpose: to increase the power, wealth, and prestige of their progenitors. No one is really interested in providing an accurate assessment of either student development or teacher competence. No, in a fetishistic desire for quantification, we’ve just allowed the Pearsons of the world to make shit up and pretend that it’s meaningful. And we’ve allowed arrogant know-nothings like Bill Gates, charlatans like Geoffrey Canada and Michelle Rhee, and corrupt bureaucrats like Arne Duncan to convince us that they care about making education better, when in fact they’re just looking to scapegoat teachers as an excuse to sell their solutions to wildly exaggerated problems.

Curmie has argued against the overuse of standardized testing in general on several occasions—here’s a piece from three years ago with internal links to several others, but there are times when the inanity of the system is very pragmatic and specific. There are the utterly stupid questions.

There’s the corruption of “cut scores,” the scores a particular state or municipality decides is acceptable. If you want to show that your snazzy new system is working, you lower the cut score; if you want to show the status quo isn’t working, you raise it. And, because you’re either a corporation trying to sell a product or a politician trying to sell an idea (and usually a product licensed by one of your campaign contributors), you do this after the scores are in. So if last year’s cut score was 60% and you want to show that your spiffy new approach is working, you lower the cut score to 50% and voilà!, a higher percentage of this year’s students passed. Or, more recently, vice versa: “Oooohhh… we need to adopt this new strategy because fewer students got a 70 this year than got a 60 last year. Education is free-fall, and we need to fix it.”

There are the teachers who are evaluated on the basis of how students they never had in class fare on standardized tests. The reigning Curmie Award winners are the administrators at Rhame Elementary School, who removed 4th grade teacher Vuola Coyle from the classroom because her students’ test scores were too high, making it difficult if not impossible for her 5th grade colleagues to demonstrate that they, too, know how to teach.

It doesn’t matter if you do.  It’s not going to be good enough.
And now we get a variation on that theme. In Florida and, Curmie fears, other places as well, teachers can be punished even for students who earn perfect scores on standardized tests. You see, if they got a perfect score last year, they should show improvement by getting better than a perfect score this year, or the teacher involved is clearly an incompetent. Well, somebody is an incompetent, but Curmie suspects it’s not the teacher.

Ah, but the mouthpiece for the state education department assures us that counties are allowed to adjust their VAM (that’s Value-Added Model, for those of you unfamiliar with corporatized educationese) to accommodate such situations. Except, of course, they don’t (or at least didn’t). Well, at least those perfect scores are so rare it really doesn’t matter. Except that there are tens of thousands of perfect scores, meaning that the chances that some teacher was penalized because one or more of his/her students wasn’t better than perfect is roughly equal to the chance that Donald Trump will say something stupid and boorish in the next 24 hours: ontological certitude, in other words.

The most significant problem here isn’t that hosts of teachers are being treated unfairly because of a glitch in the system—there’s a pretty good likelihood that it’s been fixed by now, after all. And, once middle school teacher Luke Flynt made headlines by challenging the stupid rule, it’s a reasonable surmise that other Florida teachers were alerted to the possibility.

Nor is the central issue the fact that teachers are being evaluated by a system that is rife with flaws, doesn’t measure anything worth measuring, and pays no attention whatsoever to the fact that different teachers are more effective with different students, that student populations vary—sometimes radically—from year to year or even from class to class, and completely disregards significant factors in students’ performance completely unrelated to the classroom: economic and social issues, for example. All of these criticisms are legitimate, and all of them are damning. But it’s all sort of to be expected. Regardless of what you do for a living, there are two kinds of people who are guaranteed to believe they know how to do your job better than you do: rich people and politicians… and God save us from those who are both.

The core (or Core… ) problem, however, is not that a particular district has a silly procedure in place, or that this or that teacher is being punished for circumstances that no rational person, let alone an educational professional, would consider legitimate. Those cases are awful, but they are anecdotal. The central issue, however, is systemic, and it permeates into every part of our society, not merely the education sector.

The problem is that we, the electorate, keep voting for lazy, arrogant, morally bankrupt, and utterly irresponsible politicians at every level. No sentient adult would consciously vote for the kind of idiocy that passes for public education policy in Florida (or New York, or Pennsylvania, or Texas, or…). I started to include “stupid” in the list of adjectives in the first sentence of this paragraph. The problem is, that’s not accurate. They aren’t stupid. They know better, but they’re too beholden to campaign donors, too myopic to actually contemplate the implications of proposed legislation, too interested in supporting a general concept (core curriculum, privileging good teachers over bad ones, etc.) without examining the details (which is where the Devil lurks) of a proposal. And they’re trigger-happy about their pet projects. Yes, delaying tactics are very much a part of the political arsenal of both parties. But the rush to pass legislation for the sake of appearing to solve a problem is what gives us violations of basic human rights (the PATRIOT Act), prolix and convoluted gibberish (Obamacare), and knee-jerk hysteria (the rush to bar the Confederate battle flag from national parks on one or two specially designated days a year
even at Civil War cemeteries ).

The average state legislator knows precisely nothing about education, and the average education secretary knows (and cares) less than that. But as long as we, collectively, keep electing people unfit to clean the toilets in a public school, there’s a very real sense that we get what we deserve. There is no Curmie nomination forthcoming on the issue of punishing teachers for students’ perfect scores: you can’t embarrass a profession to which you don’t belong. You can only embarrass yourself. And politicians of both parties across the country are doing a very good job of that, indeed.

NOTE: Curmie changed his mind. People who control the entire educational system of a state are Curmie-eligible.

Sunday, March 31, 2013

What Educationists Could (Really) Learn from the NFL Combine

Four events from Curmie’s (much) younger days.

1. I was maybe ten. Someone at my school had bought a contraption (an Ur-version of a modern computer) that measured not only how accurately but how quickly students responded to a series of math and vocabulary questions. I was accused of cheating (how and why would I do that on a test that didn’t count for anything?) because I answered all the questions correctly in a time that would have been considered excellent for a college student.

2. I was in college. A friend was doing a psychology experiment on the effects of caffeine in various quantities. She’d read off a number and my job was to add 17 to it. She charted response time and accuracy. Then she bought me a coke, I drank it, and we repeated the process. I don’t remember how many times we did this. I do remember that every time, I answered correctly to every question and did so in a time a quantum step or two faster than any of the math majors she tested.

3. This was also in college, but in a different venue. Another friend would read off a three digit number, then start carrying on a conversation. After an appointed length of time—30 seconds, a minute, 3 minutes—she’d ask me to repeat the numbers. I did, accurately, every time. Then we repeated the test with letters: same result.

4. Because I did my MA in England, I was actually already teaching college when I took the GRE. My preparation for the math section was minimal: I did a little brushing up on algebra, but I didn’t take any test prep courses, and I hadn’t been in a math classroom in about five years. The night before the exam, I was at a party until about 4:00 before the 8:00 a.m. test, and yes, I’d had a couple of beers. With this unimpeachable regimen, I proceeded to get a perfect score on the math section of the GRE.


I don’t know why these related memories clicked into my mind yesterday morning, but I suspect it might have something to do with the ongoing debate about standardized testing: in particular, the comments of one John Barker. I’ve made it pretty clear over the years what I think of the increasing emphasis on standardized testing and the accompanying teach-to-the-test mentality that has infested public education in recent years. That is, whereas I grant the “objectivity” of these exams, I am skeptical of their accuracy, their relevance, their potential either to measure outcomes or to predict future performance, and even—in light of the cheating scandals we know about (and the certainty that there are those we don’t)—the integrity of the process. I’ve made these points repeatedly on this blog and its predecessor over a period of nearly eight years: here, here, here, here, here, and here, for example.

Mr. Barker, the chief accountability officer for the Chicago Public Schools, thinks otherwise. Well, duh. His job is to legitimize his own salary—well over twice mine, by the way—and to pretend that his ultimate boss, the despicable Rahm Emanuel, is something other than the venal corporate meat puppet he truly is.

The Barker quotation that’s attracting all the attention from the teaching profession is this: “My philosophy has always been that if it's a good test, teach to it.” This inanity encapsulates precisely the sort of folksy pseudo-sensibility that characterizes the accountability crowd. The implicit underpinnings of this argumentation are two-fold: that every student, everywhere, ought to be learning not merely the same basic concepts, but precisely the same thing, and a smug, unspoken assertion that the “good test” in question not only exists, but is employed universally. Needless to say, these self-serving rationales have something of the aroma of merde de taureau about them.

But I want to concentrate attention on another of Barker’s pronouncements: “I was watching the NFL combine last night, and these guys are running 40-yard dashes. If they haven't trained for that particular test, they're going to have a problem. But as you run the race in that particular test, you get better and you know more about yourself.” Seriously, he said that. Look, Gentle Reader, you and I both know that the comparison is inane. I’m pretty sure we can take as given that the sprints in question aren’t designed to increase participants’ self-knowledge.

And here’s where we return to the stories of Curmie’s youth. What, after all, was determined by the fact that I had scores that were (literally) off the charts on a couple of tests? That I was some sort of genius? Hardly. I was a bright enough guy, but all that was really determined was that I can do easy math really quickly and that I employed some basic mnemonic devices in memorizing number or letter sequences, even when I wasn’t intending to do so. I recall, for example, that one of the letter sequences happened to be a friend’s initials, a fact I noticed immediately and couldn’t “unremember” even though I’d been instructed not to try too hard to get the answers correct. And, of course, as an actor, I was used to finding (or creating) connections between seemingly unique data to aid in the process of (wait for it) memorization.

How’d I do so well on the GRE? Well, for one thing, they didn’t ask anything hard: no analytical geometry or calculus or probability, much less stuff I’d never seen before. Nobody does easy math better than I do. Lots of people do hard math better than I do. If you want somebody to add 17 to a two-digit number in a hurry, I’m your man. But when you start talking about natural logarithms and second derivatives, I’d really suggest that you look elsewhere. What these tests, any or all of them, didn’t show was whether I could memorize long strings of letters and numbers, or to do higher order math: addition (or even basic algebra) is a long way from calculus or set theory.

Thus, I’d like to look at the parallelism between that scouting combine and a standardized test in a slightly different light: namely, what is done with the data collected. Those NFL scouts at the combine understand that a time in the 40 provides a single piece of objective but only marginally relevant data. No offensive tackle is going to be asked to run 40 yards on a single play, and certainly not in a straight line. Even receivers are going to be sent in motion or not, to be positioned as a wide-out or in the slot, etc. Quickness means more than speed, and pure speed without strength (also objectively measurable) and savvy (not measurable) doesn’t amount to much.

That is, whereas a scout or a general manager might be interested that this player is two-tenths of a second faster than that one for forty yards, those times will be factored in with dozens of other bits of data—some objective, some subjective—in determining whom the team should draft. Crucially, the scouts analyze everything they can about a prospect: his strength and speed, sure, but also his work ethic, sense of teamwork, flexibility (in both the physical and attitudinal senses of the term), knowledge of the techniques of the game, etc. No player will be drafted or not based solely on his speed in the 40, and no one is going to be stupid enough to judge his college coach on the basis of that number.

Moreover, whereas it may be true that a particular prospect does a little better on the test because he’s “trained for it,” and someone else is “going to have a problem,” that very fact diminishes if not completely undermines the legitimacy of the result: since no one is going to be asked to run, unobstructed, for 40 yards (or to bench press stationary weights, or whatever) in a football game, the purpose of the test is to approximate the skill set actually required by a top-notch player in a manner that is objective and at least reasonably accurate. That a player “trains” for the test so that, hypothetically, he gets a faster start from a body position he’ll never use in a game situation only distorts the test results, rendering them even less valuable than they’d already been. The analogy to test preparation services—which, assuming they work at all, of course, benefit those who can afford them at the relative expense of those who can’t—seems obvious.

Yes, it tells you something about a student that s/he does really well on some exam, the same way it tells you something about a football player if he gets a great time in the 40. The difference is that NFL teams are smart enough to know that they’re seeing only a snapshot, whereas the educationists—especially those, like the good Mr. Barker, who have apparently never spent a day as an actual educator—place increasingly higher emphasis on those isolated moments in time. Would I rather have a football player who runs the 40 in 4.5 seconds than one who runs a 4.7? All other things being equal, sure. But all other things are never equal. Never. Allow me to repeat: never.

Not only are strength, explosiveness, balance, and a host of other variables just as important as speed, but there’s one more factor that in fact occupies the very center of the discussion, although Barker may be too dim-witted to understand. A 4.7 is a great time for a lineman; a 4.5 is average (by NFL standards) for a wide receiver. That doesn’t mean the wide receiver is “better,” only different. Don’t make a quarterback try to decide between the guy who’ll make great catches and the guy who’ll keep him vertical to throw the pass at all.

Similarly, whereas standardized testing can provide useful insight into a student’s preparation, such exams not only don’t measure everything (try creating an objective test for poetry, or intellectual curiosity, or kindness), they don’t really even measure what they measure. Looked at intelligently, which is to say skeptically, they provide some useful information. But that requires both work ethic and wisdom, two attributes conspicuously absent in most educationists. Student populations differ—class to class, neighborhood to neighborhood, year to year. The data I, as a university professor, compile for our bullshit assessment tells us far more about the caliber of students we’re attracting than about either my skills or the course structure.

A grain of salt would come in handy, in other words. Otherwise, we end up with wicked fast players who can’t… you know… catch or block or tackle. And that’s no way to win.



Sunday, June 10, 2012

A Few Thoughts on High-Stakes Testing

One of the Facebook pages I “like,” “Wear Red for Ed,” posted a link to this article by Dan DeWitt, “FCAT pressure helps students and teachers to achieve,” published on the website (at least) of the Tampa Bay Times. WRFE asks, quasi-rhetorically, “Anyone care to comment on this article?” Well, yes; yes, I would.

Not that I haven’t done so, in effect, on numerous occasions in the past (in one way or another, I’ve written about one of the various forms of standardized testing here, here, here, here, here, here, and here), but on the day when my alma mater ruined the good news of giving an honorary degree to one of my heroes, Johnny Clegg, by giving one also to Teach for America pseudo-educator Wendy Kopp, another chance to stand up for real education cannot be allowed to pass by.

OK, I’m no authority on Florida’s specific version of high-stakes testing, but I’ve seen enough in other states to have a reasonable understanding of the issues. First, let’s clarify some terms. There are two kinds of high-stakes testing: one kind—the ACT, the AP and the SAT and all their respective variations on the theme—serves a useful purpose, even if the instrument is flawed. As I have said repeatedly in the past, I have a voice in how our departmental scholarship money is spent, and having some idea of how to compare students of fundamentally different backgrounds is very helpful. True, I need to be aware of the fact that scores skew towards students who are affluent, male, white, and go to good schools. I need to know, too, that cheating is far more pervasive than the testing agencies admit. But as a part of an evaluation of a prospective student’s potential, these tests are probably a net positive.

The other kind of high-stakes testing is more insidious. Students prepare all year for a single test, and their ability to pass on to the next grade level is contingent upon a satisfactory score. Worse yet, teachers and school districts are held accountable, not for whether their students actually learned anything this year, but whether they did well on that exam.

Even apart from the expense, the possibility of corruption, the vindictiveness of those who would see public education perish altogether, the tests (wait for it) are at best a partial indicator of a student’s achievement. Discounting an entire year’s work based on a single test score is like criticizing Michael Jordan for missing a single jump-shot.

More importantly, whereas scores have “improved” under FCAT, this is a meaningless statistic in all sorts of ways, not least that, in the words of a Miami Herald headline last month, “After FCAT scores plunge, state quickly lowers the passing grade.” Yeah, see, more people are passing because not enough were, so we decided that the test was too hard. Now more students are passing, so we must be doing great.

Moreover, it only stands to reason, DeWitt’s protestations to the contrary notwithstanding, that students will do better on a test when they know what is going to be on it, and what strategy to employ. I recall an incident in my own past: I’m pretty sure I’ve written about this before, so please forgive me if you’ve heard this. I was pretty good at math when I was in high school: good enough that I was one of a handful of students chosen to take some exam distributed by the Mathematical Association of America or some such organization.

I should mention here that there are three fundamental ways of scoring a multiple choice exam: the standard model (your score reflects only how many questions you get right, thereby encouraging guessing), the SAT model (your score is lowered slightly for incorrect answers, meaning that random guessing is discouraged, but if you can narrow the field a little, it’s to your advantage to guess), and the Jeopardy model (you lose as many points for a wrong answer as you get for a right one). So, hypothetically, a student who takes a 50-question test, gets 30 right, 10 wrong, and leaves 10 blank, would get 30, 27, and 20 points, respectively. Of course, this student would be compared to other students who tests were scored the same way, so a 25 on the Jeopardy model might be “better” than a 30 on the standard model.

This test was on the Jeopardy model, up to and including have more points at stake for harder questions. I took the test as a junior. I got a negative score. Yes, a negative score. I had more incorrect answers than correct ones, or at least I had more points worth of wrong responses than of right ones. A year later, having endured about ¾ of a school year with the worst (for me) math teacher I ever had and ¼ of a year with the next-worst, I took the test again. I got the highest score in the history of my school’s participation in the program. Did I know any more math as a senior than I did as a junior? Maybe a little. But what I really learned was how to approach the test: not to guess, to skim over the test to see if there were high-point questions I was confident in my ability to handle, to completely ignore low-point questions I wasn’t absolutely sure about, and so on.

This is the lesson of high-stakes testing, and I consciously employ it in reverse in my freshman-level classes at my university. I know students have been taught to look for keywords like “always” or “never” on exams, figuring that they’re keys to test-taking. Few things are “always” or “never” true, so “usually” or “sometimes” are, in Bubbletestland, nearly always (see?) the correct answer. I therefore am careful to include a question or two for which the right response is indeed “always” or “never.” I also like to go 10 or 15 questions in a row with no C’s, or with 6 B’s in a row, or whatever. It generally doesn’t take too long to convince students that understanding the material is actually a superior strategy to trying to out-think the test.

This is a struggle, however, for students who, increasingly, want to be told the “answer,” not a means of arriving at the truth. I periodically get a course evaluation from an irate student who objects to my asking questions “backward” (not “who was Marlowe?”, but “who was the most important English pre-Shakespearean playwright?”) or that I’d actually expect both a definition and an example of a term.

Corollary to this is the simple fact that test scores measure only test scores, not skills. It is more than a little troubling that when a year’s worth of classes and a day’s worth of test-taking yield different results, we seem to concede the superiority of the test-providers’ commercial product, especially given that the tests are often both written and graded by people who couldn’t get jobs as teachers. So the fact that test scores are going up, even when it’s true (unlike in Florida), means precisely that: test scores are going up. Let’s not pretend it means anything else.

Moreover, be it noted that Advanced Placement tests are an entirely different phenomenon. How many students take them, or how well they do on them, is completely unrelated to basic skills tests. The same do-well-on-the-exam mentality may prevail, but the students taking AP exams aren’t the ones who are fretting about summer school if they don’t pass the Big Scary Test at the end of the year. It’s a good thing if more students are prospering on the AP, but no one who understands education or educational testing sees much of a correlation between that process on the one hand and the FCAT and similar projects on the other, although of course the same person might embrace both strategies.

Finally, there’s the inane argument that “for these tests to have any meaning at all, good scores must be rewarded and poor ones punished. High stakes are the whole point.” Boy, does that sound better in theory than in practice. Yes, it is reasonable to have some sort of standardized testing, and it is reasonable to have scores on those tests count in some appreciable way. Scores could factor into a student’s class rank, could be used internally to compare teachers, etc. But there’s a huge caveat here: there are enormous variations in student competencies even among what would normally be described as the same population, meaning that measuring teachers or districts by student performance on these tests is even more fraught with peril than measuring students by that imperfect yardstick is.

I have occasionally taught two sections of the same class in the same semester. Not infrequently, one class is great and the other horrible, at least in relative terms. Like other universities, we are subject to our own form of legislative meddling, namely “assessment.” I need to compile statistics about how well how many of my students do on certain prescribed tasks and assignments. Compare the results of the assessments from my 2011 and 2010 Play Analysis classes without providing for context, and you’ll think I suddenly learned how to teach (after 30+ years in the classroom). Those scores sure looked a lot better. Of course, as a group, the 2011 classes were comprised of better students: more intelligent, more self-motivated, more engaged… and they came to class more often. Funny thing, they did better. There were, of course, some good students in 2010 and some lesser ones in 2011, but as groups there was no comparison.

I know this to be utterly commonplace. I know this, of course, because, unlike Jeb Bush or Dan DeWitt, I am an educator. I know that any assessment of my abilities in the classroom that doesn’t take into account what I’ve got to work with is, by definition, useless at best and counter-productive at worst. I know that a student who performs adequately but only adequately on some assessment tool might be a credit to my pedagogical skills or an indictment of them.

Similarly, a public school teacher who gives a kid a D is probably doing a better job if the kid fails the standardized test than if s/he passes. That teacher has failed to motivate the student whose skill level exceeds his/her grade; conversely, the teacher has accurately measured the commitment, intelligence, focus, etc. of the student who goes down in flames on the standardized test. Yet I know of no organized attempt—I’m sure some competent principals do this on a ad hoc basis—to compare student test scores to how they did in the actual classroom. That’s insane… which means it sounds great to the average state legislator.

There is a veneer of truth to DeWitt’s observations. But it’s only that. If you really want students who are college-ready, take it from someone who has seen three decades' worth of college freshmen: that test-driven model may seem turbo-charged, but when it comes down to the ability to actually think, trust me, there’s a Honda engine in that Porsche… and I’m not sure it isn’t from a lawnmower.



Saturday, May 12, 2012

Just When You Thought Standardized Testing Couldn't Get Any Stupider

When I first read about this story on the Ethics Alarms blog, I thought my netfriend Jack Marshall had found his way to the recreational chemicals, mistaken an article in the Onion for a real news story, or otherwise repeated silliness as if it were true. Surely, even the simultaneously vapid and arrogant pseudo-educators who run the educational testing industry wouldn’t think it appropriate to ask a 3rd-grader to reveal a secret and then detail why it was difficult to keep, right?

Ah, but they did. About 4000 8- and 9-year olds were in fact required to do exactly that on the New Jersey Assessment of Skills and Knowledge (NJ ASK) exam. Dr. Richard Goldberg, who has twin sons who were asked the question, also has far more sense than, apparently, the entirety of the New Jersey educational elite. Here’s his take:
I was kind of shocked because it was just a very–it was an outrageous question… to ask an 8-year-old, a 9-year-old to start revealing secrets in the middle of an exam—I thought was really inappropriate… these children—they want to answer the question, they want to ask it correctly, they don’t want to get a bad grade—but at the same time… think about the things a child might know–about themselves or their family.

[Whoever] put this question forward really needs to be called to account…. I find it incredible that someone could not possibly understand how dangerous or how uncomfortable a question like this might be… somebody was either very stupid or very arrogant.
Dr. Goldberg, in fact, articulates only the tip of a Titanic-sinking iceberg of reasons this question should never have made the first cut, let alone have appeared on a standardized test. Yes, it is ridiculously, hubristically, obscenely intrusive. Most of the consternation seems to be about logistics—what is a school’s responsibility if a child reports illegal activity, for example. This is because most school administrators are more concerned with covering their collective ass than in stewarding children. Yes, it’s a problem if little Johnny talks about something private. Who reads these exams? What are that person’s legal and ethical responsibilities? And so on.

But the real victims here are the children, not the parents or administrators, however much those groups might like to pretend otherwise. Child psychologist Dr. Stephen Tobias comes closer to the mark before veering off into his own exegesis on pragmatics:
I think it’s bound to cause a lot of anxiety in some kids–3rd graders tend to be very rule-governed so if an authority figure asks a question on a test—most would feel obligated to answer it… let’s say there are secrets about abuse or drug use or things like that—then what’s the responsibility of the school in terms of reporting it… I think a question like this, for the family, the child the school opens up a whole can of worms that I’m not sure people really want to deal with in this way… if the school finds out anything that might hint at abuse—they’re responsible for reporting it… you never really know how the kid is reading this question—it’s ill-advised to ask a question like this.
Dr. Tobias is getting warm, at least at the beginning. If a student reads that question, what happens? Well, s/he might just lie, make something up. This is, in practical terms, the best response because the student can blithely move on to the next question without disruption. Indeed, the test seems to encourage students to lie—any wonder plagiarism is so rife?

The student’s other option—the one faced by students who are actually developing a moral and ethical sensibility—is a sort of ethical dissonance. Even a 9-year-old knows that secrets are secrets for a reason. Yet, at the same time, s/he feels a need to be forthright: surely this Very Big Test wouldn’t be asking this question if it weren’t Really Important to Tell the Truth, right? This student is therefore conflicted, regardless of what s/he writes down, and is agitated, no doubt, for the rest of the test.

If the student doesn’t respond the way s/he thinks the test-graders expect, anxiety sets in. But if that 9-year-old dutifully tells a family secret, here comes the guilt for that lapse in promise-keeping. The good kids, in other words, are almost universally going to have mountains of stress added to the already high stakes of this kind of exam.

Such anxiety is hardly confined to small children, and hardly insignificant, especially given the consequences of these tests to students and teachers alike. I’ve had soon-to-be summa cum laude college graduates in my office telling me about how they’d panicked on a math or history test after getting “de-railed” (one student’s term) by a question early in the exam: “I just couldn’t think after that.”

This is why it isn’t enough not to count the question this time around (it was apparently only being “field-tested,” anyway) and to throw it out for subsequent exams, as is apparently being done. The entire test needs to be pitched, and, since it wasn’t the kids’ fault, they shouldn’t be required to take the exam again. Everybody passes. But every single person who signed off on this question, even as a test run, should be fired, and the testing company (Measurement, Inc. – how cute) should refund to the schools the cost of administering such slop. Governor Christie and all his colleagues in positions of political power—mostly but by no means exclusively Republicans—who think standardized testing ought to be more, not less, a part of the educational system, need to be called on the carpet.

Are we going to let the idiots who thought this even could be a legitimate question have a greater say in our educational system than teachers, principals, and other actual professionals? Or than parents and students, for that matter? The politicians need to acknowledge their culpability, and the press needs to shine the light on them. The minions at Measurement, Inc. aren’t the problem. They’re just rather stupid folks who can’t get a teaching job. The blame falls squarely on the shoulders of the politicians and educational bureaucrats who hire this gaggle of yahoos, probably because they’re cheaper than their not-quite-so-incompetent competition. This is only one of many reasons high-stakes testing is a horrible idea, but it’s one of the most objectively provable.

This question, presumably, was green-lighted by the Measurement, Inc. people, a content expert at the Department of Education (another teacher who couldn’t get hired), a teachers advisory board (guess what kind of teacher serves on this… the ones who aren’t spending their time preparing lessons and improving their skills, of course). That’s a whole lot of people, not one of whom, presumably, had sense (or courage) enough to wonder aloud whether asking children to betray normative standards of morality was a really good idea.

Facepalm.

Sunday, April 22, 2012

A Defense of Pineapples and Hares

There is copious brouhaha of late about a series of reading comprehension questions on a standardized test administered to 8th graders in New York State. I suspect that Daniel Pinkwater, the author of the story from which the exercise was adapted (but not of the piece itself), was not responsible for the headline in the New York Daily News that appears over his name, decrying “the world’s dumbest test question,” but he certainly gives the impression that he endorses the sentiment. An editorial in the same paper proclaims the selection “flunks the test of common sense,” refers to a “dumb question,” a “now-infamous question,” and a “lollapalooza of a damaging embarrassment.”

The New York Times is a little subtler: “When Pineapple Races Hare, Students Lose, Critics of Standardized Tests Say.” NPR pronounces two questions in particular as “bizarre.” New York magazine suggests the exam “gets trippy.” You get the idea.

The situation is confused even further by the fact that there are two different versions of the test question floating around. Both appear in the NPR article, linked above. NPR also points out that the Daily News changed the version that appears on its website, apparently without comment. Here’s a link to, presumably, the story exactly as it appears in the exam. I mean, if you can’t trust the authenticity of a pdf file of a document that explicitly forbids reproduction, what (and whom) can you trust?

OK, so where does that leave us? Well, there are problems, to be sure. A variation on the theme of this series of questions has apparently been making the rounds on Pearson-devised tests in other states for some time, and there has been outcry every time. It would be reasonable to suggest, therefore, that Pearson is more than a little arrogant (who knew?) in re-cycling a question that has been challenged previously. And the edits of Pinkwater’s original story show why he writes books for children and young adults, and the editor… erm… works for Pearson. Giving authorship “credit” to Pinkwater after rendering his tale less interesting and less coherent isn’t necessarily doing him any favors.

All that said, the furor is way overblown, especially if the “real” reading selection was what I presume it to be (since the other one has punctuation errors, I’m guessing it’s the fake). I say this as a staunch opponent of high-stakes testing and the accompanying teach-to-the-test mentality. The story is fine. Five of the six questions are fine, and the sixth… well, I got it right without much thought, but a). I’ve got a little more education than the average 8th-grader, b). interpreting text is what I do, and c). I wasn’t 100% confident of my answer.

One question that spawned some debate was #7: “The animals ate the pineapple most likely because they were: a). hungry b). excited c). annoyed d). amused.” There’s no mention of hunger, excitement, or amusement in the story. But the animals had tricked themselves into cheering for the pineapple in its race against the hare. One might reasonably surmise that, two hours into a race in which their chosen hero had not yet budged would lead to a little annoyance. C: annoyed. Final answer.

The question that has generated the biggest kerfuffle, however, is #8: “Which animal spoke the wisest words? a). the hare b). the moose c). the crow d). the owl.” I can see where someone reading the first version of the question to be distributed to the public (almost certainly not the one actually used on the test) would be confused, as there’s no owl at all in that story. But the no doubt real version is actually fairly clear. True, what the hare says isn’t dumb, but neither is it very profound. The moose and crow speak out of a paranoid suspicion of the Other, and are proved wrong in their beliefs. The leaves the owl, whose observation that “Pineapples don’t have sleeves” is true and both the literal and symbolic level, and is the stated moral of the story. The owl has the wisest words to say.

This is not to discount the uproar, but rather to re-contextualize it. The problem with this question isn’t that it’s “dumb,” or “bizarre,” or “trippy.” It may not be a great question, but it’s a good one. The problem is that it tests precisely what ought to be tested: not rote memorization or even mere vocabulary, but actual comprehension. Moreover, it commits the cardinal sin of being told in a rather flippant manner, with talking animals and headstrong flora. It is, in other words, even in the Pearson-ized version, kind of fun. You know, like the works of Aesop, and Aristophanes, and Jonathan Swift, and Alexander Pope, and Molière, and Lewis Carroll, and P.G. Wodehouse, and James Barrie, and….

You may color me unsurprised, Gentle Reader, that both sides of the debate over high-stakes testing, or educational accountability, or whatever you want to call it, are on the wrong side of this one. Both sides, you see, are more interested in winning political points than in being right. So those who would corporatize the educational system are using this phony issue to argue that the educationists need to be removed from power: look at the crazy stuff they want our kids to have to answer! Meanwhile, the teachers unions and similar folk revel in the same opportunity to claim that standardized tests themselves are the problem. I need hardly mention that I’m a lot closer to the latter mentality, and have been for a long time (here’s a link to a blog piece I wrote nearly seven years ago).

But the consternation over wise owls and annoyed animals isn’t justified. What has happened is in fact an indictment of the status quo, but not for the reasons being advanced by anyone I’ve seen weigh in on the issue. The problem is that anything outside the norm, anything that treats learning as a process rather than a collection of data, anything, in short, that has the slightest bit to do with actual education: this isn’t what students or teachers expect to see on tests. The T^4 (Teach To The Test) mentality cannot function in a world in which creativity, reasoning, and openness are prioritized over memorization, formula application, and test-taking strategies.

There is, moreover, nothing wrong with asking a question most appropriate to a grade level slightly above that of the students taking the test. How else are we to distinguish the truly accomplished from the merely proficient? And if identifying excellence isn’t a goal of these tests, it bloody well should be.

The crux of the situation is that, however much they bellow to the contrary, virtually everyone—teachers, students, education critics, journalists, everyone—really thinks education works best if it turns out Jeopardy champions instead of novelists, physicists, and sociologists. It’s neater, cleaner, less work, and far easier to measure. No one knows how to react when the paradigm is a little different than what’s been experienced already. “That wasn’t on the practice test” is somehow regarded as an inherently legitimate complaint.

The question about which animal speaks the wisest words is mildly problematic because of its failures as a determinant of student success: the second-best answer is a little too appealing. But ultimately, that’s not what the ululation is about. The real problem with the question in the minds of the critics is that it seeks to find out precisely what it should be seeking to find out. It demonstrates by contrast everything that is wrong with the country’s fetishistic craving for standardization in education. And that heresy cannot be tolerated.

Monday, December 26, 2011

Standardized Testing and the Myth of the Meritocracy

Curmie is a little behind on his reading, or at least at writing about his reading, so it’s only now we discuss an article that appeared some three weeks ago in the Washington Post (or at least on their website). The original post, by Marion Berry (no, not the crooked DC mayor, that was Marion Barry), provides a single anecdote which, he claims, provides “A concise summary of what’s wrong with present corporately driven education change: Decisions are being made by individuals who lack perspective and aren’t really accountable.”

Well, as I became notorious for saying to my Asian Theatre class a few years ago, yes and no.

Let’s start with Berry’s credentials. According to Valerie Strauss at WashPo, he’s a “veteran teacher, administrator, curriculum designer and author.” According to his own website, he began his teaching career in 1952. That would make him roughly 80 years old. The last time he was actually in a classroom? Can’t say, but I’m willing to bet it wasn’t in this millennium. So while Strauss may see Berry as an authority, I see—at first glance, at least—an old guy who wants to sell more books and collect more speaking fees.

I really don’t want to denigrate Mr. Berry’s credentials, especially since I think he has a lot that’s good to say. But, much as I agree with his general assessment of standardized testing, it’s important to realize that a lot of his argument here is sheer crap.

Berry’s friend, Rick Roach, oh-so-courageously (Gasp!!!) took the Florida Comprehensive Assessment Test. The version Roach took was administered to sophomores. (Roach’s identity and the test in question are revealed here.) Miraculously, Roach lived to tell the tale. Of course, Berry reveals a lot about himself in his description of what makes Roach “successful”: “His now-grown kids are well-educated. He has a big house in a good part of town. Paid-for condo in the Caribbean. Influential friends. Lots of frequent flyer miles.” Ooooohh… I’m impressed.

Anyway, here’s Roach’s commentary:
I won’t beat around the bush. The math section had 60 questions. I knew the answers to none of them, but managed to guess ten out of the 60 correctly. On the reading test, I got 62%. In our system, that’s a ‘D,’ and would get me a mandatory assignment to a double block of reading instruction.

It seems to me something is seriously wrong. I have a bachelor of science degree, two masters degrees, and 15 credit hours toward a doctorate. I help oversee an organization with 22,000 employees and a $3 billion operations and capital budget, and am able to make sense of complex data related to those responsibilities....

It might be argued that I’ve been out of school too long, that if I’d actually been in the 10th grade prior to taking the test, the material would have been fresh. But doesn’t that miss the point? A test that can determine a student’s future life chances should surely relate in some practical way to the requirements of life. I can’t see how that could possibly be true of the test I took.
At first blush, this would seem a pretty damning indictment of the test.

Trouble is, some of those questions were revealed in a subsequent post: the reading test here and the math test here. OK, we can, perhaps, criticize these questions for relevance, but not for difficulty. We’re encouraged to see how we’d do. So I checked them out. Needless to say, I got them all right. That means I’m sufficiently well-educated to survive as an above-average high school sophomore. Somehow I don’t feel the urge to add that line to my CV. I do confess that I actually had to think about one of the math questions a little, but I ask you, Gentle Reader, to remember that I haven’t been in a math class at any level in almost 37 years.

Here’s the point: if those reading questions are in fact indicative of the difficulty of the test, then Mr. Roach’s score is indeed appalling. His claim not to have known any of the math questions reveals one of two things: he is—multiple Masters degrees, well-educated children, and Caribbean condo notwithstanding—a blithering idiot, or he’s a mendacious dirtbag, willing to say anything to advance an agenda. Either way, the fact that the likes of Mr. Roach occupy any sort of leadership role in the education establishment certainly dispels any myth of a meritocracy in these matters.

Roach’s comments also include this:
If I’d been required to take those two tests when I was a 10th grader, my life would almost certainly have been very different. I’d have been told I wasn’t “college material,” would probably have believed it, and looked for work appropriate for the level of ability that the test said I had.
From what I’ve seen of Mr. Roach, I kinda wish that had happened.

I’m also impressed (not!) by this quotation from a piece by Michael Winerip in the New York Times, cited with approbation by Mr. Berry: “As of last night, 658 principals around the state [New York] had signed a letter—488 of them from Long Island, where the insurrection began—protesting the use of students’ test scores to evaluate teachers’ and principals’ performance.” I am, in fact, disturbed by using standardized testing as the sole criterion to measure, well, anything or anyone: teachers, principals, schools, and students alike. But I’m more distressed by the fact that someone writing in the NYT would employ such an egregiously misplaced modifier, and that such a preposterous grammatical error would be quoted without comment by someone who purports to have solutions to all that ails American education. Irony abounds, to say the least.

Berry and Roach are absolutely correct in two areas. First, the tests are indeed designed by educationists (often Education majors who couldn't get an actual teaching job) who are granted virtual free reign, accountable to no one. Secondly, performance on a standardized test should never completely override a student’s achievements (or lack of them) for the academic year as a whole. I’ve been thinking about both these issues for a while—here’s a blog post I wrote six and a half years ago on the subject; my views haven’t changed in the interim.

Moreover, to say that cheating on these exams is endemic is rather like saying Tim Tebow is annoying. A couple of reports on recent cases are here and here.

But this doesn’t mean that standardized tests ought to disappear. At their best, they do distinguish between students from different schools, different backgrounds, different priorities. As I said recently, “there is [a] wide disparity between schools and their populations—being at the top of a weak class might be better or worse than being in the middle of a strong one. It’s useful to have some means of comparing such students.” Yes, it would be nice if the questions were devised by actual educators, if the scorers for these exams had real credentials, if there were legitimate oversight. But it’s not the test’s problem if Mr. Roach can’t handle algebra problems I was doing in 7th grade.

For all this, thanks are due to Mssrs. Roach and Berry for calling attention to some of the inadequacies in the system. Even though I strongly suspect that both of these guys are charlatans.

Monday, November 28, 2011

More Incompetence in SAT-Land, But This One's Not All on Them

As an educator at the post-secondary level, I have deeply ambivalent feelings about standardized testing. The teach-to-the-test mentality engendered by high-stakes exams like the fetishistic stupidity propagated in my adopted state of Texas, for example, is as antithetical to real education as it is possible to be. There also are, or at least have been, serious concerns about cultural bias. Still, I won’t pretend that I don’t look pretty carefully at ACT and SAT scores when evaluating prospective students for admission into our program or scholarships.

This is a function of two things. First, there is the wide disparity between schools and their populations—being at the top of a weak class might be better or worse than being in the middle of a strong one. It’s useful to have some means of comparing such students.

Second, most recommendations written by high school teachers and administrators are utterly useless. I do understand that writing an effective, accurate recommendation takes time; I’ve written a fair number of them, after all. But in my time as coordinator of Theatre Day (that’s our tri-annual on-campus recruitment event) I’ve read recs that are a single sentence long, that are incoherent and/or ungrammatical, that praise a student in the bottom 20% of her class for excelling in the classroom (without any explanation why the little darling consistently gets C’s and D’s, including in that teacher’s classes), that talk more about the student’s parents than about him. One teacher in Dallas obviously has a template singing the praises of some generic student; she just swaps out one student name for another and sends it along. (We once had two students at the same Theatre Day with recs from her. The letters were identical except for the students’ names and the gender-specific pronouns.)

Standardized test scores aren’t the only criterion by which to measure a student’s academic skills, of course. Grades (and class rank relative to class size) count a lot; if we happen to get a well-crafted, individualized, recommendation that actually gives us insight into the student, that’s incredibly helpful. The résumé counts: both how it’s structured and what it contains. If you want to be a designer and your layout is cluttered, unimaginative, and unattractive, you’ve got some work to do. If you want to be an actor, your teacher can talk about how wonderful you are all she wants, but if you’ve never played anything but Chorus in a musical or Third Spear-Carrier from the Left, she’s really not all that impressed with your work.

Nor do minor differences in scores matter. A 550 and a 560 in Critical Reading are, for our purposes, identical. A 440 and a 650 aren’t. If you’re smart and can really handle the language, it doesn’t matter if your acting skills are marginal: a). you’re teachable and b). there are plenty of non-acting jobs in theatre (mine, for instance). On the other hand, you may audition well with carefully-chosen pieces, but if we’re not reasonably certain you can handle the work of a freshman-level Play Analysis course, we’re not going to invest our limited scholarship money on you.

Good scores matter, in other words. And whereas a fair number of high school seniors don’t seem to comprehend much else, they do get that much. All of which means there’s a lot of pressure to perform well. For those with more money than brains or morality, that means an increased temptation to cheat: specifically, to hire a similarly immoral, but intelligent, surrogate to take the test in your stead. And that is what happened, repeatedly, in an upscale suburban area of Long Island over at least a three-year period.

The New York Times reports that some 20 people have been arrested over the past two months for being at one or the other end of such transactions: either paying someone to take the exam or accepting payment to impersonate someone else in order to take the test. The group faces felony charges of scheming to defraud, as well as misdemeanor charges of falsifying business records and criminal impersonation.

The victims here are of two kinds: the middling but honest students whose chances of getting entrance to the university of their choice or receive a scholarship to do so have been compromised by a competition that cheats, and (now) every honest and gifted kid in the area, whose own good scores, the product of native intelligence and hard work, are now (alas, quite reasonably) viewed as suspect.

There are two problems here. One is that the College Board/Educational Testing Service folks (who administer the SAT), in particular, have been more than somewhat less than diligent in preventing or prosecuting cheating. An earlier NYT article quotes Bernard Kaplan, principal of Great Neck North High School, lambasting the administrators of the SAT: “The procedures E.T.S. uses to give the test are grossly inadequate in terms of security. Furthermore, E.T.S.’s response when the inevitable cheating occurs is grossly inadequate. Very simply, E.T.S. has made it very easy to cheat, very difficult to get caught.” Mr. Kaplan adds:
It is ridiculously easy to take the test for someone else. That’s why when E.T.S. says this kind of impersonation is a rare occurrence, you just have to laugh. How would they know? All they can say is they are unaware of a large number of impersonations. I’m sure, that’s true. They are most assuredly unaware.
OK, you know how this blog consistently reams high school principals and school superintendents? Credit where it’s due: Mr. Kaplan seems to be the exception that proves the rule. It was his initial investigation that led to the arrests: “I think it’s [cheating] widespread across the country,” he said Tuesday. “We were the school that stood up to it.” And to say that the security has been lax with respect to ensuring that test-takers are who they claim to be is sort of like saying that NBA power forwards tend to be large men. [Side note: sign off on the damned agreement and play ball!]

After all, Samuel Eshaghoff is accused of impersonating six different students, including a girl to take the SAT for them at $3500 a pop, earning scores up to the 97th percentile. (I guess it wouldn’t have been worth it for me to hire him, as I did better than that on my own.) And, write the Times’s Jenny Anderson and Winnie Hu, “Currently, if a score is suspect, E.T.S. investigates. If cheating is uncovered, the score is canceled and the student is permitted to get a refund and take the test again. Neither the student’s high school nor any college is notified.” Seriously? No penalty at all?

Here’s the College Board’s attitude:
You got sick and couldn’t make the test? Sorry, no refund. Need to change the date of your test more than two weeks in advance? Well, OK, but it’ll cost you $25. Oh, you’re an immoral, self-entitled little weasel who thinks rules are for other people? Sorry, sir/madam, we thought you might be an honest kid with a legitimate excuse. Of course you can have a refund. Of course we won’t notify anyone who might be interested in whether you’ve already proven that you have no intention of actually learning anything at their academic institution. Would you like us to give you the names of a few really smart, dishonest people whom we’re far too stupid to catch if they impersonate you next time?
Not being a lawyer, I don’t know whether this qualifies as “criminal negligence” according to our legal system. But in a just universe, the cretinous yahoos at the CB/ETS who decided on this policy would lose their jobs, have “unethical moron” branded into their foreheads, and be publicly pilloried. Preferably literally.

Still, the fact that everyone in the ETS hierarchy seems to be terminally incompetent does not absolve the little urchins from taking responsibility for their actions. The recent economic meltdown may have been facilitated by de-regulation, but it was caused by specific immoral, greedy wheeler-dealers. Whether what these students did was criminal may be up for debate (I can’t imagine that there’s a legitimate legal argument there, but Dickens’s Mr. Bumble is, alas, too often correct in his assertion that “the law is an ass”). If not, every college that admitted one of these lying little bastards based on false information should immediately expel those students and sue the ETS. Actually, they should do that, anyway.

The defendants’ lawyers are paid to defend their clients, and they should do so as vigorously as they can. They are not, however, entitled to spew forth such rubbish as this from Gerald McCloskey: “My feeling is that it should be handled administratively by the College Board and the school board, not criminally, especially when the county is experiencing budget issues and resources are limited to begin with.” How noble of him to be so concerned with the county’s fiscal restraints. Translation: “My client is guilty as sin. Obfuscate!” Or Melvin Roth: “I think this is overkill for this kind of thing.” Yes, stealing is only stealing when someone other than Mr. Roth’s client does it. These guys are why there are lawyer jokes.

I strongly suspect that the ETS’s estimate of perhaps 150 cases of cheating a year is off by a factor of 100 or so. So it’s up to people like me to make sure no one gets a college diploma without earning one. Time to go to work…