An insufficient information-theoretic exploration of a proposed use of end-of-course survey data (#1439)
Topics/tags: Grinnell
At an upcoming faculty meeting, the Grinnell faculty will vote on a proposal on how to use our new end-of-course surveys [1] in faculty reviews. Believe it or not, I was trying not to express a strong opinion. In part, I think early-career faculty should take the lead, or at least the faculty should reflect closely on the likely concerns and preferences of the younger faculty. In part, I respect the committee members, find them thoughtful, and know how hard they worked to develop a solution that balances various interests and goals. In part, it’s that I was a member of the previous committee charged with exploring how to use end-of-course surveys in faculty reviews, and our recommendation was Don’t
. Unfortunately, that recommendation seems to have been judged unacceptable.
As I’ve noted in the past [2], I worry that end-of-course surveys (or evaluations) have two seemingly incompatible goals. On the one hand, they should inform faculty members on what areas of their course they might want to improve. On the other, they are also used to evaluate faculty. Or they might be; that’s a question we are asked to address. If I’m interested in improving my course, I’d like detailed (but sensible) critiques. If it affects my salary or my chance of retaining my job, I want only positive comments.
In any case, I found myself reflecting on a colleague’s comment at the most recent discussion of the EOCSs that I attended, which was approximately, It’s important that student voices be represented in faculty reviews.
I found myself asking, How much information do we really get from each student under the proposal?
What’s the proposal? The proposal is that for each of the six four-point Likert-scale questions in each course, department chairs, review committees, and the Personnel Committee will receive aggregated agree/disagree data.
Somewhere in the back of my head, an esteemed colleague is telling me to do an information-theoretic analysis of what this strategy implies. I’m not an information theorist, so my analysis will be inadequate [3]. For now, I’m just going to count the bits
.
If we’re reducing each answer to agree
or disagree
, we get one bit of data per answer, or six bits per student. Not very much.
But wait! The results are aggregated. So for each course, all we really need to represent is the size of the course (five bits, since most/all classes are 32 students or fewer) and the number of students who agree with each question (another five bits per question). That’s 35 bits per class. For a typical Grinnell class of sixteen students, that’s under three bits per student.
How little is three bits? Very little. If we don’t worry about efficient encoding, a typical three-letter word requires twenty-four bits. Even with careful encoding, we are unlikely to get a typical three-letter word to use fewer than eight bits. Three student responses correspond to about one three-letter word; each student is one letter. As I said, that’s very little information.
What should we do?
From my perspective, student voices should be best represented by the SEPC [4] reports. However, I must admit that those vary wildly in quality and depth, which means that they don’t necessarily represent the full array of student voices. The Dean’s surveys have a fairly low response rate, which suggests that the information is likely biased toward students who loved or hated the professor.
In the end, what the data are most likely to show are low outliers. How will we know that they are low outliers? Ooh! That’s a hard one, since we aren’t planning on gathering comparative data. It’s also not clear what comparative data we should use. We already know that, on average, students rate classes in the sciences lower than classes in the humanities, large classes [5] lower than small classes, and lower than men.
In the end, I’m not sure whether or not we use the aggregated scores really matters. Given how little information they provide (or should provide), they seem unlikely to affect decisions.
In any case, I plan to speak neither for nor against the proposals at the meeting.
Postscript: Given how little this piece conveys, one might argue that we could represent it with eight bits or fewer.
Postscript: While I will not speak in favor of nor against the proposals, I may, however, ask what the implications are if we vote No
on both proposals, particularly given that the Personnel Committee is permitted to request additional information during their deliberations.
Never mind. I’ll just stay out of it. Someone else can ask those kinds of questions.
Postscript: I didn’t want to write this. My muse forced me.
[1] We used to call them End-of-course evaluations
, but renamed them to make it clear that we were surveying students, rather than evaluating faculty.
[2] Sorry, I’m too burnt out to find the prior musing tonight.
[3] Hence the title.
[4] Student Educational Policy Committee.
[5] At Grinnell, large classes have about twenty-four students.
Version 1.0 of 2026-09-14.
