Bibliographic record
Abstract
“Evidence-based” has become the mantra of our times. Gone are the days when the adjective “scientific” was enough to grant legitimacy. Science now has to be evidence-based, as though several centuries of scientific inquiry was practiced by mere charlatans. I wonder whether the practitioners of this “newer learning” may have much to offer for those of us uncertain about the value of reference letters. Over the last 20+ years, I have written more than my fair share of such letters. In part, this stems from my involvement with problem-based learning or inquiry courses where I get to know more about students than merely their names or their grades. I have done this cheerfully and as part of my normal duties. More recently, I have moved to a province where freedom of information and privacy rules seem to weigh heavily. Students in this university have to give me their permission to write references on their behalf. This was a new experience for me, and at first I found it merely amusing. It seemed as though the university felt that I had so little work to do that I was desperately grabbing students to write references for them, even though they had not requested it. However, that amusement has given way to some reflection on the process itself and a consideration of the responsibilities of the parties involved. The more I thought about it, the less certain I became of my role, hence this commentary. Though there are three parties in this interaction, the student, the giver (letter writer), and the recipient, I will consider only the perspectives of the latter two. The student's needs are relatively simple, to get the best possible letter with minimal fuss. The situation for the giver and the recipient are more problematic. Well written reference letters provide information that cannot be gleaned from standardized scores or curriculum vitae [1], but many seem to be nothing more than testimonials of little value from a fantasy land where “No one is ever poor, fair, or average; they are all ‘very good’ or ‘excellent’ ” [2]. This land is also occupied by those who have extraordinarily high grades. Both maladies may have similar etiologies [3], at least in the North American context, “the legacy of the 1960s, an absence of clear standards, pressures to accommodate student customers, and the like. As with grade inflation, the problem is systemic; once inflationary rhetoric becomes normative, it is difficult for faculty members to do otherwise.” Those who receive letters of recommendation would like some comments that let them decide between several candidates. A highly regarded leader in business and management argues that the hiring and promoting of individuals should be on the basis of integrity, motivation, capacity, understanding, knowledge, and experience, in that order [2,4]. These are rarely evident from transcripts or curriculum vitae. The letter writers have multiple obligations: they have to ensure that the reference they give helps the students attain their goals, and they must provide useful information to the recipient and ensure that their reference does not come back to haunt their own institutions (either as lawsuits or embarrassment). A number of guidelines are available, and universities provide help to their faculty as well. A recent report [2] notes that the process of writing letters that “are truly useful to the recipients and fair to the party described is difficult.” The authors provide a list of items that are worth considering. Standardized letters of reference may be useful under some conditions [5]. A study that took a process approach showed that readers considered the tone, content, and structure of the reference letter to be significant elements in coming to decisions [6]. All these suggestions and guidelines are useful up to a point. Clearly, letters of reference must be based on evidence. But the evidence that is useful for one party may be marginal for another. Selection of information is crucial. Checklists are pointless. Comparator statements are useful, but what yardstick should one use? Above average in one situation could well be the norm elsewhere. Unless the recipient is aware of the characteristics of the group, comparator statements are as meaningless as stand-alone comments. For those statements to be useful, a more detailed description of the cohort itself must be given. One cannot write a generic letter for all occasions because the needs of the recipients must be considered. Qualities such as flair, imagination, and ability to think tangentially may be more suited for academic programs rather than strictly professional ones. So to state clearly that a student does all the right things but is not too imaginative may be an important point to mention in a reference letter for graduate school, but may not be worth putting down in a letter to a medical school. Although checklists are useful, what would really help me (as the letter writer) is some indication that my letters made a difference, i.e. whether the comments were useful, useless, or incidental. Very rarely do I get any inkling on that score. I get form letters acknowledging receipt of the letter and nothing much more. True, recipients are swamped and once the selection process is over would love to merely close the file. That does not help the giver of references write any better letters. The futility of the cycle gets perpetuated. So what I would like is a “thick” description of the process, a natural history of the recommendation letter as it were. I have tried hard to be candid in my assessments, pointing out that a particular student was far from compliant or that another was self indulgent or was acerbic. The students thus described do not appear to have suffered because they have done well. I would like to believe that my candor helped them or at least did not hinder them. I could well be wrong there. There is evidence that given the fact of letter inflation, even the hint of a negative comment can be a red flag, signaling the reader to read between the lines and gauge the writer's true feelings [6, 7]. Honesty then may be far from the best policy in this inflationary universe. Can we ever test the hypothesis that letters of recommendation are inconsequential? Consider a study that randomly allocated a population of students with good academic skills but identified major weaknesses into two groups. One group (the control) will be given glowing reference letters lauding their academic skills but not mentioning any concerns. Their fortunes will be compared with the other group who are given “truthful” letters of recommendation. The recipients obviously should be unaware and so should the students. Such a venture may never receive ethics approval or even funding, but at least it will try to gather evidence in an approved fashion (using the buzz words random, controls, blinded) and should therefore be lauded at least for effort.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".