Bibliographic record
Abstract
Lupus (2007) 16, 849‐851 http://lup.sagepub.com Many years ago the late Peter Cook and Dudley Moore wrote a sketch in which the characters spoke to each other mostly in abbreviations along the lines . . . ‘Good evening JB wasn’t it you I saw driving the BMW in SW2 last week? Are you OK?’ ‘Yes, PC that was me I was in a bit of hurry; spot of D&V coming off the P&O cruise from NYC …’ and so on …. Reading the current lupus literature in relation with lupus disease assessment draws an inevitable comparison given the plethora of abbreviations and acronyms especially for disease activity indices. This situation, however, should not be regarded as wholly bad given what went before, which was frankly anarchy. As has been pointed out 1 in the twenty years or so leading up to the mid-80s some sixty attempts at defining disease activity in patients with lupus were published that were of little value since the authors did not attempt to validate them, demonstrate their reproducibility or show they could be used in large numbers of collaborating centres. I should perhaps, put my hand up here as being at least as guilty as anybody else, as three of these early attempts were mine! Things began to change in the mid-80s when groups in the UK, Toronto and Boston began to take the assessment of lupus patients seriously. The Boston and Toronto groups went along a conventional global score path attempting to develop accurate and reliable global scores that were arrived at by consensus, tested on real and paper patients and in a variety of collaborating centres. Thus Matt Liang developed the Systemic Lupus Activity Measures (SLAM) global score 1 and
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.008 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.005 | 0.007 |
| Insufficient payload (model declined to judge) | 0.005 | 0.018 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".