Prediction of KIR3DL1/Human Leukocyte Antigen binding
Bibliographic record
Abstract
Abstract KIR3DL1 is a polymorphic inhibitory Natural Killer (NK) cell receptor that recognizes Human Leukocyte Antigen (HLA) class I allotypes that contain the Bw4 motif. Structural analyses have shown that in addition to residues 77-83 that span the Bw4 motif, polymorphism at other sites throughout the HLA molecule can influence the interaction with KIR3DL1. Given the extensive polymorphism of both KIR3DL1 and HLA class I, we built a machine learning prediction model to describe the influence of allotypic variation on the binding of KIR3DL1 to HLA class I. Nine KIR3DL1 tetramers were screened for reactivity against a panel of HLA class I molecules which revealed different patterns of specificity for each KIR3DL1 allotype. Separate models were trained for each of KIR3DL1 allotypes based on the full amino sequence of exons 2 and 3 encoding the α 1 and α 2 domains of the class I HLA allotypes, the set of polymorphic positions that span the Bw4 motif, or the positions that encode α 1 and α 2 but exclude the connecting loops. The Multi-Label-Vector-Optimization (MLVO) model trained on all alpha helix positions performed best with AUC scores ranging from 0.74 to 0.974 for the 9 KIR3DL1 allotype models. We show that a binary division into binder and non-binder is not precise, and that intermediate levels exist. Using the same models, within the binder group, high- and low-binder categories can also be predicted, the regions in HLA affecting the high vs low binder being completely distinct from the classical Bw4 motif. We further show that these positions affect binding affinity in a nonadditive way and induce deviations from linear models used to predict interaction strength. We propose that this approach should be used in lieu of simpler binding models based on a single HLA motif.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".