ProtBert-BFD token classification improves intrinsically disordered protein region prediction
Bibliographic record
Abstract
Abstract Background Intrinsically disordered regions (IDRs) lack stable tertiary structures yet are crucial to transcriptional regulation, signal transduction, and molecular recognition. Experimental annotation of IDRs remains limited—only ~25,000 disordered proteins are cataloged compared with hundreds of millions of known sequences—highlighting the need for scalable computational prediction. We present a systematic evaluation of classical, deep, and transformer-based models for IDR prediction and introduce a fine-tuned ProtBert-BFD token classification framework that achieves high accuracy and generalizability in data-limited settings. Methods Using manually curated annotations from the DisProt database, redundant sequences were filtered with CD-HIT (<30% similarity) and restricted to ≤526 residues, producing ~700 high-quality sequences. The dataset was partitioned (70:15:15) into training, validation, and test sets. We compared multiple architectures: (1) logistic regression and multilayer perceptron as classical baselines; (2) bidirectional and Seq2Seq LSTMs capturing sequential dependencies; and (3) ProtBert-BFD models leveraging pretrained protein-language embeddings for residue-level classification. Results Classical models showed limited predictive power (AUC 0.56–0.69), while LSTM variants improved recall but overfitted due to data imbalance. ProtBert-BFD fine-tuning substantially enhanced accuracy (75.6%), recall (68.4%), and F1-score (64.0%). The token classification variant achieved precision 0.816, recall 0.823, and F1 0.815—surpassing all internal baselines and approaching the state-of-the-art PROFbval model (recall 0.835). Average inference time was only 1.8 s per 530-residue sequence, enabling proteome-scale deployment. Conclusion These findings demonstrate that pretrained transformers effectively capture disorder-related sequence contexts and outperform conventional deep learning in low-data regimes. By integrating transfer learning with efficient inference, the ProtBert-BFD token classification model provides a robust framework for large-scale IDR annotation. Future work will expand datasets, incorporate physicochemical descriptors, and analyze error distributions to refine understanding of protein disorder in health and disease. References Hu, G., Katuwawala, A., Wang, K. et al. flDPnn: Accurate intrinsic disorder prediction with putative propensities of disorder functions. Nat Commun 12, 4438 (2021).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".