MétaCan
Menu
← Retour à la cohorte
Enregistrement W6968141960 · doi:10.5281/zenodo.13769879

openproblems-bio/openproblems: v1.0.0

2024· other· en· W6968141960 sur OpenAlexaff

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2024
Typeother
Langueen
Domaine
Thématique
Établissements canadiensTellabs (Canada)
Organismes subventionnairesnon disponible
Mots-clésRanking (information retrieval)WorkflowJSONParsingPython (programming language)MetadataComputationFunction (biology)

Résumé

récupéré en direct d'OpenAlex

Note: This changelog was automatically generated from the git log. New functionality Added cell2location to the spatial_decomposition task. Added nearest-neighbor ranking matrix computation to _utils. Datasets now store nearest-neighbor ranking matrix in adata.obsm["X_ranking"]. Added support for parsing Nextflow output and generating benchmark results for the website. Added max_samples parameter to qlocal, qglobal, qnn_auc, lcmc, qnn, and continuity metrics to allow for subsampling of data for faster computation. Added new scArches based methods: scarches_scanvi_xgb_all_genes and scarches_scanvi_xgb_hvg. Added prediction_method parameter to _scanvi_scarches to specify prediction method. Added _pred_xgb function to perform XGBoost prediction based on latent representations. Added obsm parameter to _xgboost function to allow specifying the embedding space for XGBoost training. Major changes Updated scvi-tools to version 0.20 in both Python and R environments. Updated datasets to include nearest-neighbor ranking matrix. Modified dimensionality reduction task to include nearest-neighbor ranking matrix computation in dataset generation. The website update workflow was refactored to use a new workflow using json instead of markdown. Updated the website generation process to remove duplicate BibTex entries. Added a new parse_metadata.py script for generating metadata for the website. Added a new function to openproblems.utils.py to get the member ID of a task, dataset, method or metric. Removed the redundant computation and storage of the nearest-neighbor ranking matrix in datasets. Minor changes Updated method names to be shorter and more consistent across tasks. Improved method summaries for clarity. Updated JAX and JAXlib versions to 0.4.6. Updated dependencies to support new versions of Snakemake and GitPython. Removed code related to "nbt2022-reproducibility" repo and merged it into the main website. Updated the schema for benchmark results to include submission time, code version, and resource usage metrics. Improved error handling and added logging to the parsing script. Removed the "raw.json" file from the results directory and merged all data into a single "results.json" file. Updated the workflow to upload the final results to the website's results directory instead of the data directory. Removed unnecessary code and refactored the parsing script for better readability. Added unit tests for the new parsing script. Updated the run_tests workflow to skip testing on the test_website branch. Updated the run_tests workflow to skip testing on the test_process branch. Updated the create-pull-request step to set the author for the pull request. Updated the run_tests workflow to skip testing on pull request reviews. Updated the update_website_content workflow to update the website on the main branch. Updated the main.bib file to fix a typo. Removed extraneous headings from task README files. Updated generate_test_matrix.py to use the new openproblems.utils.get_member_id function. Updated the website generation process to copy BibTex files to the correct location. Updated the process_requires section in setup.py to include gitpython. Updated git commit hash generation for openproblems functions. Modified _xgboost to allow for specifying tree_method. Modified _scanvi_scarches to consistently use unlabeled_category. Modified _scanvi_scarches to remove unnecessary copying of labels. Removed _scanvi_scarches functions that were redundant with _scanvi_scarches. Removed unused _scanvi functions. Modified _scanvi_scarches to allow for specifying prediction_method and handle unlabeled_category consistently. Documentation Improved the documentation of the auprc metric. Improved the documentation of the cell2location methods. Document sub-stub task behaviour Bug fixes Fixed an error in neuralee_default where the subsample_genes argument could be too small. Fixed an error in knn_naive where the is_baseline argument was set to False. Fixed calculation of ranking matrix in _utils to include ties. Fixed a bug in load_tenx_5k_pbmc() where a warning about non-unique variable names was being raised. Removed the unused _utils.py file. Removed the X_ranking entry from the obsm attribute of datasets. The _fit() function in nn_ranking.py now subsamples the data if max_samples is specified. The nn_ranking metrics now use subsampling in the _fit() function to improve performance. Fixed the git hash generation for openproblems functions Fixed a warning about pkg_resources being deprecated Removed unnecessary fetch-depth: 1 from workflow Fixed potential issue in _scanvi_scarches where labels_pred could be overwritten Fixed potential issue in _pred_xgb where num_round wasn't being used correctly Fixed an issue where baseline methods were not being filtered correctly from the benchmark results. Fixed an issue where metrics with all NaN values were not being removed from the benchmark results. Fixed an issue where some metrics were not being parsed correctly from the Nextflow output. Fixed an issue where the "mean_score" field was not being calculated correctly for each method. Fixed an issue where the "code_version" field was not being populated correctly for each method. Fixed an issue where the "submission_time" field was not being populated correctly for each method. Fixed an issue where the resource usage metrics were not being parsed correctly from the Nextflow output. Updated the run_tests workflow to skip testing on the test_website branch. Updated the run_tests workflow to skip testing on the test_process branch. Updated the create-pull-request step to set the author for the pull request. Updated the run_tests workflow to skip testing on pull request reviews. Updated the `update_website_ Full Changelog: https://github.com/openproblems-bio/openproblems/compare/v0.8.0...v1.0.0

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,015
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesCharge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Logiciel · Signal consensuel: Logiciel
Score de désaccord entre enseignants0,340
Score d'incertitude au seuil0,941

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0040,015
Méta-épidémiologie (sens strict)0,0040,004
Méta-épidémiologie (sens large)0,0030,005
Bibliométrie0,0030,003
Études des sciences et des technologies0,0010,001
Communication savante0,0050,005
Science ouverte0,0090,006
Intégrité de la recherche0,0030,008
Charge utile insuffisante (le modèle a refusé de juger)0,3400,407

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,040
Tête enseignante GPT0,265
Écart entre enseignants0,226 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreLogiciel

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2024
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)→Travaux en français237 207→