Bibliographic record
Abstract
This is a minor release with several bugfixes and no new features. The new version is tested for Python 3.8-3.11 (but should also work with Python 3.12). This release requires pandas≥1.5. We recommend scipy≥1.11.0. What's Changed Minor typo fix in docs by @musicinmybrain in https://github.com/raphaelvallat/pingouin/pull/329 clip r values by @remrama in https://github.com/raphaelvallat/pingouin/pull/342 fix: deprecated parameter by @bitsnaps in https://github.com/raphaelvallat/pingouin/pull/341 hotfix: CI crash in test_power_chi2 [WIP] by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/344 hotfix: plot_rm_corr crash with specific column names by @remrama in https://github.com/raphaelvallat/pingouin/pull/351 Add check for noncentrality parameters. by @agkphysics in https://github.com/raphaelvallat/pingouin/pull/347 Use pyupgrade by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/364 fix groupby.mean for only numeric values by @jajcayn in https://github.com/raphaelvallat/pingouin/pull/363 Function test fails for np.mean by @gedeck in https://github.com/raphaelvallat/pingouin/pull/380 Fix in flatten_list for Python 3.12 by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/370 corr(): fix CI95% column name in returned dataframe by @kraktus in https://github.com/raphaelvallat/pingouin/pull/382 Replace None in dataset to fix unit tests by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/388 Remove outdated + bump pandas 1.5 by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/389 Fix doctests by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/390 Fix warnings by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/391 Remove non-centrality check (solved in scipy 1.11) by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/392 Use numeric_only=True in DataFrame.corr() and cov() by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/393 Add numeric_only=True in remaining pandas functions by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/396 Release 0.5.4 by @raphaelvallat in https://github.com/raphaelvallat/pingouin/pull/397 New Contributors @musicinmybrain made their first contribution in https://github.com/raphaelvallat/pingouin/pull/329 @bitsnaps made their first contribution in https://github.com/raphaelvallat/pingouin/pull/341 @agkphysics made their first contribution in https://github.com/raphaelvallat/pingouin/pull/347 @jajcayn made their first contribution in https://github.com/raphaelvallat/pingouin/pull/363 @kraktus made their first contribution in https://github.com/raphaelvallat/pingouin/pull/382 Full Changelog: https://github.com/raphaelvallat/pingouin/compare/v0.5.3...v0.5.4
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.015 |
| Meta-epidemiology (narrow) | 0.005 | 0.004 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.008 | 0.007 |
| Research integrity | 0.002 | 0.006 |
| Insufficient payload (model declined to judge) | 0.373 | 0.496 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".