Bibliographic record
Abstract
Release notes Version 0.8 refactors much of the layout module. It drops the grabbit dependency, overhauls the file indexing process, and features a number of other improvements. However, changes to the public API are very minimal, and in the vast majority of cases, 0.8 should be a drop-in replacement for 0.7.*. API-breaking changes Changes to (rarely-used) BIDSLayout initialization arguments: include and exclude have been replaced with ignore and force_index. Paths passed to ignore will be ignored from indexing; paths passed to force_index will be forcibly indexed even if they are otherwise BIDS-non-compliant. force_index takes precedence over ignore. Most querying/selection methods add a new scope argument that controls scope of querying (e.g., 'raw', 'derivatives', 'all', etc.). In some cases this replaces the more limited derivatives argument. No more domains: with the grabbit removal (see below), the notion of a 'domain' has been removed. This should impact few users, but those who need to restrict indexing or querying to specific parts of a BIDS project should be able to use the scope argument more effectively. Other changes FIX: Path indexing issues in get_file() (#379) FIX: Duplicate file returns under certain conditions (#350) FIX: Pass new variable args as kwargs in split() (#386) @effigies TEST: Update naming conventions for synthetic dataset (#385) @effigies REF: The grabbit package is no longer a dependency; as a result, much of the functionality from grabbit has been ported over to pybids. REF: Required functionality from six and inflect is now bundled with pybids in bids.external, minimizing external dependencies. REF: Core modules have been reorganized. Key data structures and containers (e.g., BIDSFile, Entity, etc.) are now in a new bids.layout.core module. REF: A new Config class has been introduced to house the information found in bids.json and other layout configuration files. REF: The file-indexing process has been completely refactored. A new hierarchy of BIDSNode objects has been introduced. While this has no real impact on the public API, and isn't really intended for public consumption yet, it will in future make it easier for users to work with BIDS projects in a tree-like way, while also laying the basis for a more sensible approach to reading and accessing associated BIDS data (e.g., .tsv files). MNT: All invocations of pd.read_table have been replaced with read_csv.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.031 |
| Meta-epidemiology (narrow) | 0.004 | 0.007 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.010 | 0.014 |
| Open science | 0.012 | 0.009 |
| Research integrity | 0.004 | 0.008 |
| Insufficient payload (model declined to judge) | 0.343 | 0.467 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".