Evaluating Approximations and Heuristic Measures of Integrated Information
Bibliographic record
Abstract
Integrated information theory (IIT) proposes a measure of integrated information (Φ) to capture the level of consciousness for a physical system in a given state. Unfortunately, calculating Φ itself is currently only possible for very small model systems, and far from computable for the kinds of systems typically associated with consciousness (brains). Here, we consider several proposed measures and computational approximations, some of which can be applied to larger systems, and test if they correlate well with Φ. While these measures and approximations capture intuitions underlying IIT and some have had success in practical applications, it has not been shown that they actually quantify the type of integrated information specified by the latest version of IIT. In this study, we evaluated these approximations and heuristic measures, based not on practical or clinical considerations, but rather based on how well they estimate the Φ values of model systems. To do this, we simulated networks consisting of 3–6 binary linear threshold nodes randomly connected with excitatory and inhibitory connections. For each system, we then constructed the system’s state transition probability matrix (TPM), as well as its state transition matrix (STM) over time for all possible initial states. From these matrices, we calculated, approximations to Φ, and measures based on state differentiation, state entropy, state uniqueness, and integrated information. All measures were correlated with Φ in a state dependent and state independent manner. Our findings suggest that Φ can be approximated closely in small binary systems by using one or more of the readily available approximations (r > 0.95), but without major reductions in computational demands. Furthermore, Φ correlated strongly with measures of signal complexity (LZ, rs = 0.722), decoder based integrated information (Φ*, rs = 0.816), and state differentiation (D1, rs = 0.827), on the system level (state independent). These measures could allow for efficient estimation of Φ on a group level, or as accurate predictors of low, but not high, Φ systems. While it’s uncertain whether the results extend to larger systems or systems with other dynamics, we stress the importance that measures aimed at being practical alternatives to Φ are at a minimum rigorously tested in an environment where the ground truth can be established.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.144 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.006 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".