Interpretability of Neural Networks with Probability Density Functions
Bibliographic record
Abstract
Abstract It is an interesting topic to interpret artificial neural networks (ANNs) by considering some change various approaches. This paper explores the relationship between the input and output units of the simplest ANN, a single layer perceptron for the binary classification problem, from the probability point of view. If the feature variables of datasets follow independent normal distribution and outputs are activated by sigmoid function or smooth Relu function, we advocate that the probability density function (pdf) of the output variable is an exponential family distribution. Furthermore, by introducing an intermediate variable, the pdf of the output variable can be written as a linear combination of three normal distributions with same spread but different centers. Based on these results, the probability of the predicted class label can be written as a standard normal cumulative distribution function (cdf). The originality of this paper comes with interesting theoretical results to provide ANNs with a new description of the relationship between input variables to output variables, which can enable ANNs to be understood from a new perspective. Extensive experiments based on one artificial synthesized dataset and ten real‐world benchmark datasets validate the reasonability of those results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.035 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".