Identification of Fuzzy Rule-Based Models With Output Space Knowledge Guidance
Bibliographic record
Abstract
In this article, we advocate that a knowledge tidbit residing in the output space could be helpful in improving the performance (accuracy) of the fuzzy rule-based model. It states thatif two outputs are far apart from each other,it is advisable to place their corresponding inputs in different clusters when forming subspaces of the input space. Considering this knowledge guidance mechanism, we propose two different methods to partition the input space. In the first method, input data are first partitioned with the use of the standard clustering algorithm, say fuzzy C-means; here, a constructed partition matrix is reflective of the structure present in the input space. Then, the knowledge tidbit is used to adjust the entries of the original partition matrix in such a way that those input data whose corresponding output data are far apart from each other are assigned with low values of proximity. In the second method, we propose two strategies to modify the distance between input data and a prototype (cluster center) identified in the input space. The crux of this method is that if there are many input data (which, in virtue of the knowledge tidbit, are regarded as being far-apart from the input data of interest) around a certain prototype, the distance between the input data of interest and this prototype should be penalized. Thus, the membership of these input data to the prototype is reduced. The comprehensive experimental studies carried out on both synthetic and publicly available data are used to examine the usefulness of the proposed methods.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".