Design and in‐situ evaluation of a mixed‐initiative approach to information organization
Bibliographic record
Abstract
Organizing personal information by folders or tags has proved to be effective for finding, remembering, and understanding information. However, past studies have shown that the cost of organization can be too high for some users to be worth the effort. Mixed‐initiative approaches attempt to reduce the burden of manual organization by automatically identifying and suggesting organizational units such as folders to users. However, little is known about how such mixed‐initiative approaches influence users' organizational experiences. In this paper, we explore a mixed‐initiative approach that suggests high‐level organizational units to users to facilitate e‐mail organization. In 2 in‐situ experiments with 34 knowledge workers, we study how our mixed‐initiative approach influenced users' experience with organization. We show that our approach made it easier to create organizational units without negatively affecting recall of those units, and led to the creation of units that otherwise would have not been created. Our findings suggest ways computers and people can most effectively work together to organize information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.031 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.012 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".