Concurrent Consideration of Human and Machine Reliability in Human-Machine Systems - A Virtual Environment Approach
Bibliographic record
Abstract
Reliability is an important concept contributing to building trust in human-machine systems (HMS). Existing studies have reported separate assessment of human and machine reliability. Thus, there is a gap on considering human and machine reliability concurrently in HMS. To fill the gap, this study investigated the feasibility of such concurrent consideration by using a virtual environment (VE) approach to simulate an HMS. In a developed VE, each human participant performed a task of exploring an invisible surface to perceive its shape, followed by his/her response to a recommendation about the shape made by the VE setting (the machine). Related to human reliability, the perception might be disrupted through a mismatch between the actual shape and force feedback delivered to the participant’s hand. Associated with machine reliability, the recommendation could be incorrect to induce a fault in the setting. Thus, the shape of the invisible surface became an instrument to combine human and machine reliability. The outcomes of the study confirmed the feasibility of combining human and machine reliability in the HMS. Moreover, human reliability might be dominant in the HMS to accomplish the task.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".