Utility of Store and Forward Teledermatology for Skin Patch Test Readings
Bibliographic record
Abstract
BACKGROUND: Teledermatology (TD) is the use of imaging technology to provide dermatology services at a distance. To date, studies assessing its application for grading skin patch test reactions have been lacking. OBJECTIVES: The aim was to compare conventional, in-person (IP) grading of skin patch test reactions with store-forward TD. METHODS: Patients undergoing patch testing to the North American Contact Dermatitis Group (NACDG) screening series were invited to participate in this repeated-measures study. Photographs of the NACDG screening series patch sites were obtained at 2 time points (48-hour and final readings). Teledermatology assessments were completed by the same staff dermatologist who performed the IP readings; 48-hour and final TD photographs were viewed at weeks 4 and 8 after the IP encounter, respectively, to prevent recall bias. Staff dermatologists were blinded to IP grading results. The main outcome was percent agreement. Eight categories of agreement were created according to possible pairings of TD and IP reading results. Three final outcome groups of "success," "indeterminate," and "failure" were defined based on clinical significance. RESULTS: One hundred one participants completed the study. There were 7070 comparison points between IP and TD final readings. Excluding negative/negative agreement, there was "success" of TD in 54% of final readings. "Indeterminate" agreement with possible clinical significance was present in 40% of final readings. There was "failure" (definite clinical significance) in 6% of final readings. CONCLUSIONS: Teledermatology may be a viable option for grading skin patch test reactions, particularly for clinicians who perform limited patch testing. However, a clinically significant "failure" rate of 6% and practical barriers to TD implementation may preclude its widespread use for skin patch testing in tertiary referral centers where large numbers of patches are tested per patient.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".