MétaCan
Menu
Back to cohort
Record W6931840515 · doi:10.5281/zenodo.7463745

Outcomes of the sixth DELAD Workshop held on 22-23 September 2022

2022· article· en· W6931840515 on OpenAlexaboutno aff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2022
Typearticle
Languageen
FieldPsychology
TopicLanguage Development and Disorders
Canadian institutionsnot available
Fundersnot available
KeywordsSteering committeeZoomText messaging

Abstract

fetched live from OpenAlex

What is DELAD? DELAD stands for Database Enterprise for Language And speech Disorders, and is also Swedish for SHARED. DELAD is an initiative to share corpora of speech of individuals with communication disorders (CSD) among researchers. See the DELAD website. The venue was a zoom room that CLARIN kindly offered and professionally hosted. What was the workshop about? This workshop was the sixth in a row that started in 2015, and it was the fourth organised under the CLARIN umbrella. The workshop was announced via the DELAD website and via tweets and the CLARIN Newsletter. The workshop was held in the afternoons of 22 and 23 September. About 30 participants registered and attended the meeting although some of them attended only part time due to other obligations which are difficult to avoid when not physically meeting. In the meeting were 18-20 participants mostly. Participants came from all over Europe amongst others from the Netherlands, Finland, Ireland, Poland, Italy, France, Cyprus, the UK, and Canada, and had backgrounds in language and speech pathology, linguistics and phonetics, speech technology, data archiving, ICT, and law specialists. This is exactly the mix that makes DELAD attractive and suited for discussing and taking care of sharing CSD. The workshop was organised by the DELAD steering group together with Esther Hoorn from CLARIN’s CLIC. The aim of this workshop was to: Extend DELAD network with new participants Recent developments at DELAD and CLARIN K-Centre on Atypical Communication Expertise (ACE) Follow up on voice pseudonimisation Sharing clinical data: Obtaining clinical data via hospitals vs Obtaining data outside clinical institutes from alternative organisations The impact of the Data Governance Act & Data Altruism An overview of the workshop’s program can be found here. Day 1 On the first day we started out with four presentations: “Speech STAR - An outline ultrasound- tongue-imaging-based curated corpus of disordered and non-disordered speech” (Eleanor Lawson, Joanne Cleland, & Jane Stuart-Smith) “Using video‑based pose estimation for automated analysis of interaction” (Satu Saalasti) “Rethinking Language: Computational approaches to speech and language assessment for clinical studies” (Chiara Barattieri di San Pietro) Eleanor Lawson showed the website https://www.seeingspeech.ac.uk/ with UTI and MRI recordings (database). Aimed for teaching and supporting intervention and use of the methods. Speech disorder database is online with search/filter options: https://www.seeingspeech.ac.uk/speechstar/ with all permissions in place to share the data, so they can also be downloaded via Joan Cleland. Satu Salaasti showed the use of OpenPose for realtime movement estimations. Research into self and other initiated repairs in typical and atypical conversations. Chiara Barattieri di San Pietro presented the https://www.gatekeeper-project.eu/ They work with a company: ab.acus SRL SWETALY. Mentoring programme: dementia detection. After the tea break there were three more presentations: Recent developments at ACE / DELAD for support of sharing CSD (DELAD Steering group) Follow up on voice Pseudonymisation (Rob van Son) (pre-recorded presentation) (Chair: Henk van den Heuvel) “Crosslinguistic data in typical and protracted phonological development” (Barbara Bernhardt) The DELAD steering group highlighted a couple of aspects: CSD annotation tools and techniques Guidelines for consent & storage: https://delad.ruhosting.nl/wordpress/guidelines-consent-storage/ DELAD DPIA Roleplay material: https://delad.ruhosting.nl/wordpress/dpia-role-play-with-video/ The voice pseudonymisation discussion of the previous workshop was continued based on a presentation at Interspeech 2021 by Rob van Son. In her presentation Barbara Bernhardt formulated some interesting questions to DELAD about the link of DELAD, TLA and Phonbank, and the desired annotation languages for lingual resources. Day 2 Also on 23 September we started with three research presentations “Developments in sensitive data processing at CSC” (Martin Matthiesen) “The assumptions for the analysis of speech disorders and primary functions using the CARSTENS AG501 articulograph and an acoustic field distribution analyzer” (Katarzyna Klessa, Anita Lorenc, Łukasz Mik, Agnieszka Borowiec, & Daniel Król) “Licensing and distributing language resources via the Language Bank of Finland” (Mietta Lennes) Martin Matthiesen presented the SD toolbox for sensitive data which ensures that data never leaves the encrypted/safe environment. It is free to use for Finnish institutes and collaborating partners, see https://docs.csc.fi/data/sensitive-data/. Katarzyna Klessa presented the research goals and data collections for articulatory analysis using the CARSTENS AG501 articulograph, and for acoustic-phonetic analyses using the AFDA camera (Acoustic Field Distribution Analyzer). Mietta Lennes gave an overview of CLARIN licences used at the Language Bank of Finland: PUB, ACA, RES with extra features for ACA and RES licences. They are now working towards v2.1 of the licences, see https://www.kielipankki.fi/support/clarin-eula/. See also: https://www.clarin.eu/content/clarin-license-category-calculator The rest of the afternoon was devoted to legal aspects of data sharing inspired by two presentations: The impact of the Data Governance Act & Data Altruism (Pawel Kamocki) Lessons learnt from the use of the data sharing template of Dutch Academic hospitals (Anne Jan Sikkema) Pawel Kamocki gave a detailed insight into the Data Governance Act which is in force but not in application yet (24 Sep. 2023). He briefly mentioned other interesting acts: AI Act (close to being adopted), Digital services act, Common European Language Data Space. He also spent some time on introducing the Data Intermediation Services (DIS) Framework for Data Altruism European Data Innovation Board Anne Jan Sikkema, finally, presented a data sharing template of Dutch Academic hospitals which is a separate controllership agreement: supplying existing data upon request of other party for their research, see https://elsi.health-ri.nl/categorieen/verzamelen-en-uitgeven-van-data-en-lichaamsmateriaal/waar-vind-ik-een-voorbeeld-van-0. Click DTA

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.009
metaresearch head score (Gemma)0.010
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: Other
Teacher disagreement score0.151
Threshold uncertainty score0.504

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0090.010
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0060.001
Scholarly communication0.0100.003
Open science0.0020.009
Research integrity0.0030.004
Insufficient payload (model declined to judge)0.1510.090

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.033
GPT teacher head0.279
Teacher spread0.247 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicLanguage Development and DisordersFrench-language works237,207