Attribute Controlled Dialogue Prompting
Bibliographic record
Abstract
Prompt-tuning has become an increasingly popular parameter-efficient method for adapting large pretrained language models to downstream tasks.However, both discrete prompting and continuous prompting assume fixed prompts for all data samples within a task, neglecting the fact that inputs vary greatly in some tasks such as open-domain dialogue generation.In this paper, we present a novel, instancespecific prompt-tuning algorithm for dialogue generation.Specifically, we generate prompts based on instance-level control code, rather than the conversation history, to explore their impact on controlled dialogue generation.Experiments on popular open-domain dialogue datasets, evaluated on both automated metrics and human evaluation, demonstrate that our method is superior to prompting baselines and comparable to fine-tuning with only 5%-6% of total parameters. * Work done during an internship at Huawei.overhead.We present results on both intent and persona controlled dialogue.2 Related Work GPT-3 (Brown et al., 2020) introduces prompting, a method to steer a frozen PLM by transforming inputs into cloze-style phrases with task description and some task examples.Though it is memoryefficient since one single copy of the PLM can be shared across different tasks, the model's performance is largely restricted by the maximum conditional input length, the model size and manual guesswork for prompts (Zhao et al., 2021; Schick and Schütze, 2021a,b;Jiang et al., 2020).Other works focus on automatically searching for better discrete prompts (Jiang et al., 2020;Shin et al., 2020;Gao et al., 2021;Ben-David et al., 2021).Recently, there has been an increased interest in continuous prompts / prompt-tuning, which bridges the gap between prompting and fine-tuning, while remaining efficient during training (Lester et al., 2021;Li and Liang, 2021;Liu et al., 2021Liu et al., , 2022)).Continuous prompts extend prompt selection to the entire space of embeddings, including vector embeddings that do not correspond to any humaninterpretable natural language tokens.Hence, soft prompts are more expressive than discrete prompts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.019 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.011 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".