A Comparative Study of Text-to-Image Generative Models
Bibliographic record
Abstract
Recent advancements of deep learning (DL) techniques have revolutionized various fields such as computer vision, image processing, artificial intelligence, and natural language processing. One notable application that showcases the power of these algorithms is the field of image synthesis, where new images are created from textual descriptions. Generative models play a crucial role in this process, enabling the generation of novel data based on patterns learned from the training set. Diffusion models are a distinctive class of generative models that operate by introducing random noise to existing data, and subsequently learning to reverse this diffusion process. This technique is particularly valuable in scenarios where the transformation of data over time or through sequential steps is a critical aspect of the generation process. The ability to translate textual descriptions into visual representations offers new possibilities for human-computer interaction and creative expressions. This paper provides a comparison and analysis of generative adversarial networks (GANs) and diffusion models within the domain of “text-to-image generation” to understand the strength and weaknesses of different models in specific contexts. For this purpose, we are using a combination of Vector-Quantized GAN (VQGAN) and Contrastive Language-Image Pre-training (CLIP) model. This combination provides a powerful integration of two distinct machine learning (ML) techniques for the purpose of creating images from textual input. Guided Language to Image Diffusion for Generation and Editing (GLIDE) is the diffusion model used in this study. For both models, text input from the MS-COCO data set is used. Evaluation of generated images is performed using Fréchet Inception Distance (FID) and Inception Score (IS) metrics. Semantic object accuracy score (SOA) is also used as a metric to add an additional layer for analysis by considering the relevance of the generated images to the provided captions during the image generation process. This metric is helpful not only to assessing visual quality of the generated images but also their alignment with the intended semantic content.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".