Bibliographic record
Abstract
pandoc (2.4) [new features] New input format man (Yan Pashkovsky, John MacFarlane). [behavior changes] --ascii is now implemented in the writers, not in Text.Pandoc.App, via the new writerPreferAscii field in WriterOptions. Now the write* functions for Docbook, HTML, ICML, JATS, LaTeX, Ms, Markdown, and OPML are sensitive to writerPreferAscii. Previously the to-ascii translation was done in Text.Pandoc.App, and thus not available to those using the writer functions directly. --ascii now works with Markdown output. HTML5 character reference entities are used. --ascii now works with LaTeX output. 100% ASCII output can't be guaranteed, but the writer will use commands like \"{a} and \l whenever possible, to avoid emiting a non-ASCII character. For HTML5 output, --ascii now uses HTML5 character reference entities rather than numerical entities. Improved detection of format based on extension (in Text.Pandoc.App). We now ensure that if someone tries to convert a file for a format that has a pandoc writer but not a reader, it won't just default to markdown. Add viz. to abbreviations file (#5007, Nick Fleisher). AsciiDoc writer: always use single-line section headers, instead of the old underline style (#5038). Previously the single-line style would be used if --atx-headers was specified, but now it is always used. RST writer: Use simple tables when possible (#4750). CommonMark (and gfm) writer: Add plain text fallbacks. (#4528, quasicomputational). Previously, the writer would unconditionally emit HTML output for subscripts, superscripts, strikeouts (if the strikeout extension is disabled) and small caps, even with raw_html disabled. Now there are plain-text (and, where possible, fancy Unicode) fallbacks for all of these corresponding (mostly) to the Markdown fallbacks, and the HTML output is only used when raw_html is enabled. Powerpoint writer: support raw openxml (Jesse Rosenthal, #4976). This allows raw openxml blocks and inlines to be used in the pptx writer. Caveats: (1) It's up to the user to write well-formed openxml. The chances for corruption, especially with such a brittle format as pptx, is high. (2) Because of the tricky way that blocks map onto shapes, if you are using a raw block, it should be the only block on a slide (otherwise other text might end up overlapping it). (3) The pptx ooxml namespace abbreviations are different from the docx ooxml namespaces. Again, it's up to the user to get it right. Unzipped document and ooxml specification should be consulted. With --katex in HTML formats, do not use the autorenderer (#4946). We no longer surround formulas with \(..\) or \[..\]. Instead, we tell katex to convert the contents of span elements with class "math". Since math has already been identified, this avoids wasted time parsing for LaTeX delimiters. Note, however, that this may yield unexpected results if you have span elements with class "math" that don't contain LaTeX math. Also, use latest version of KaTeX by default (0.9.0). The man writer now produces ASCII-only output, using groff escapes, for portability. ODT writer: Add title, author and date to metadata; any remaining metadata fields are added as meta:user-defined tags. Implement table caption numbering (#4949, Nils Carlson). Captioned tables are numbered and labeled with format "Table 1: caption", where "Table" is replaced by a translation, depending on the value of lang in metadata. Uncaptioned tables are not enumerated. OpenDocument writer: Implement figure numbering in captions (#4944, Nils Carlson). Figure captions are now numbered 1, 2, 3, … The format in the caption is "Figure 1: caption" and so on (where "Figure" is replaced by a translation, depending on the value of lang in the metadata). Captioned figures are numbered consecutively and uncaptioned figures are not enumerated. This is necessary in order for LibreOffice to generate an Illustration Index (Table of Figures) for included figures. RST reader: Pass through fields in unknown directives as div attributes (#4715). Support class and name attributes for all directives. Org reader: Add partial support for #+EXCLUDE_TAGS option. (#4284, Brian Leung). Headers with the corresponding tags should not appear in the output. Log warnings about missing title attributes now include a suggestion about how to fix the problem (#4909). Lua filter changes (Albert Krewinkel): Report traceback when an error occurs. A proper Lua traceback is added if either loading of a file or execution of a filter function fails. This should be of help to authors of Lua filters who need to debug their code. Allow access to pandoc state (#5015). Lua filters and custom writers now have read-only access to most fields of pandoc's internal state via the global variable PANDOC_STATE. Push ListAttributes via constructor (Albert Krewinkel). This ensures that ListAttributes, as present in OrderedList elements, have additional accessors (viz. start, style, and delimiter). Rename ReaderOptions fields, use snake_case. Snake case is used in most variable names, using camelCase for these fields was an oversight. A metatable is added to ensure that the old field names remain functional. Iterate over AST element fields when using pairs. This makes it possible to iterate over all ield names of an AST element by using a generic for loop with pairs`: for field_name, field_content in pairs(element) do ... end Raw table fields of AST elements should be considered an implementation detail and might change in the future. Accessing element properties should always happen through the fields listed in the Lua filter docs. Note that the iterator currently excludes the t/tag field. Ensure that MetaList elements behave like Lists. Methods usable on Lists can also be used on MetaList objects. Fix MetaList constructor (Albert Krewinkel). Passing a MetaList object to the constructor pandoc.MetaList now returns the passed list as a MetaList. This is consistent with the constructor behavior when passed an (untagged) list. Custom writers: Custom writers have access to the global variable PANDOC_DOCUMENT(Albert Krewinkel, #4957). The variable contains a userdata wrapper around the full pandoc AST and exposes two fields, meta and blocks. The field content is only marshaled on-demand, performance of scripts not accessing the fields remains unaffected. [API changes] Text.Pandoc.Options: add writerPreferAscii to WriterOptions. Text.Pandoc.Shared: Export splitSentences. This was previously duplicated in the Man and Ms writers. Add ToString typeclass (Alexander Krotov). New exported module Text.Pandoc.Filter (Albert Krewinkel). Text.Pandoc.Parsing Generalize gridTableWith to any Char Stream (Alexander Krotov). Generalize readWithM from [Char] to any Char Stream that is a ToString instance (Alexander Krotov). New exposed module Text.Pandoc.Filter (Albert Krewinkel). Text.Pandoc.XML: add toHtml5Entities. New exported module Text.Pandoc.Readers.Man (Yan Pashkovsky, John MacFarlane). Text.Pandoc.Writers.Shared Add exported functions toSuperscript and toSubscript (quasicomputational, #4528). Remove exported functions metaValueToInlines, metaValueToString. Add new exported functions lookupMetaBool, lookupMetaBlocks, lookupMetaInlines, lookupMetaString. Use these whenever possible for uniformity in writers (Mauro Bieg, #4907). (Note that removed function metaValueToInlines was in previous released versions.) Add metaValueToString. Text.Pandoc.Lua Expose more useful internals (Albert Krewinkel): runFilterFile to run a Lua filter from file; data type Global and its constructors; and setGlobals to add globals to a Lua environment. This module also contains Pushable and Peekable instances required to get pandoc's data types to and from Lua. Low-level Lua operation remain hidden in Text.Pandoc.Lua. Rename runPandocLua to runLua (Albert Krewinkel). Remove runLuaFilter, merging this into Text.Pandoc.Filter.Lua's apply (Albert Krewinkel). [bug fixes and under-the-hood improvements] Text.Pandoc.Parsing Make uri accept any stream with Char tokens (Alexander Krotov). Rewrite uri without withRaw (Alexander Krotov). Generalize parseFromString and parseFromString' to any streams with Char token (Alexander Krotov) Rewrite nonspaceChar using noneOf (Alexander Krotov) Text.Pandoc.Shared: Reimplement mapLeft using Bifunctor.first (Alexander Krotov). Text.Pandoc.Pretty: Simplify Text.Pandoc.Pretty.offset (Alexander Krotov). Text.Pandoc.App Work around HXT limitation for –syntax-definition with windows drive (#4836). Always preserve tabs for man format. We need it for tables. Split command line parsing code into a separate unexported module, Text.Pandoc.App.CommandLineOptions (Albert Krewinkel). Text.Pandoc.Readers.Roff: new unexported module for tokenizing roff documents. New unexported module Text.Pandoc.RoffChar, provided character escape tables for roff formats. Text.Pandoc.Readers.HTML: Fix htmlTag and isInlineTag to accept processing instructions (#3123, regre
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.004 | 0.003 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.317 | 0.304 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".