A new probabilistic method to identify fire-igniting lightning events
Bibliographic record
Abstract
Lightning-induced ignitions play a major role shaping the frequency, patterns and characteristics of wildfires in several regions across the globe, including extreme wildfire events (e.g., Góis wildfire in 2017 in Portugal) and fire seasons, such as 2019-20 in Australia, 2020 in California, and 2023 in Canada. The attention to lightning-ignited wildfires has been growing in recent years. Studies on LIWs frequently associate lightning and wildfire data to discern or approximate the place and moment of fire ignition. This typically requires to select the lightning strike responsible for the ignition.Currently, several methods are applied to select the most likely lightning strike causing the ignition. However, this selection is complicated by, at least, two aspects. First, the spatial uncertainty of fire and lightning data (e.g., the location errors of detected lightning events). Second, the holdover phenomenon. Holdover time, commonly defined as the time between lightning-induced fire ignition and fire detection, can range from a few minutes to several days, and more rarely to some weeks or even months. Long holdover times are associated to the presence of a smoldering phase that hinders the detection of these lightning fires.Here, we present a novel method that uses location accuracy information from lightning location networks, as well as expected distributions of holdover time, to assess the probabilities of lightning igniting wildfires. Our method computes a probability metric, which is the product of two independent probabilities: a spatial and a temporal probability. The spatial component assesses the probability of a cloud-to-ground lightning event striking within a given area surrounding the fire discovery point, while the temporal component evaluates the probability of a lightning-ignited wildfire undergoing a certain holdover time. The lightning event with the maximum probability metric value is then selected as the most likely ignition source. We applied this method in three study areas: Switzerland, Catalonia (Spain), and California and Nevada (USA). The results were compared with lightning selections identified by the index of proximity, one of the currently most common methods to select the most likely ignition source of lightning-induced wildfires.The initial results indicate that the probability metric yields a different selection of lightning events, in comparison with the index of proximity, for a great proportion of wildfires, with considerable differences across the study areas. We suggest that the probability metric provides a solid alternative to current methods. The probability metric offers some advantages: (1) it simplifies some methodological decisions despite the need for additional computations; (2) it is flexible and can be adapted to different types of lightning and fire data (e.g., fire perimeters); (3) it has a more robust theoretical basis than current methods; and (4) the lightning selection can be enhanced over time due to continuous improvements in lightning and fire databases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".