How to calculate forecast accuracy for stocked items with a lumpy demand : A case study at Alfa Laval
Bibliographic record
Abstract
Inventory management is an important part of a good functioning logistic. Nearly all the literature on optimal inventory management uses criteria of cost minimization and profit maximization. To have a well functioning forecasting system it is important to have a balance in the inventory. But, it exist different factors that can results in uncertainties and difficulties to maintain this balance. One important factor is the customers’ demand. Over half of the stocked items are in stock to prevent irregular orders and an uncertainty demand. The customers’ demand can be categorized into four categories: Smooth, Erratic, Intermittent and Lumpy. Items with a lumpy demand i.e. the items that are both intermittent and erratic are the hardest to manage and to forecast. The reason for this is that the quantity and demand for these items varies a lot. These items may also have periods of zero demand. Because of this, it is a challenge for companies to forecast these items. It is hard to manage the random values that appear at random intervals and leaving many periods with zero demand. Due to the lumpy demand, an ongoing problem for most organization is the inaccuracy of forecasts. It is almost impossible to predict exact forecasts. It does not matter how good the forecasts are or how complex the forecast techniques are, the instability of the markets confirm that the forecasts always will be wrong and that errors therefore always will exist. Therefore, we need to accept this but still work with this issue to keep the errors as minimal and small as possible. The purpose with measuring forecast errors is to identify single random errors and systematic errors that show if the forecast systematically is too high or too low. To calculate the forecast errors and measure the forecast accuracy also helps to dimensioning how large the safety stock should be and control that the forecast errors are within acceptable error margins. The research questions answered in this master thesis are: How should one calculate forecast accuracy for stocked items with a lumpy demand? How do companies measure forecast accuracy for stocked items with a lumpy demand, which are the differences between the methods? What kind of information do one need to apply these methods? To collect data and answer the research questions, a literature study have been made to compare how different researchers and authors write about this specific topic. Two different types of case studies have also been made. Firstly, a benchmarking process was made to compare how different companies work with this issue. And secondly, a case study in form of a hypothesis test was been made to test the hypothesis based on the analysis from the literature review and the benchmarking process. The analysis of the hypothesis test finally generated a conclusion that shows that a combination of the measurements WAPE, Weighted Absolute Forecast Error, and CFE, Cumulative Forecast Error, is a solution to calculate forecast accuracy for items with a lumpy demand. The keywords that have been used to search for scientific papers are: lumpy demand, forecast accuracy, forecasting, forecast error.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".