Main Article Content

Abstract

Weather-dependent decision making, such as planning an outdoor barbecue (BBQ), benefits from short-term forecasts that are both accurate and honestly evaluated. This study addresses two overlooked risks in applied weather classification: label leakage from same-day rule-based targets, and validation-set overfitting caused by repeated model-selection decisions. Using the ECA&D Basel daily weather records (2000–2010), the original BBQ-weather label was found to be fully determined by same-day precipitation, so the task was reframed as next-day forecasting through one-day lag features and target shifting. Data were split chronologically into training, validation, and test sets (60:20:20) to preserve temporal independence. Five heterogeneous classifiers (CatBoost, LightGBM, RUSBoost, Nearest Centroid, SGDClassifier) were compared, tuned with a Genetic Algorithm, and combined through three ensemble strategies: Weighted Voting via Dirichlet-distributed random search, Stacking, and Greedy Ensemble Selection. Weighted Voting achieved the best validation F1-score (0.6841), with RUSBoost receiving the largest weight (0.6700). A 7-feature subset, selected via SHAP, native feature importance, and linear coefficients, was statistically indistinguishable from the full 22-feature model (McNemar test,  = 0.8388) and was adopted as the final model for parsimony. On the held-out test set, the final model achieved F1 = 0.6776, ROC-AUC = 0.9076, and PR-AUC = 0.6961, with only a 0.0065 gap from validation performance, confirming strong generalization. These results demonstrate that rigorous chronological splitting and formal statistical testing can materially change both the interpretation and the trustworthiness of ensemble classification results in weather-dependent decision support.

Keywords

Class Imbalance Ensemble Learning Feature Selection McNemar Test Weather Forecasting

Article Details

How to Cite
Amigo, R., Ksatria Brahmacarya, R., & Hilman Rizaldi, M. (2026). BBQ Weather Prediction in Basel Using Ensemble Machine Learning. JURNAL ILMIAH MATEMATIKA DAN TERAPAN, 23(1), 54 - 63. https://doi.org/10.22487/2540766X.2026.v23.i1.18187

References

  1. 1. Caruana, R., Niculescu-Mizil, A., Crew, G., and Ksikes, A. 2004. "Ensemble Selection from Libraries of Models." Proceedings of the 21st International Conference on Machine Learning, 18.
  2. 2. Cawley, G.C. and Talbot, N.L.C. 2010. "On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation." Journal of Machine Learning Research, 11, 2079-2107.
  3. 3. Chen, T. and Guestrin, C. 2016. "XGBoost: A Scalable Tree Boosting System." Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785-794.
  4. 4. Dietterich, T.G. 2000. "Ensemble Methods in Machine Learning." Lecture Notes in Computer Science, 1857, 1-15.
  5. 5. Dorogush, A.V., Ershov, V., and Gulin, A. 2018. "CatBoost: Gradient Boosting with Categorical Features Support." arXiv preprint arXiv:1810.11363.
  6. 6. Ferri, C., Hernandez-Orallo, J., and Modroiu, R. 2009. "An Experimental Comparison of Performance Measures for Classification." Pattern Recognition Letters, 30(1), 27-38.
  7. 7. Kaufman, S., Rosset, S., Perlich, C., and Stitelman, O. 2012. "Leakage in Data Mining: Formulation, Detection, and Avoidance." ACM Transactions on Knowledge Discovery from Data, 6(4), 1-21.
  8. 8. Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y. 2017. "LightGBM: A Highly Efficient Gradient Boosting Decision Tree." Advances in Neural Information Processing Systems, 30, 3146-3154.
  9. 9. Klein Tank, A.M.G. and Coauthors. 2002. "Daily Dataset of 20th-Century Surface Air Temperature and Precipitation Series for the European Climate Assessment." International Journal of Climatology, 22, 1441-1453.
  10. 10. Kotz, S., Balakrishnan, N., and Johnson, N.L. 2000. "Continuous Multivariate Distributions, Volume 1: Models and Applications." 2nd ed. John Wiley & Sons, New York.
  11. 11. Lundberg, S.M. and Lee, S.-I. 2017. "A Unified Approach to Interpreting Model Predictions." Advances in Neural Information Processing Systems, 30, 4765-4774.
  12. 12. McNemar, Q. 1947. "Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages." Psychometrika, 12(2), 153-157.
  13. 13. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. 2011. "Scikit-learn: Machine Learning in Python." Journal of Machine Learning Research, 12, 2825-2830.
  14. 14. Saito, T. and Rehmsmeier, M. 2015. "The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets." PLoS ONE, 10(3), e0118432.
  15. 15. Seiffert, C., Khoshgoftaar, T.M., Van Hulse, J., and Napolitano, A. 2010. "RUSBoost: A Hybrid Approach to Alleviating Class Imbalance." IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans, 40(1), 185-197.