AIWiki
Malaysia
Back to all articles
AI Foundationscross-validationmodel-evaluationk-fold

Cross-validation

4 min readUpdated October 2026
Cross-validation
Type
Model evaluation technique
Field
Statistical machine learning
Introduced
1970s (Stone, Geisser)
Common variants
k-fold, stratified k-fold, leave-one-out, nested CV
Related
Hold-out validation, bootstrapping, hyperparameter tuning
Cross-validation is a model evaluation strategy in which the available data is partitioned into complementary training and validation subsets, the model is fitted and scored repeatedly across those subsets, and the results are aggregated into a single performance estimate. It is used to check how well a model generalises to unseen data and to compare models without relying on a single, possibly lucky, train-test split.

History

The statistical roots of cross-validation trace to the work of Morris Stone and Playfax Geisser in the early 1970s on cross-validatory choice and prediction, which formalised using held-out data to assess how well a fitted model predicts new observations.[1] Leave-one-out cross-validation — refitting the model once per data point — was among the earliest variants and remains common in small-sample statistics.

As machine learning industrialised in the 1990s and 2000s, k-fold cross-validation became the default evaluation protocol in research and practice, codified in standard references such as The Elements of Statistical Learning and implemented directly in libraries like scikit-learn.[2][3] Its importance grew further when researchers demonstrated that selecting model hyperparameters on the same data used to report final performance produces optimistically biased results, motivating nested cross-validation.[4]

Key Concepts and Technology

In k-fold cross-validation, the dataset is split into k roughly equal folds. The model trains on k−1 folds and is evaluated on the remaining fold, repeating k times so every fold serves as the validation set once; the k scores are averaged and their spread reported. k = 5 or k = 10 is the usual default, balancing bias and variance: small k leaves little training data per fit (higher bias), while large k costs more computation and yields highly correlated folds (higher variance).

Important variants address specific data problems:

  • Stratified k-fold preserves the class distribution in each fold, which matters for imbalanced tasks such as fraud detection.
  • Leave-one-out uses k equal to the number of samples — low bias but expensive and high-variance.
  • Nested cross-validation runs an inner loop for hyperparameter tuning inside an outer loop for evaluation, preventing the tuning step from leaking information into the reported score.[4]
  • Time-series cross-validation uses rolling or expanding windows with no shuffling, because randomly mixing future observations into training data creates look-ahead leakage.
A common source of error is preprocessing leakage: scaling, feature selection or embedding fitting must happen inside each training fold, not once on the full dataset, or validation scores are inflated.

Applications or Impact

Cross-validation underpins model selection, benchmark reporting and uncertainty estimation across machine learning. Practitioners use it to compare algorithms, tune hyperparameters and detect overfitting before committing to a final hold-out or test set. Its main cost is compute — k fits instead of one — which is why large deep-learning runs often substitute a single fixed validation split during training and reserve cross-validation for the final model comparison.

>See Also

🇲🇾Malaysian Context

🇲🇾 Relevance to Malaysia: Malaysian banks and fintechs building fraud-detection or credit-scoring models are expected to demonstrate robust out-of-sample performance under Bank Negara Malaysia's risk management and AI guidance, and cross-validation is the standard evidence for that claim. The Malaysia AI Governance Framework likewise emphasises validation and monitoring of AI systems before deployment.

Local datasets are often small — palm-oil yield prediction, SME credit scoring and public-sector analytics rarely reach the sample sizes of global benchmarks — so k-fold cross-validation extracts the most reliable estimate possible from limited labelled data. Universities and MDEC-supported research programmes teaching applied machine learning routinely make k-fold the first evaluation method in their curricula, since it requires no extra data and catches overfitting early.

References

  1. ↑Wikipedia. Cross-validation (statistics) — history of the method. https://en.wikipedia.org/wiki/Cross-validation_(statistics)
  2. ↑Hastie, T., Tibshirani, R. and Friedman, J. The Elements of Statistical Learning — Model Assessment and Selection. https://web.stanford.edu/~hastie/ElemStatLearn/
  3. ↑scikit-learn. Cross-validation: evaluating estimator performance. https://scikit-learn.org/stable/modules/cross_validation.html
  4. ↑Cawley, G. and Talbot, N. (2010). On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation. Journal of Machine Learning Research, 11. https://jmlr.org/papers/v11/cawley10a.html