A single train-test split can mislead. The model might get lucky, or unlucky, depending on which examples land in each set. Cross-validation reduces that risk by rotating which data serves as the test set.
In k-fold cross-validation, the data is divided into k equal parts. The model trains on k-1 parts and validates on the remaining one. Repeat k times, each fold serving once as validation. Average the results for a more stable estimate.
Common variants
- K-fold, typically five or ten folds
- Stratified k-fold, preserving class proportions
- Leave-one-out, k equals the number of samples
- Time series split, respecting temporal order
Cross-validation costs compute. Training k models takes k times longer than training one. On large datasets, a single holdout set is often enough. On small ones, cross-validation is worth the expense.
Comments
No comments yet. Be the first to share a thought.
Leave a comment