agentsclimarketplace

Cross validation

Skill Amey-Thakur/AI-SKILLS/skills/machine-learning/cross-validation

Pick CV schemes that respect time and grouping, nest them for tuning, and read the variance, not just the mean. Use when data is too small for a single split or when validating tuning claims.From its SKILL.md

Install
npx -y skills add Amey-Thakur/AI-SKILLS --skill cross-validation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

3.0 KB, 639 tokens by cl100k_base, as published. Nobody here has run it

Cross-validation

CV answers "how much does my estimate depend on which rows landed where" by averaging over K splits. It only works when every fold respects the same boundaries deployment will face; a leaky fold structure launders leakage into confidence.

Method

  1. Choose the scheme from the data's dependence structure. IID rows: stratified K-fold (K=5-10; stratify on the label, especially when imbalanced; see imbalanced-data). Repeated measures per entity: GroupKFold on the entity so no user/patient/device spans folds (see train-test-discipline). Temporal data: walk-forward (expanding or sliding window, train always before test); shuffled K-fold on time series is the canonical way to publish a model that cannot predict anything.
  2. Run the entire pipeline inside each fold. Scaling, imputation, encoding, feature selection, resampling: all fit on the fold's training portion only (pipeline objects make this automatic; see feature-engineering). Preprocessing fit on the full dataset before CV is the most common silent inflation in applied work.
  3. Nest when CV both tunes and evaluates. Inner loop selects hyperparameters (see hyperparameter-tuning), outer loop estimates generalization of the whole selection procedure; the outer score is the honest one. Tuning and reporting on the same folds overstates performance by exactly the amount the search exploited them.
  4. Report the spread with the mean. Per-fold scores, mean, and standard deviation; overlapping spreads between model A and B is a tie (see model-evaluation on uncertainty). One anomalous fold is information: inspect it (a segment? a time period?), do not average it away (see ml-error-analysis).
  5. Keep a final untouched test set anyway. CV replaces the validation role at small scale, not the test role: model and protocol decisions accumulate against the CV estimate, so the last word comes from data no fold ever saw (see train-test-discipline).
  6. Match the final refit to the protocol. Standard practice: select config by CV, refit on all training data with that config, confirm on test. For iterative learners early-stopped per fold, set the final iteration count from the folds' median rather than re-early-stopping on test.

Boundaries

  • At large data scale, a single well-constructed temporal/group split estimates as well as CV at a tenth the compute; CV's value concentrates where data is scarce.
  • CV variance understates true uncertainty when folds share an entity, a time regime, or collection artifacts; the scheme can only respect boundaries you know to encode.
  • Leave-one-out is high-variance and expensive for most learners; prefer K=5-10 unless a statistical reason says otherwise.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most quality gates skills give in 639 tokens

Counted across 1,524 of the 2,830 authors here whose files we hold, read 2026-09-06

  • Read full output and check exit codein 45 of 1524, across 40 files
  • Verify output confirms the claimin 44 of 1524, across 39 files
  • Identify the command that proves the claimin 43 of 1524, across 39 files
  • Execute the full verification commandin 36 of 1524, across 30 files
  • Produce a verification reportin 34 of 1524, across 18 files
  • Review git diff changesin 30 of 1524, across 16 files
  • Fix build failures immediatelyin 29 of 1524, across 9 files
  • Group findings by severityin 28 of 1524
  • State claim only with evidencein 27 of 1524, across 22 files
  • Verify regression tests with red-green cyclein 26 of 1524, across 22 files
  • Run the full test suitein 26 of 1524, across 25 files
  • Run test suite with coveragein 25 of 1524, across 10 files

Said here and by no other author read

  • Choose CV scheme based on data dependence structure
  • Run entire pipeline inside each fold
  • Fit preprocessing on training portion only
  • Use nested CV for tuning and evaluation
  • Report per-fold scores mean and standard deviation
  • Inspect anomalous folds

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.