What is Cross-Fitting?

Cross-Fitting is a sample-splitting procedure that trains nuisance models on one subset and evaluates each observation's estimating score with predictions from a model that did not train on that observation, then aggregates scores across folds.

Quick Facts

SpecificationOfficial Specification

How It Works

Out-of-fold predictions are the core artifact

Persist one nuisance prediction per observation together with its fold ID, training-data boundary, feature version, model version, and eligibility checks. The score for fold k may use only nuisance models trained outside fold k. Generating predictions from a final all-data model, tuning on held-out scores, or preprocessing globally before splitting reintroduces own-observation leakage.

Orthogonality and Cross-Fitting solve different parts of the problem

Chernozhukov and colleagues combine Neyman-orthogonal scores with Cross-Fitting in Double/Debiased Machine Learning. Orthogonality reduces first-order sensitivity to nuisance error; Cross-Fitting controls empirical dependence between nuisance fitting and score evaluation. Neither one repairs an unidentified estimand, missing overlap, post-treatment features, or arbitrary model failure.

Fold boundaries must preserve the real sampling unit

Random row folds are invalid when rows from one user, session, household, experiment cluster, or future time period can leak across train and held-out sets. Assign folds at the independent unit, use blocked or forward-time designs when required, and repeat the full procedure with declared seeds if split variability matters. Hyperparameter tuning must occur within each training side or in a separate layer.

Key Characteristics

  • Produces out-of-fold nuisance predictions for every observation
  • Separates nuisance fitting from estimating-score evaluation
  • Differs from predictive Cross-Validation and final-model training
  • Pairs naturally with orthogonal and Doubly Robust scores
  • Requires group- or time-aware folds when observations are dependent
  • Does not fix overlap, identification, feature leakage, or bad outcomes

Common Use Cases

  1. Estimating treatment effects with flexible outcome and propensity models
  2. Computing Doubly Robust policy values from logged decisions
  3. Reducing own-observation bias in semiparametric machine learning
  4. Constructing out-of-fold residuals for orthogonal moment equations
  5. Auditing leakage boundaries in causal or OPE pipelines

Example

loading...
Loading code...

Frequently Asked Questions

How is Cross-Fitting different from Cross-Validation?

Cross-Validation primarily estimates predictive performance or selects hyperparameters. Cross-Fitting creates out-of-fold nuisance predictions used inside an estimating equation for a target parameter. They can share folds, but their objective, aggregation, and refitting rules differ.

Why does Cross-Fitting help Doubly Robust estimators?

It prevents each observation from influencing the nuisance model that predicts that observation. With an orthogonal score and suitable nuisance convergence, this reduces own-observation overfitting effects and supports valid target-parameter inference with flexible learners.

How many folds should Cross-Fitting use?

Five or ten folds are common, but no count is universally optimal. More folds enlarge each training set and increase compute; fewer folds simplify dependence and repeated fitting. Choose in advance based on sample size, rare groups, cluster boundaries, and computational budget.

Can preprocessing be fitted before the Cross-Fitting split?

Only transformations fixed independently of the observed sample are safe globally. Learned imputation, normalization, feature selection, representation learning, calibration, and hyperparameter selection must be fitted within each training side or they leak held-out information.

Does Cross-Fitting guarantee unbiased or causal estimates?

No. It addresses reuse of observations during nuisance fitting and score evaluation. Identification still depends on the estimand, treatment or action consistency, overlap, confounding assumptions, correct timing, and a score with the required statistical properties.

Related Terms