What is Principal Component Analysis?
Principal Component Analysis (PCA) is a linear dimensionality-reduction method that projects centered observations onto orthogonal directions ordered by the variance they explain.
Quick Facts
| Created | 1901 by Karl Pearson; statistical formulation by Harold Hotelling in 1933 |
|---|---|
| Specification | Official Specification |
How It Works
Center first, then choose covariance or correlation geometry
PCA is fitted to centered data. Applying Singular Value Decomposition directly to the centered sample matrix avoids explicitly forming the covariance matrix and yields the same principal axes under the usual formulation. Current scikit-learn documentation centers inputs but does not scale each feature.
Standardizing features before PCA is a modeling choice, not a universal rule. Use covariance geometry when units and variance magnitudes are meaningful; use standardized or correlation geometry when incomparable units should receive comparable influence. Persist the training means, scales, component order, solver, and sign convention with the projection.
Treat explained variance as a compression diagnostic
Explained variance ratio reports how much sample variance lies along each retained axis. A cumulative threshold such as 95% describes variance retention on the fitting distribution; it does not guarantee classification accuracy, retrieval recall, robustness, or semantic fidelity.
Select k inside the training and validation protocol, then measure reconstruction error and downstream metrics on untouched data. Fit centering, scaling, and PCA only on each training fold. Fitting the transform before a split leaks the evaluation distribution even though PCA never reads target labels.
Audit whitening, outliers, and unstable components
Whitening rescales retained component scores to unit variance and removes their relative variance scale. It can help a downstream method whose assumptions need that geometry, but it discards amplitude information and can amplify weak directions.
Classical PCA is sensitive to outliers and captures only linear subspaces. Component signs are arbitrary, and nearly equal eigenvalues permit rotations within an almost tied subspace, so individual loadings may change while the represented subspace remains similar. Compare subspaces, reconstruction, and downstream behavior across samples or time rather than requiring element-wise identical loadings.
Key Characteristics
- Produces orthogonal linear components ordered by explained variance
- Equivalently maximizes projected variance and minimizes squared reconstruction error
- Requires centering and is sensitive to the chosen feature scaling
- Provides an explicit transform for compatible new observations
- Can be computed through covariance eigendecomposition or centered-data SVD
- Does not guarantee preservation of labels, neighborhoods, rare signals, or nonlinear structure
Common Use Cases
- Compressing correlated numeric features before a downstream model
- Building a deterministic linear baseline for embedding reduction
- Visualizing dominant linear variation with two or three components
- Denoising when weak directions have been validated as nuisance variation
- Monitoring covariance and subspace drift across dataset versions
Example
Loading code...Frequently Asked Questions
Is PCA feature selection or feature extraction?
PCA is feature extraction. Each retained component is a linear combination of the original features, so the output coordinates do not preserve the identity of selected input columns. Feature selection instead keeps a subset of original variables.
Should data always be standardized before PCA?
No. Centering is fundamental, but unit-variance scaling changes the question from covariance structure to correlation structure. Standardize when feature units are incomparable and equalized influence is intended; retain original scales when their variance magnitudes are meaningful.
Does retaining 95% explained variance preserve model accuracy?
No. Explained variance measures input dispersion, not target information or operational utility. Select the component count within the training protocol and evaluate the downstream metric, rare slices, calibration, latency, and reconstruction on untouched data.
Can PCA transform new data?
Yes. Subtract the training means, apply the same training scales if used, and multiply by the frozen component matrix. Never refit preprocessing on a test batch, and version the entire transform with the downstream model or vector index.
Why can PCA loadings change between runs or samples?
Every component sign is arbitrary, and components with nearly equal eigenvalues can rotate within their shared subspace. Compare explained subspaces and downstream behavior, align signs for reporting, and avoid assigning stable semantic meaning to a single loading vector without uncertainty analysis.