What is Support Vector Machine?

Support Vector Machine (SVM) is a supervised learning method that selects a decision boundary with a large margin between classes, using the training points that constrain that margin as support vectors.

Quick Facts

Created1995 by Corinna Cortes and Vladimir Vapnik
SpecificationOfficial Specification

How It Works

Optimize a margin objective with explicit conventions

For binary labels y_i in {-1,+1}, the functional margin is y_i(w^T x_i+b). The hard-margin problem requires every margin to be at least one and minimizes ||w||^2; the soft-margin form adds slack penalties when data overlap or contain noise.

The original support-vector network paper develops the maximum-margin and kernel construction. Libraries differ in loss normalization, intercept regularization, class weighting, and whether C multiplies a sum or mean, so numeric hyperparameters are not portable without the full objective.

Let support vectors determine the fitted boundary

Karush-Kuhn-Tucker conditions imply that only examples with nonzero dual coefficients contribute directly to the kernel expansion. In a linear model, the weight vector can be reconstructed from those coefficients; in a nonlinear model, every prediction evaluates kernels against retained support vectors.

A large support-vector fraction can signal overlap, weak features, a very flexible kernel, or aggressive regularization choices. It also increases inference latency and model size. Support vectors are influential training examples under the fitted objective, not automatically mislabeled records or causal explanations.

Validate preprocessing, decisions, and operating cost

Feature scale strongly affects Euclidean distances, dot products, and RBF bandwidth. Fit scaling, feature selection, kernel parameters, C, and class weights inside each training fold. Use nested or untouched evaluation data when model selection itself must be assessed.

Current scikit-learn SVM guidance notes that SVM training scales at least quadratically with sample count in common implementations. Decision scores are signed distances or kernel-expansion values, not probabilities; probability estimates require a separate fitted calibration step and independent validation.

Key Characteristics

  • Maximizes a geometric margin while penalizing permitted violations
  • Represents the fitted boundary through support vectors
  • Supports linear and valid positive-semidefinite kernel functions
  • Uses C to trade margin regularization against training penalties
  • Produces decision scores rather than calibrated probabilities by default
  • Can incur quadratic-or-worse training cost on large sample sets

Common Use Cases

  1. Classifying high-dimensional text or sparse feature vectors
  2. Building nonlinear classifiers for small and medium datasets
  3. Creating a maximum-margin baseline before larger neural models
  4. Handling unequal class costs with validated sample or class weights
  5. Comparing exact kernel boundaries with linear or approximate features

Example

loading...
Loading code...

Frequently Asked Questions

What does a Support Vector Machine maximize?

A binary SVM maximizes geometric separation subject to its soft- or hard-margin objective. In the common soft-margin form, it trades a smaller weight norm against penalties for points inside the margin or on the wrong side of the boundary.

What does C control in a Support Vector Machine?

C controls the relative cost of training violations under a specific objective convention. Larger C usually penalizes violations more strongly and may create a tighter, less regularized fit; smaller C usually accepts more violations for a wider or smoother margin.

Why must features be scaled before an SVM?

Large-scale features dominate dot products and distance-based kernels, changing both the margin and kernel bandwidth. Fit the scaler only on each training fold, apply it unchanged to validation and production data, and tune it as part of the model pipeline.

Does an SVM decision score represent probability?

No. Its sign selects a side of the boundary and its magnitude follows the fitted margin or kernel expansion. If probability estimates are required, fit a calibration model without reusing evaluation labels and verify reliability on representative held-out slices.

When should a linear model replace a kernel SVM?

Prefer a linear baseline when samples are numerous, features are high-dimensional, latency is strict, or nonlinear gains are small. Compare held-out quality, support-vector count, memory, training time, and prediction latency rather than choosing from training accuracy.

Related Terms

Related Articles