What is Support Vector Regression?

Support Vector Regression (SVR) is a supervised regression method that seeks a regularized function while ignoring residuals within an epsilon-wide tolerance tube and penalizing deviations beyond it.

Quick Facts

Created1996 by Harris Drucker, Christopher Burges, Linda Kaufman, Alex Smola, and Vladimir Vapnik
SpecificationOfficial Specification

How It Works

Optimize epsilon-insensitive error and function flatness

The epsilon-insensitive loss is max(0, |y-f(x)|-epsilon). A common primal SVR objective combines ||w||^2/2 with C times positive and negative slack, so residuals inside the tube do not affect the loss and large outside residuals grow linearly.

The SVR tutorial by Smola and Schölkopf derives primal, dual, kernel, and loss variants. Objective scaling and parameter names differ across texts and libraries; record whether loss is summed or averaged before comparing C.

Interpret epsilon, C, and support vectors together

Increasing epsilon widens the tolerance tube and often reduces the number of active constraints and support vectors, but may hide errors that matter to the application. Increasing C usually punishes outside-tube errors more strongly and can produce a less regularized fit.

These controls are coupled with kernel bandwidth, target units, sample weighting, and noise. Standardizing a target changes what a numeric epsilon means. Choose an application-level error tolerance first when one exists, then validate neighboring settings rather than treating support-vector sparsity as the goal.

Evaluate against regression baselines and deployment cost

Fit feature scaling, target transforms, kernel parameters, epsilon, and C within training folds. Report MAE or another business-aligned metric in original target units, inspect residual slices and tails, and compare linear regression, tree models, and Kernel Ridge Regression.

Current scikit-learn SVR guidance warns that common kernel implementations scale worse than quadratically with sample count. Prediction evaluates retained support vectors, so measure their count, model size, and latency in addition to held-out error.

Key Characteristics

  • Uses epsilon-insensitive loss with a zero-penalty tolerance tube
  • Balances function regularity against outside-tube penalties through C
  • Represents nonlinear predictions with kernels against support vectors
  • Makes epsilon meaningful only relative to the target scale
  • Penalizes large residuals linearly in the standard epsilon-SVR form
  • Carries support-vector-dependent prediction cost

Common Use Cases

  1. Predicting continuous targets with an explicit acceptable error band
  2. Fitting nonlinear relationships on small and medium datasets
  3. Creating a robust-to-small-noise alternative to squared-loss regression
  4. Using domain-specific sequence, molecular, or graph kernels for regression
  5. Comparing sparse kernel expansions with Kernel Ridge Regression

Example

loading...
Loading code...

Frequently Asked Questions

What does epsilon mean in Support Vector Regression?

Epsilon is the tolerated absolute residual under the target's current units or transformed scale. Predictions inside the tube receive zero epsilon-insensitive loss. It is not a confidence interval and does not guarantee that a chosen fraction of future targets will fall inside.

How do C and epsilon interact in SVR?

Epsilon determines which residuals activate the loss, while C controls how strongly active violations compete with regularization. Wider tubes often retain fewer support vectors; larger C often fits outside-tube deviations more aggressively. Tune both with kernel and scale.

How does SVR differ from Kernel Ridge Regression?

Standard epsilon-SVR uses an epsilon-insensitive linear penalty and often yields a sparse support-vector expansion. KRR uses squared loss and generally retains coefficients for all training references. They also differ in optimization, sensitivity, and parameter conventions.

Should targets be standardized before SVR?

Target scaling can improve numeric conditioning, but it changes the meaning of epsilon and reported errors. Fit the transform on training data only, invert predictions before business evaluation, and document epsilon in both transformed and original units.

Why can kernel SVR be slow at prediction time?

Each query typically evaluates the selected kernel against every retained support vector. A flexible fit or noisy overlap can retain many vectors. Measure count and latency, then compare linear SVR, explicit kernel approximations, or other regressors.

Related Terms

Related Articles