What is Bayesian Optimization?

Bayesian Optimization is a sequential, model-based method for optimizing an expensive black-box objective by updating a probabilistic surrogate and using an acquisition function to choose each next evaluation.

Quick Facts

SpecificationOfficial Specification

How It Works

Run a closed, auditable optimization loop

Start with a space-filling or otherwise defensible initial design, evaluate the real objective, fit the surrogate, optimize the acquisition function, evaluate the selected point, and update the history. Frazier's tutorial formalizes this loop for expensive derivative-free global optimization. Every record should preserve parameters, objective values, noise estimates, failures, runtime, feasibility, surrogate version, and selection policy so that a recommendation can be reproduced.

Match the surrogate and policy to the problem

A Gaussian Process with an appropriate kernel is effective for many low-dimensional continuous spaces, but mixed variables, conditional parameters, high dimension, nonstationarity, unknown constraints, and heterogeneous costs can require transformations, trust regions, alternative surrogates, or multi-fidelity policies. Snoek, Larochelle, and Adams show that kernel choice, hyperparameter treatment, variable cost, and parallel evaluation materially affect optimizer behavior. The method name alone does not guarantee sample efficiency.

Evaluate optimization quality, not only the final incumbent

Compare simple regret or best feasible value against cumulative evaluation cost and wall-clock time across repeated seeds and representative tasks. Report failures, infeasible trials, initialization cost, surrogate fitting time, and acquisition optimization time. Use random or quasi-random search as a minimum baseline, preserve an untouched final evaluation when hyperparameters tune a predictive model, and stop using a predeclared budget or decision threshold rather than a favorable transient result.

Key Characteristics

  • Uses a probabilistic surrogate to represent an expensive objective
  • Chooses evaluations sequentially with an acquisition function
  • Balances predicted objective value, uncertainty, cost, and constraints
  • Supports noisy, batch, multi-fidelity, and multi-objective extensions
  • Depends on search-space, kernel, noise, and policy assumptions
  • Must be judged against equal-budget baselines over repeated runs

Common Use Cases

  1. Tuning costly machine-learning training configurations
  2. Selecting scientific or industrial experiments under a fixed budget
  3. Calibrating expensive simulators with noisy measurements
  4. Optimizing engineering designs with black-box constraints
  5. Allocating evaluations across multiple fidelities or compute costs

Example

loading...
Loading code...

Frequently Asked Questions

When should Bayesian Optimization be used?

Use it when evaluations are expensive, the budget is limited, and a probabilistic surrogate can learn useful structure from relatively few observations. It is less attractive when evaluations are cheap, gradients are reliable, or the search space is extremely high-dimensional without exploitable structure.

Is Bayesian Optimization only for hyperparameter tuning?

No. Hyperparameter tuning is one application. The same loop can select laboratory experiments, engineering designs, simulator inputs, treatment policies, or other costly configurations, provided the objective, constraints, noise, and evaluation budget are explicitly defined.

Does Bayesian Optimization always use a Gaussian Process?

No. Gaussian Processes are common because they provide a posterior mean and covariance from limited data. Tree ensembles, neural surrogates, density models, and task-specific probabilistic models can also support BO if their uncertainty is useful for the acquisition decision.

How is Bayesian Optimization different from Active Learning?

Bayesian Optimization usually chooses evaluations to optimize an unknown objective. Active Learning usually chooses examples or queries to obtain labels that improve a predictive task. Both use sequential acquisition, but their actions, utilities, costs, and evaluation protocols differ.

How should a Bayesian Optimization system be evaluated?

Compare best feasible value or simple regret against cumulative cost and wall-clock time across repeated seeds. Include initialization, failed trials, surrogate fitting, acquisition optimization, and parallel resource use, and retain random or quasi-random search as an equal-budget baseline.

Related Terms

Related Articles