Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To train a classifier in dlib, represent each example as a numeric sample vector, encode binary labels as −1 and +1, configure an svm_c_trainer, call train(), and validate predictions on data the trainer did not see. For more than two classes, wrap that binary trainer with dlib’s one-vs-one or one-vs-all multiclass trainers. This guide covers the complete workflow, including CMake builds, feature scaling, model selection, validation, and the changes in dlib 20.0.

What dlib provides for classification

Dlib is a modular C++ toolkit that includes supervised-learning algorithms, support-vector machines, feature-processing utilities, and multiclass classification tools. Its SVM interfaces are suitable when you want a native C++ model with explicit control over kernels, regularization, and validation rather than a high-level scripting interface.

The current official release identified for this workflow is dlib 20.0, released May 27, 2025. That release added auto_train_multiclass_svm_linear_classifier(), which searches for linear-SVM settings automatically. You can still construct and tune trainers manually when you need a particular kernel or reproducible hyperparameters.

Build the examples with CMake

The official examples use CMake and require a compiler with C++14 support.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. cd examples
  2. mkdir build
  3. cd build
  4. cmake ..
  5. cmake --build . --config Release

On multi-configuration generators, --config Release selects the release configuration. On single-configuration generators, set the build type during configuration if required by your toolchain. The project README also documents installation through vcpkg with vcpkg install dlib; package-manager versions can change, so check the package metadata when reproducing a build.

Train a binary classifier

Choose a sample type and label contract

Each observation must be a fixed-length numeric vector. A small two-dimensional example can use dlib::matrix<double,2,1>. The binary SVM contract requires two classes, conventionally represented by labels -1 and +1. Keep the sample and label containers aligned: item i in the sample vector must have label i.

#include <dlib/svm.h>
#include <vector>

using sample_type = dlib::matrix<double, 2, 1>;
using kernel_type = dlib::radial_basis_kernel<sample_type>;

std::vector<sample_type> samples;
std::vector<double> labels;

sample_type a, b;
a = 1, 2;  samples.push_back(a); labels.push_back(+1);
b = -1, -2; samples.push_back(b); labels.push_back(-1);

Real datasets normally need preprocessing before this point. Put features on comparable scales, handle missing values explicitly, and apply exactly the same transformation to training and test samples. Without scaling, a large-unit feature can dominate a kernel distance and change the decision boundary.

Configure svm_c_trainer

svm_c_trainer is dlib’s binary C-SVM trainer. It uses sequential minimal optimization (SMO) and returns a decision function after training. The regularization parameter C controls the penalty for training errors: larger values usually fit the training set more aggressively, while smaller values permit a wider margin and more violations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
kernel_type kernel(0.1);              // RBF gamma-like parameter

dlib::svm_c_trainer<kernel_type> trainer;
trainer.set_kernel(kernel);
trainer.set_c(10.0);

auto decision_function = trainer.train(samples, labels);

double score = decision_function(samples[0]);
double predicted_label = score >= 0 ? +1 : -1;

The sign of the returned score determines which side of the learned boundary a sample occupies. The magnitude is a margin-like decision value, not a calibrated probability. If you need probabilities, add a calibration procedure using validation data rather than treating the raw score as one.

Validate the binary model

Do not judge a classifier from the training samples alone. Reserve a held-out test set or use cross-validation. Tune C and kernel parameters only inside the training portion of each split; otherwise information from the test fold leaks into model selection.

  • Report the confusion matrix: true positives, true negatives, false positives, and false negatives.
  • Include precision, recall, or class-specific error rates when the costs of mistakes differ.
  • Use a separate final test set when you have enough data for a three-way train, validation, and test split.

Extend the trainer to multiple classes

svm_c_trainer is binary. Dlib’s multiclass wrappers train several binary models and combine their outputs. For N classes, the two principal strategies have different model counts and error behavior.

Strategy Binary models Inference behavior Important trade-offs
One-vs-one N*(N-1)/2 Each pair of classes votes; the class with the strongest aggregate vote wins. Each model sees only two classes, which can simplify boundaries, but model count grows quadratically. Pairwise errors can be diagnosed by inspecting confusing class pairs.
One-vs-all N One classifier scores each class against all other classes; the combined outputs select a class. Fewer models than one-vs-one, but each positive class may be much smaller than the negative pool, making class imbalance and score calibration important.

One-vs-one

Wrap the binary trainer in one_vs_one_trainer when pairwise boundaries are a good fit for your data or when you want errors localized to class pairs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
using binary_trainer = dlib::svm_c_trainer<kernel_type>;

binary_trainer base;
base.set_kernel(kernel_type(0.1));
base.set_c(10.0);

dlib::one_vs_one_trainer<binary_trainer> ovo;
ovo.set_trainer(base);

auto multiclass_function = ovo.train(samples, class_labels);
int predicted = multiclass_function(new_sample);

The exact sample and label types must satisfy the trainer’s template requirements, and all classes must be represented in the training data. With many classes, the quadratic number of pairwise models increases training time and model storage.

One-vs-all

Use one_vs_all_trainer when a linear number of models is preferable or when each class naturally has its own detector.

dlib::one_vs_all_trainer<binary_trainer> ova;
ova.set_trainer(base);

auto multiclass_function = ova.train(samples, class_labels);
int predicted = multiclass_function(new_sample);

Because each classifier contrasts one class with every other class, inspect per-class confusion and recall. A high overall accuracy can conceal a class that is rarely detected.

Automatic linear multiclass SVM tuning in dlib 20.0

Dlib 20.0 introduced auto_train_multiclass_svm_linear_classifier(). It is useful when a linear decision boundary is an acceptable starting point and you want dlib to search linear-SVM settings instead of selecting them manually. Validate the resulting model on held-out data; automatic tuning does not remove the need for a sound split or class-specific error analysis.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Evaluate multiclass predictions correctly

Dlib’s API includes cross_validate_multiclass_trainer for cross-validation of multiclass trainers. Cross-validation is especially useful when the dataset is too small to sacrifice a large test partition.

  1. Choose the number of folds before looking at results.
  2. Fit preprocessing parameters, such as means and standard deviations, inside each training fold.
  3. Run the multiclass trainer on each fold.
  4. Aggregate predictions into a confusion matrix.
  5. Report per-class recall and error counts, not just one accuracy number.

Dlib’s geometric multiclass example demonstrates the API mechanics with three classes. It is a learning example, not a benchmark for production data, latency, memory use, or expected accuracy. Real results depend on feature quality, kernel choice, hyperparameters, class balance, and validation design.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Kernel and parameter choices

Linear kernels

A linear SVM is often the first model to try for many numeric or high-dimensional feature sets. It trains and predicts without the expanded pairwise geometry of a nonlinear kernel and is the basis of dlib 20.0’s automatic multiclass SVM routine.

Radial-basis kernels

An RBF kernel can represent nonlinear boundaries. Its scale parameter and C interact strongly with feature scaling, so select them with cross-validation rather than copying values between datasets.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value

What to compare during tuning

  • Validation performance and per-class errors.
  • Training time and prediction time on your target hardware.
  • Model size and serialization requirements.
  • Sensitivity to feature scaling and outliers.

Common failure modes

Invalid labels or mismatched containers

A binary trainer cannot learn three labels, and sample and label vectors must have identical lengths. Convert labels deliberately and verify the unique-label set before calling train().

Unscaled features

Large differences in feature units can distort RBF distances and margin optimization. Fit scaling on the training data, persist the parameters, and reuse them at inference time.

Training accuracy mistaken for generalization

A model can separate its training samples while failing on new data. Use held-out evaluation or cross_validate_multiclass_trainer and retain the confusion matrix.

Imbalanced one-vs-all classes

One-vs-all creates a potentially small positive class against many negatives. Examine recall for every class and adjust sampling, features, or training strategy when minority classes are missed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Academic citation

The canonical academic reference is Davis E. King, “DLIB-ML: A Machine Learning Toolkit,” Journal of Machine Learning Research, volume 10, pages 1755–1758, 2009. The project describes dlib as “a modern C++ toolkit containing machine learning algorithms and tools for creating complex software in C++ to solve real world problems.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.