Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A k-nearest neighbors (k-NN) classifier predicts a new example’s class from the labels of nearby training examples. To build one reliably, split your data first, scale features when their units make distances incomparable, and choose the number of neighbors, weighting, and distance metric using validation data—not a universal rule.

How k-nearest neighbors classification works

Scikit-learn describes nearest-neighbor methods as finding a predefined number of training samples closest to a new point and predicting from their labels. In classification, the default approach is a majority vote among those neighbors. The model retains the training examples rather than learning a compact set of parameters, so prediction depends on comparing new observations with stored data. Scikit-learn’s nearest neighbors guide

The parameter n_neighbors, usually called k, sets how many examples vote. A small k makes predictions more local and can be sensitive to noisy or atypical examples. A larger k tends to suppress noise, but produces less distinct decision boundaries. The best value depends on the dataset and should be selected through validation.

Build a k-NN classifier in scikit-learn

  1. Separate features and labels. Put input columns in a feature matrix X and the class to predict in a target vector y. Split into training and test data before fitting; use the test set for final evaluation rather than for choosing settings.
  2. Scale numeric features where appropriate. Distances can be dominated by a feature with a much larger numeric range. If using Euclidean distance, scale numeric columns so that units and ranges do not inadvertently determine which examples count as nearest. Fit any scaler on training data only, then apply it to validation and test data. Scikit-learn’s scaling example explains why scaling matters for a Euclidean k-neighbors model.
  3. Choose an initial configuration. Set n_neighbors explicitly and consider the distance weighting and metric. The example below uses a pipeline so scaling is learned within each training fold during cross-validation.
  4. Compare candidates with validation. Evaluate plausible k values, weights='uniform' and weights='distance', and relevant metrics using cross-validation or a separate validation set. Keep the test set out of this selection process.
  5. Fit and evaluate once selected. Fit the chosen pipeline on the training data, then assess it on held-out test data using measures that reflect the class balance and cost of errors.
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

model = make_pipeline(StandardScaler(), KNeighborsClassifier())
search = GridSearchCV(
    model,
    {
        "kneighborsclassifier__n_neighbors": [3, 5, 7, 11],
        "kneighborsclassifier__weights": ["uniform", "distance"],
        "kneighborsclassifier__metric": ["minkowski"],
        "kneighborsclassifier__p": [1, 2],
    },
    cv=5,
    scoring="accuracy",
)
search.fit(X_train, y_train)

best_model = search.best_estimator_
print(search.best_params_)
print("Test accuracy:", best_model.score(X_test, y_test))

This example assumes the features are numeric and uses accuracy as the cross-validation scoring measure. For imbalanced classes or unequal error costs, choose a more suitable scoring measure and inspect class-specific results rather than relying on accuracy alone. If your data includes categorical or mixed-type features, preprocessing must match those feature types; standardizing arbitrary category codes does not make them meaningful Euclidean distances.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose k, weights, and a distance metric

Number of neighbors

Compare several plausible values of k with the same validation procedure. Smaller values emphasize very local examples; larger values average across a broader neighborhood and can smooth away meaningful local structure. Scikit-learn notes that the optimal k is highly data-dependent, so neither an odd k nor any fixed default is a reliable rule for every dataset. An odd k can reduce a simple two-class voting tie, but it does not eliminate all ties in multiclass problems.

Uniform or distance weighting

With weights='uniform', each selected neighbor has equal influence. With weights='distance', closer neighbors receive more influence, with weights proportional to inverse distance in scikit-learn. Compare both on validation data: distance weighting can let the closest examples matter more, but it is not automatically better when nearby observations are noisy or unrepresentative. Scikit-learn’s weighting documentation

Rank #2
The New Real Book
  • Used Book in Good Condition

Distance metric

The API exposes metric and, for Minkowski distance, the parameter p. Minkowski with p=2 is Euclidean distance; p=1 is Manhattan distance. A metric is meaningful only in the context of feature representation and scaling. Compare reasonable alternatives through validation rather than choosing one solely because it is common. KNeighborsClassifier API reference

Search settings and practical costs

algorithm='auto' lets scikit-learn select among brute-force search, KD-tree, and Ball-tree approaches. The API also exposes leaf_size, which affects tree-based search behavior. These are computational choices, not substitutes for validating predictive quality; the best search strategy depends on the data and metric. Since k-NN retains training examples and compares new points against them, consider prediction speed and memory use as well as test performance when assessing a configuration. KNeighborsClassifier API reference

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When radius neighbors or another model may fit better

If examples are sampled at uneven densities, a fixed k forces every prediction to use the same number of neighbors even when those neighbors span very different distances. RadiusNeighborsClassifier instead uses a fixed distance radius, allowing the number of neighbors to vary locally. It is an alternative to compare when local density varies; its radius also needs validation. Scikit-learn’s nearest neighbors guide

Neighbor methods also become less effective in high-dimensional spaces because of the curse of dimensionality: distances may become less useful for distinguishing which examples are genuinely close. If your feature space is very wide, evaluate whether the distance-based neighborhood remains meaningful rather than assuming that adding features will improve k-NN.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluation checks and edge cases

  • Match metrics to the task. Accuracy can conceal poor performance on a minority class. Review a confusion matrix and, where errors have asymmetric consequences, use measures such as precision and recall that reflect those costs.
  • Watch for distance ties. Scikit-learn warns that when the k-th and (k+1)-th neighbors have identical distances but different labels, the result can depend on the ordering of training data. This matters when repeated or quantized feature values create equal distances. Scikit-learn’s nearest neighbors guide
  • Keep validation honest. Any scaling, feature selection, or other preprocessing learned from data belongs inside the cross-validation process or must be fitted only on the training partition. Repeatedly tuning against the test set turns it into part of model selection.
  • Inspect neighbor evidence. Because predictions are based on stored examples, examine the neighbors behind surprising predictions. Their labels and distances can reveal poor scaling, unrepresentative training data, or a local class overlap.

How to select a configuration

Choose the configuration that performs well on held-out data and fits the way the classifier will be used. Compare predictive quality, sensitivity to scaling and metric choice, behavior as k changes, computational cost and memory, class-imbalance performance, and whether the supporting neighbors make sense. There is no single best k or distance metric independent of the dataset.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.