A k-nearest neighbors (k-NN) classifier predicts a new example’s class from the labels of nearby training examples. To build one reliably, split your data first, scale features when their units make distances incomparable, and choose the number of neighbors, weighting, and distance metric using validation data—not a universal rule.
How k-nearest neighbors classification works
Scikit-learn describes nearest-neighbor methods as finding a predefined number of training samples closest to a new point and predicting from their labels. In classification, the default approach is a majority vote among those neighbors. The model retains the training examples rather than learning a compact set of parameters, so prediction depends on comparing new observations with stored data. Scikit-learn’s nearest neighbors guide
The parameter n_neighbors, usually called k, sets how many examples vote. A small k makes predictions more local and can be sensitive to noisy or atypical examples. A larger k tends to suppress noise, but produces less distinct decision boundaries. The best value depends on the dataset and should be selected through validation.
Build a k-NN classifier in scikit-learn
- Separate features and labels. Put input columns in a feature matrix
Xand the class to predict in a target vectory. Split into training and test data before fitting; use the test set for final evaluation rather than for choosing settings. - Scale numeric features where appropriate. Distances can be dominated by a feature with a much larger numeric range. If using Euclidean distance, scale numeric columns so that units and ranges do not inadvertently determine which examples count as nearest. Fit any scaler on training data only, then apply it to validation and test data. Scikit-learn’s scaling example explains why scaling matters for a Euclidean k-neighbors model.
- Choose an initial configuration. Set
n_neighborsexplicitly and consider the distance weighting and metric. The example below uses a pipeline so scaling is learned within each training fold during cross-validation. - Compare candidates with validation. Evaluate plausible k values,
weights='uniform'andweights='distance', and relevant metrics using cross-validation or a separate validation set. Keep the test set out of this selection process. - Fit and evaluate once selected. Fit the chosen pipeline on the training data, then assess it on held-out test data using measures that reflect the class balance and cost of errors.
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
model = make_pipeline(StandardScaler(), KNeighborsClassifier())
search = GridSearchCV(
model,
{
"kneighborsclassifier__n_neighbors": [3, 5, 7, 11],
"kneighborsclassifier__weights": ["uniform", "distance"],
"kneighborsclassifier__metric": ["minkowski"],
"kneighborsclassifier__p": [1, 2],
},
cv=5,
scoring="accuracy",
)
search.fit(X_train, y_train)
best_model = search.best_estimator_
print(search.best_params_)
print("Test accuracy:", best_model.score(X_test, y_test))
This example assumes the features are numeric and uses accuracy as the cross-validation scoring measure. For imbalanced classes or unequal error costs, choose a more suitable scoring measure and inspect class-specific results rather than relying on accuracy alone. If your data includes categorical or mixed-type features, preprocessing must match those feature types; standardizing arbitrary category codes does not make them meaningful Euclidean distances.
#1 Best Overall
Choose k, weights, and a distance metric
Number of neighbors
Compare several plausible values of k with the same validation procedure. Smaller values emphasize very local examples; larger values average across a broader neighborhood and can smooth away meaningful local structure. Scikit-learn notes that the optimal k is highly data-dependent, so neither an odd k nor any fixed default is a reliable rule for every dataset. An odd k can reduce a simple two-class voting tie, but it does not eliminate all ties in multiclass problems.
Uniform or distance weighting
With weights='uniform', each selected neighbor has equal influence. With weights='distance', closer neighbors receive more influence, with weights proportional to inverse distance in scikit-learn. Compare both on validation data: distance weighting can let the closest examples matter more, but it is not automatically better when nearby observations are noisy or unrepresentative. Scikit-learn’s weighting documentation
Rank #2
- Used Book in Good Condition
Distance metric
The API exposes metric and, for Minkowski distance, the parameter p. Minkowski with p=2 is Euclidean distance; p=1 is Manhattan distance. A metric is meaningful only in the context of feature representation and scaling. Compare reasonable alternatives through validation rather than choosing one solely because it is common. KNeighborsClassifier API reference
Search settings and practical costs
algorithm='auto' lets scikit-learn select among brute-force search, KD-tree, and Ball-tree approaches. The API also exposes leaf_size, which affects tree-based search behavior. These are computational choices, not substitutes for validating predictive quality; the best search strategy depends on the data and metric. Since k-NN retains training examples and compares new points against them, consider prediction speed and memory use as well as test performance when assessing a configuration. KNeighborsClassifier API reference
Rank #3
When radius neighbors or another model may fit better
If examples are sampled at uneven densities, a fixed k forces every prediction to use the same number of neighbors even when those neighbors span very different distances. RadiusNeighborsClassifier instead uses a fixed distance radius, allowing the number of neighbors to vary locally. It is an alternative to compare when local density varies; its radius also needs validation. Scikit-learn’s nearest neighbors guide
Neighbor methods also become less effective in high-dimensional spaces because of the curse of dimensionality: distances may become less useful for distinguishing which examples are genuinely close. If your feature space is very wide, evaluate whether the distance-based neighborhood remains meaningful rather than assuming that adding features will improve k-NN.
Rank #4
Evaluation checks and edge cases
- Match metrics to the task. Accuracy can conceal poor performance on a minority class. Review a confusion matrix and, where errors have asymmetric consequences, use measures such as precision and recall that reflect those costs.
- Watch for distance ties. Scikit-learn warns that when the k-th and (k+1)-th neighbors have identical distances but different labels, the result can depend on the ordering of training data. This matters when repeated or quantized feature values create equal distances. Scikit-learn’s nearest neighbors guide
- Keep validation honest. Any scaling, feature selection, or other preprocessing learned from data belongs inside the cross-validation process or must be fitted only on the training partition. Repeatedly tuning against the test set turns it into part of model selection.
- Inspect neighbor evidence. Because predictions are based on stored examples, examine the neighbors behind surprising predictions. Their labels and distances can reveal poor scaling, unrepresentative training data, or a local class overlap.
How to select a configuration
Choose the configuration that performs well on held-out data and fits the way the classifier will be used. Compare predictive quality, sensitivity to scaling and metric choice, behavior as k changes, computational cost and memory, class-imbalance performance, and whether the supporting neighbors make sense. There is no single best k or distance metric independent of the dataset.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

