Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

K-nearest neighbors (KNN) predicts a new value by finding the k closest training examples: classification returns the most common label, while regression returns an average of their targets. It stores the examples rather than fitting a parametric model, so distance choice, feature scaling, and the value of k directly shape each prediction.

What the implementation needs to do

For training data, use a numeric feature matrix X with shape (n_samples, n_features) and a target array y with one target per row. A prediction query is one row with the same number of features. The model retains X and y; at prediction time it measures the query against the stored rows and uses the nearest k.

The implementation should reject inconsistent inputs: X must be two-dimensional, its row count must match the number of targets, each query must have the expected feature count, and k must be between 1 and the training-set size. The algorithm below uses NumPy for array handling and implements the neighbor search and prediction logic directly.

Measure distance and select neighbors

Squared Euclidean distance is convenient for ranking points: because square root is monotonic, sorting by squared distance gives the same neighbor order as sorting by Euclidean distance, without computing a square root for every training row. For rows a and b, it is sum((a[j] - b[j]) ** 2).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This readable baseline computes every distance and fully sorts the training rows. Its per-query cost is O(n_train log n_train), in addition to calculating the distances. For larger workloads, selecting only the smallest k distances or vectorizing calculations can reduce overhead; indexed search structures are another option discussed below.

Build a KNN classifier

The classifier below supports uniform voting and inverse-distance voting. For uniform voting, it chooses the label with the most neighbors. If vote counts tie, it chooses the smallest label according to NumPy’s sorted unique values, making this implementation’s rule deterministic for compatible label types. Stable sorting preserves training-row order when distances tie at the neighbor cutoff; consequently, equal-distance points may still make the selected neighborhood depend on training-row order.

Rank #2
Airbition Talking Flash Cards for Toddlers Ages 1‑4, 510 Words English Blue
  • 510 Words, 31 Themes: This learning toy for toddlers aged 1-3 years old adds to 31 topics, covering almost all aspects of daily life, including numbers, shapes, colors, animals, transportation, food, etc. Help children recognize and distinguish things
  • Professional Clear Voice: This talking flash cards reader has a clear voice with a standard American accent
  • Montessori Education: This Montessori material simply requires inserting cards, allowing toddlers to use it independently. Utilizing the Montessori education stimulates children's independent learning ability while enhancing their attention and concentration
  • Enhance Language Development: Presenting images and words through the card machine can help children learn new vocabulary and strengthen language comprehension, which can help children in teaching and language development
  • Good for Kids Aged 1-6: It comes in a cute reusable box, suitable as a birthday, Easter, Christmas, Thanksgiving present for kids aged 1-6 years old
import numpy as np

class KNNClassifier:
    def __init__(self, k=5, weights="uniform"):
        if weights not in ("uniform", "distance"):
            raise ValueError("weights must be 'uniform' or 'distance'")
        self.k = k
        self.weights = weights

    def fit(self, X, y):
        X = np.asarray(X, dtype=float)
        y = np.asarray(y)
        if X.ndim != 2:
            raise ValueError("X must be a 2D matrix")
        if len(X) != len(y):
            raise ValueError("X and y must have the same number of rows")
        if not isinstance(self.k, (int, np.integer)) or not 1 <= self.k <= len(X):
            raise ValueError("k must be an integer from 1 to n_samples")
        if len(X) == 0:
            raise ValueError("training data must not be empty")
        self.X = X
        self.y = y
        return self

    def _neighbors(self, x):
        x = np.asarray(x, dtype=float)
        if x.ndim != 1 or x.shape[0] != self.X.shape[1]:
            raise ValueError("query must be a 1D row with n_features values")
        distance2 = np.sum((self.X - x) ** 2, axis=1)
        idx = np.argsort(distance2, kind="stable")[:self.k]
        return idx, np.sqrt(distance2[idx])

    def predict_one(self, x):
        idx, distances = self._neighbors(x)
        labels = self.y[idx]
        if self.weights == "uniform":
            values, counts = np.unique(labels, return_counts=True)
            return values[np.argmax(counts)]

        scores = {}
        for label, distance in zip(labels, distances):
            weight = 1.0 / max(distance, 1e-12)
            scores[label] = scores.get(label, 0.0) + weight
        return max(scores, key=scores.get)

    def predict(self, X):
        X = np.asarray(X, dtype=float)
        if X.ndim == 1:
            return np.asarray([self.predict_one(X)])
        if X.ndim != 2 or X.shape[1] != self.X.shape[1]:
            raise ValueError("X must have shape (n_queries, n_features)")
        return np.asarray([self.predict_one(row) for row in X])

With distance weights, a closer neighbor contributes more. The small denominator floor prevents division by zero; an exact-match training point receives a very large weight. If several exact duplicates have different labels, they can still share influence under this rule, unlike a policy that explicitly returns the exact-match label.

Build a KNN regressor

Regression uses the same neighbor search, but combines numeric targets instead of voting on labels. Uniform weighting returns their arithmetic mean. With inverse-distance weighting, an exact match is handled explicitly: if any selected neighbor has zero distance, return the mean target among the exact matches. This avoids division by zero and defines behavior when duplicate feature rows have different targets.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Aullsaty Talking Flash Cards for Toddlers 1-3, Upgraded 248 Sight Words Montessori Speech Therapy Toy, Autism Sensory Educational Learning Toys, Birthday Gift for Boys Girls (Blue)
  • [ Toddler Montessori Learning Toys ] - The toddler educational talking flash cards is designed as a cute cat card reader which attracts children's interests and includes 248 sight words covering 14 subjects like animals, vehicles, letters, numbers, foods, fruits, vegetables, clothing, nature, colors, persons, jobs, shapes and daily necessities. The speech therapy toy teaches kids to learn with Montessori way by all kinds of animals’ and vehicles’ sounds with a lot of fun and interests.
  • [ Speech Therapy Autism Sensory Toys ] - Your kids can play and interact with the autism sensory toys by themselves with a very interesting upgraded Montessori learning way. It is a also great learning opportunity for autistic children to play with their families. The combination of sound and images enhance their ability to recognize and interact with new things on the cards, which is very suitable for autistic children and speech therapy sessions for children who do not talk.
  • [ Easy to Use ] - Just put the card into the cute cat machine’s mouth ( card reader’s slot ), the American cat will pronounce the words with a standard American accent. The card reader makes a real animal or vehicle’s sound when an animal card or vehicle card is inserted. There are also letters and numbers cards for preschool children and more cards for kindergarten children, your toddler can press the repeat button to repeat the pronunciation and sound, adjust volume to 5 levels.
  • [ Perfect Gifts for Boys and Girls 1-4 Year Old ] - The ABC letters and 123 numbers as well as the cute image, animals’ and vehicles’ sounds and cat card reader is perfect gifts for preschool kids age 1-2 year old, more cute cards is perfect gifts for kindergarten kids age 3-4 year old. The learning sensory toy is a great gift for birthday, Christmas, Halloweens, Easter and back to school day. It can also be used home and in class, parents and teachers can teach little ones learning talking.
  • [ Rechargeable and Durable ] - Aullsaty toddler toy comes with a built-in rechargeable battery and a charger instead of extra batteries, It can be used up to 5 hours and no need to charge frequently. The cards is made of high quality double copper paper which is thicker and durable, not easy to bend. The toy is very portable and size is perfect for toddlers to hold and use. It is also equipped with a cute bag for easy storage of the cards and reader, perfect for children and families to travel.
class KNNRegressor(KNNClassifier):
    def predict_one(self, x):
        idx, distances = self._neighbors(x)
        targets = np.asarray(self.y[idx], dtype=float)
        if self.weights == "uniform":
            return float(np.mean(targets))

        exact = distances == 0
        if np.any(exact):
            return float(np.mean(targets[exact]))
        weights = 1.0 / distances
        return float(np.average(targets, weights=weights))

Example use:

X_train = [[1.0, 10.0], [2.0, 20.0], [3.0, 30.0], [4.0, 40.0]]
y_class = ["low", "low", "high", "high"]
y_value = [12.0, 18.0, 31.0, 39.0]

classifier = KNNClassifier(k=3).fit(X_train, y_class)
regressor = KNNRegressor(k=3).fit(X_train, y_value)

print(classifier.predict_one([2.5, 25.0]))
print(regressor.predict_one([2.5, 25.0]))

Scale features before using distance

Distance combines differences across every feature. If one column records annual income in values of tens of thousands while another records a fraction between 0 and 1, the income differences can dominate Euclidean distance simply because of their units. Scaling puts features on more comparable ranges when that is appropriate for the task.

Fit scaling statistics on the training data only, then reuse them unchanged for validation, test, or future query rows. For standardization, subtract each training feature’s mean and divide by its training standard deviation:

Rank #4
Eaever 520 ABC Sight Words Talking Flash Cards, Christmas Birthday Gift for 2 3 4 5 6 Year Old Boys and Girls, Preschool-Learning-Activities, Toddler Educational Toys for Ages 1-6 Kids, Blue
  • EASY TO USE: Simply insert the cards into the machine, it will read the cards out. Let the loud and clear readings captivate your child.
  • FUN LEARNING: Start an educational journey with a set of 520 sight words, 28 themes, from ABC letters, numbers, animals, and shapes, to colors, nature, seasons, months, etc, your child will explore a wide range of topics. Insert the animal and vehicle cards, the machine will imitate their voices in a hilarious manner.
  • AUTHENTIC SPOKEN: Experience authentic expressions and pronunciation that sets our product apart from the rest. Ideal for enriching kids' language development.
  • RECHARGEABLE & POCKET SIZES: Say goodbye to frequent charging with the built-in rechargeable battery, providing up to 4.5 hours of uninterrupted playtime. Measuring 4*3.75*0.75 inches, the card reader is perfectly sized for little hands.
  • INTERACTIVE TOYS: These Montessori toy sets have limitless possibilities! It empowers parents and teachers to teach language skills, expand vocabulary, and reinforce sight words in a captivating and interactive way.
mean = X_train.mean(axis=0)
scale = X_train.std(axis=0)
scale[scale == 0] = 1.0

X_train_scaled = (X_train - mean) / scale
X_valid_scaled = (X_valid - mean) / scale

Computing the mean or standard deviation from validation or test rows leaks information into evaluation. Scikit-learn’s feature-scaling example illustrates why scaling matters for Euclidean KNN.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose k using validation data

A small k can react strongly to individual noisy or atypical examples. A larger k averages across more neighbors, which can suppress noise but smooth away local boundary detail. There is no universally best value: select it by comparing performance on data not used to fit the neighbor set or preprocessing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Torlam Toddler Flash Cards Baby Cognitive Flashcards for Kids, Learning Alphabet, Numbers, Shapes & Colors, Animals, First Words, Body Parts, Foods, Preschool Kindergarten Activities Educational Toys
  • 【What's Included】Include 60 double-sided toddler flash cards, and 5 colored rings. Designed to teach young children foundational skills, these cards cover the alphabet, counting from 1 to 10, shapes and colors, animals, first words, body parts, foods and fruits.
  • 【Curated for Children】These baby flash cards are beautifully illustrated with vibrant colors, images, and easy-to-read fonts, allowing children to immerse themselves in a world full of fun and learning, sparking their curiosity and imagination with every flashcard.
  • 【Early Skills Development】Young learners will expand their vocabulary, develop their memory, sharpen their focus and improve recognition skills with these first words flashcards. They help children develop essential kindergarten readiness skills.
  • 【Elegant Design】Our flash cards are sized at 4" x 5", making the cards large enough for little hands to hold. All cards have rounded edges. Additionally, the set includes 5 rings for easy classification, keeping the cards neat and organized.
  • 【Ideal toy for Kids】Our flashcards can make a great toy for curious toddlers. This learning toy for kids is perfect for interactive learning activities in preschools, kindergarten classrooms, and homeschooling supplies.
  1. Split data into training and validation sets before fitting any scaler.
  2. For each candidate k, fit the scaler on the training split, transform both splits, fit the KNN model on the transformed training rows, and predict validation rows.
  3. Choose a metric suited to the task: accuracy and a confusion matrix for classification; mean absolute error (MAE) or root mean squared error (RMSE) for regression.
  4. Plot validation score or error against k and choose a value that performs well without relying on a single lucky split. Cross-validation can provide a more robust comparison when the dataset allows it.

For binary classification, an odd-numbered candidate grid can reduce tied uniform votes, but it does not eliminate all ties—for example, tied distances can still affect which rows enter the neighborhood. For multiclass classification and regression, choose a grid appropriate to the sample size and compare validation results rather than assuming odd values are necessary.

Check the scratch version and consider performance

A library comparison is a useful verification step, not proof of correctness. On the same scaled training and validation split, match the metric, k, and weighting, then compare predictions with scikit-learn. Differences may be legitimate if the implementations use different tie handling. For a meaningful comparison, also consider prediction latency and memory use, not just accuracy or error.

Scikit-learn documents brute-force, KD-tree, and Ball-tree neighbor searches. Trees may help for low-to-moderate-dimensional data, while high-dimensional data can make useful neighborhood distinctions harder to obtain. A brute-force implementation remains a clear baseline because every distance calculation is inspectable.

The KNeighborsClassifier API exposes choices including neighbor count, weighting, search algorithm, leaf size, Minkowski parameter p, and metric. In Minkowski distance, p=2 gives Euclidean distance and p=1 gives Manhattan distance; see the distance metric reference. The code here deliberately uses squared Euclidean distance for neighbor ranking. Manhattan distance or another metric can change which training examples count as neighbors, so comparisons should hold the metric constant.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.