Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For most predictive tasks, implement multinomial logistic regression with scikit-learn’s LogisticRegression inside a Pipeline. Use a multinomial-capable solver such as lbfgs, keep scaling and encoding in the pipeline, and evaluate both predicted labels and probabilities. If your priority is maximum-likelihood estimates and inferential output, use statsmodels’ MNLogit.

What multinomial logistic regression does

Multinomial logistic regression models a categorical target with three or more classes. It calculates a score for each class, then applies the softmax function to turn those scores into probabilities that sum to one. Scikit-learn describes the multiclass approach in its LogisticRegression documentation.

The model has a coefficient vector for each class. Scikit-learn uses this symmetric representation deliberately; for an unpenalized model, that parameterization can make the solution non-unique. In practice, scikit-learn applies regularization by default, which also helps control coefficient size.

Fit a multinomial model with scikit-learn

This example splits the data while preserving class proportions, fits scaling only on the training portion, and evaluates the model on held-out data. It assumes X contains numeric features and y contains the categorical target.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix, log_loss
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=42
)

model = Pipeline([
    ("scale", StandardScaler()),
    ("clf", LogisticRegression(
        solver="lbfgs",
        penalty="l2",
        max_iter=1000,
        random_state=42,
    )),
])

model.fit(X_train, y_train)
pred = model.predict(X_test)
proba = model.predict_proba(X_test)

print(classification_report(y_test, pred))
print(confusion_matrix(y_test, pred))
print(log_loss(y_test, proba))

Keeping preprocessing inside the pipeline matters: the scaler is fitted on training data rather than on the full dataset. This avoids letting information from the test set influence the transformation. Scikit-learn’s Pipeline guide demonstrates this train-then-score pattern.

Use a ColumnTransformer for mixed columns

If your data includes both numeric and categorical columns, use a ColumnTransformer to scale numeric columns and one-hot encode categorical columns, then place that transformer and the classifier in the same pipeline. This keeps each preprocessing step tied to the training data and applies the fitted transformations consistently to test or future inputs.

Choose a solver and penalty

For three or more classes, the solver must support scikit-learn’s multinomial loss. The current LogisticRegression reference identifies lbfgs as a good general-purpose default and lists lbfgs, newton-cg, newton-cholesky, sag, and saga as multinomial-capable. liblinear handles binary classification, not the true multinomial loss; it can be used for multiclass only by wrapping it in a one-versus-rest strategy.

Need Practical choice Important consideration
Stable baseline for a typical predictive task lbfgs with L2 regularization Scikit-learn describes lbfgs as a good default for a wide range of problems.
L1 sparsity or Elastic-Net with multinomial loss saga Scale features; its fast-convergence guarantee assumes similarly scaled features.
Many more samples than features multiplied by classes Consider newton-cholesky Its Hessian has quadratic memory dependence on the product of feature count and class count.
One-versus-rest multiclass setup using liblinear Wrap liblinear in a one-versus-rest strategy This is not the same as optimizing the multinomial loss directly.

Scikit-learn regularizes by default. A very large C weakens regularization and approximates an unpenalized fit, but removing regularization does not remove the potential non-uniqueness of the symmetric multinomial parameterization.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Evaluate labels and probability quality

Check class-by-class label performance

The confusion matrix shows which classes the model tends to confuse. The classification report gives precision, recall, and F1 for each class, so a strong overall score cannot hide weak performance on an important minority class.

Assess predicted probabilities with log loss

predict_proba returns a probability for every class, not just the winner. Multiclass log_loss measures the negative log-likelihood of those predicted probabilities; lower values indicate better probabilistic fit on the same evaluation set. See scikit-learn’s log_loss reference.

When a decision depends on a risk threshold, inspect calibration on a validation set as well as the class labels. A model can choose the correct class while assigning probabilities that are poorly calibrated. There is no universal accuracy figure to expect: results depend on the data, class balance, feature representation, regularization, and how the data is split.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use statsmodels when inference is the priority

Scikit-learn is a natural fit for predictive workflows, regularization, pipelines, and production-oriented evaluation. Choose statsmodels’ MNLogit when maximum-likelihood estimation, coefficient tables, and likelihood-based diagnostics are central. The MNLogit.fit documentation describes fitting by maximum likelihood; statsmodels also exposes methods such as fit_regularized, loglike, and score.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import statsmodels.api as sm

X_sm = sm.add_constant(X)
result = sm.MNLogit(y, X_sm).fit()
probabilities = result.predict(X_sm)
print(result.summary())

Before interpreting the output, document how the target is coded, which category is the reference, whether the feature matrix includes an intercept, and how features are represented. Coefficients describe changes relative to the base outcome; they are not ordinary linear-regression slopes.

Statsmodels’ MNLogit.predict documentation describes mean, linear, variance, and probability outputs. For probability output, column 0 is the base case, with remaining columns corresponding to shifted parameter rows. Confirm the returned category ordering against your target coding before assigning meaning to individual columns.

Quick decision guide

  • Use scikit-learn with lbfgs and L2 for a straightforward predictive baseline.
  • Use saga when you need L1 or Elastic-Net in a multinomial model, and scale features.
  • Keep numeric scaling and categorical encoding inside a pipeline to prevent test-set leakage.
  • Use confusion matrices and class-wise precision, recall, and F1 for class decisions, then add log loss for probability quality.
  • Use statsmodels MNLogit when likelihood-based estimation and inferential output matter more than a streamlined prediction workflow.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.