Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A scikit-learn pipeline lets you preprocess Titanic passenger data and train a classifier as one estimator. Use a ColumnTransformer to handle numeric and categorical columns differently, then place that transformer and a classifier in a Pipeline. Splitting the data before fitting keeps learned preprocessing away from the held-out test set.

Load the Titanic data and identify feature types

The official scikit-learn example fetches the Titanic dataset from OpenML and returns features in X and the survival target in y:

from sklearn.datasets import fetch_openml

X, y = fetch_openml(
    "titanic",
    version=1,
    as_frame=True,
    return_X_y=True,
)

print(X.columns.tolist())
print(X.isna().sum())

Inspect the available columns and missing values before choosing inputs. The official example uses age and fare as numeric features, and embarked, sex, and pclass as categorical features. These are illustrative selections, not a requirement to use only those columns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The target is survived. Do not include it in X: it is the outcome the model is meant to predict.

#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

The current stable scikit-learn mixed-types example shows this OpenML workflow. The versioned scikit-learn 1.6 documentation demonstrates the same general pattern. APIs can vary by release, so check the documentation for the version installed in your environment.

Split the data before fitting preprocessing

Imputers and scalers learn values from data. Split first, then fit the eventual pipeline only on the training portion. The pipeline can transform the test portion using what it learned from training, without letting test-set information influence those steps.

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

stratify=y asks the splitter to preserve the target-class proportions as closely as possible in both partitions. The chosen test fraction and random seed define this example’s split; they do not guarantee a particular model score.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Preprocess numeric and categorical columns separately

Numeric values and category labels need different handling. Numeric imputation can fill missing measurements; categorical imputation can replace missing labels. One-hot encoding represents categories as indicator columns, rather than treating labels such as class numbers as measurements on a numeric scale.

Here is one configuration using the feature groups from the official example:

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

numeric_features = ["age", "fare"]
categorical_features = ["embarked", "sex", "pclass"]

numeric_preprocessing = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="median")),
        ("scaler", StandardScaler()),
    ]
)

categorical_preprocessing = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("onehot", OneHotEncoder(handle_unknown="ignore")),
    ]
)

preprocessor = ColumnTransformer(
    transformers=[
        ("numeric", numeric_preprocessing, numeric_features),
        ("categorical", categorical_preprocessing, categorical_features),
    ]
)

ColumnTransformer routes each selected column group through its own transformer. The numeric branch fills missing values with the training-set median and scales the resulting values. The categorical branch fills missing entries with the most frequent training-set category and one-hot encodes the results. With handle_unknown="ignore", a category encountered during transformation but not seen during fitting does not cause the encoder to fail.

These choices are not universally best. Scaling is useful for many estimators that are sensitive to feature magnitudes, but not every classifier needs it. Imputation strategies also affect the data representation; choose them with the feature meaning and estimator in mind.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Combine preprocessing and the classifier

Put the transformer before a classifier in a Pipeline. The following uses logistic regression as an example; it does not imply a predicted accuracy or establish that this is the best classifier for the task.

from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline

model = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        ("classifier", LogisticRegression(max_iter=1000)),
    ]
)

model.fit(X_train, y_train)
predictions = model.predict(X_test)

Calling fit on the complete pipeline fits the preprocessing steps and classifier in sequence. Calling predict applies those fitted transformations before producing predictions, so training and inference follow the same workflow. Passing the complete estimator to evaluation or search tools also keeps preprocessing inside the fitted model.

Evaluate on held-out passengers

Choose a metric that matches the purpose of the prediction. For example, accuracy counts the fraction of predictions that match the labels, while precision, recall, and F1 summarize different trade-offs for the positive class. No single score is meaningful without the split, selected features, classifier, and metric that produced it.

from sklearn.metrics import accuracy_score

score = accuracy_score(y_test, predictions)
print(score)

This code calculates accuracy for the particular split and configuration above; the value must come from running it in your environment. For more stable model comparisons, use cross-validation on the training data and keep the test set for a final evaluation rather than repeatedly selecting models against it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Tune preprocessing and classifier settings together

Because preprocessing and classification are steps of one estimator, scikit-learn search tools can tune parameters from either step. Pipeline parameters use the step name followed by two underscores and the parameter name. The example below searches the classifier’s regularization parameter while evaluating the pipeline with cross-validation.

from sklearn.model_selection import GridSearchCV

search = GridSearchCV(
    estimator=model,
    param_grid={
        "classifier__C": [0.1, 1.0, 10.0],
    },
    cv=5,
    scoring="accuracy",
)

search.fit(X_train, y_train)
print(search.best_params_)
print(search.best_score_)

best_model = search.best_estimator_
test_predictions = best_model.predict(X_test)

best_score_ is the search’s cross-validation score on training data for the selected configuration, not the held-out test result. To tune preprocessing, add parameters using the relevant step name, such as preprocessor__numeric__imputer__strategy. Keep candidate settings sensible for the data and estimator; a larger search is not automatically a better one.

Optionally return transformed features as a DataFrame

Pipeline use does not require pandas output. If you want transformed output in a DataFrame for inspection or downstream code, scikit-learn also documents an output-format setting:

from sklearn import set_config

set_config(transform_output="pandas")

This is an optional output-format convenience, not a preprocessing step. The related official set_output example illustrates it with Titanic data. Check compatibility with the transformers and scikit-learn version you use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.