Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To classify time-series data with TensorFlow, represent each example as a tensor shaped (batch, time steps, features), create leakage-safe training, validation and test partitions, normalize using training data only, and train a model with a class-output head. A 1D convolutional neural network (CNN) is a strong baseline; a Transformer is a valid alternative when long-range attention is worth the added complexity. Choose between them with the same held-out data and metrics rather than assuming one architecture always wins.

What time-series classification means

Classification assigns a discrete label to an entire time-series example: for instance, identifying an engine condition from a sensor trace. This differs from forecasting, which estimates future numerical values or sequences. TensorFlow’s prominent time-series tutorial focuses on forecasting, but its windowing, input-pipeline and time-aware evaluation practices are useful when adapted carefully; it is not itself a classification recipe.

The Keras FordA example describes training a classifier from scratch on motor-sensor engine-noise measurements. FordA contains 3,601 training instances and 1,320 test instances in that example. Those counts describe this dataset, not a required dataset size or an expected model result.

Represent the data with the right shape

Keras recurrent, convolutional and attention layers commonly expect three dimensions:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
  • Batch: the number of examples processed together.
  • Time steps: samples in each sequence.
  • Features: channels measured at each time step.

A univariate sequence therefore has one feature, while a multivariate sequence might contain temperature, pressure and vibration channels. If every observation has a fixed length of 500, a univariate batch has shape (batch, 500, 1). Variable-length series require a deliberate strategy such as padding and masking, resampling, or extracting fixed windows.

FordA-style loading

The Keras FordA example reads separate FordA_TRAIN and FordA_TEST tab-separated files. The first column is the label, and the remaining columns form each series. It reshapes the series to add a channel dimension and converts the example’s labels from -1/1 to 0/1. FordA’s sequences are length 500 and already z-normalized; do not treat those dataset-specific details as universal requirements.

import numpy as np

# x_train and x_test: (examples, time_steps)
x_train = x_train[..., np.newaxis]
x_test = x_test[..., np.newaxis]

y_train = (y_train + 1) // 2
y_test = (y_test + 1) // 2

For your own data, document whether it is univariate or multivariate, how missing readings are handled, whether timestamps are regular, and whether sequence lengths are fixed.

How should I split and normalize time-series data?

Use separate training, validation and test roles. Fit preprocessing parameters on training data only, then apply the same transformation to validation, test and production inputs. Computing a mean, standard deviation or other scaling statistic from validation or test values leaks information into model selection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Machine Learning Using TensorFlow Cookbook: Create powerful machine learning algorithms with TensorFlow
  • Machine Learning Using TensorFlow Cookbook: Create powerful machine learning algorithms with TensorFlow
  • ABIS BOOK
  • Packt Publishing

Choose a split that matches deployment

  • Chronological deployment: keep earlier periods for training and later periods for validation and testing.
  • Independent entities: if examples come from machines, patients or users, keep related observations in one partition so the same entity does not appear on both sides.
  • Benchmark datasets: honor supplied partitions such as FordA’s predefined train and test files; create validation data from the training portion.

A random split can be reasonable only when examples are genuinely independent and deployment does not involve a future-time shift. Whatever rule you choose, use exactly the same partitions when comparing models.

Apply training-only scaling

mean = x_train.mean(axis=(0, 1), keepdims=True)
std = x_train.std(axis=(0, 1), keepdims=True)
std = np.maximum(std, 1e-8)

x_train = (x_train - mean) / std
x_val = (x_val - mean) / std
x_test = (x_test - mean) / std

Per-series normalization can remove absolute-level information, while global training-set normalization preserves relative differences between examples. Select the policy based on what the label depends on and keep it fixed at inference time.

A practical 1D CNN baseline

A fully convolutional network is a sensible first model when local temporal patterns are informative. The documented Keras FordA baseline stacks three Conv1D blocks with 64 filters and kernel size 3, followed by batch normalization and ReLU, then global average pooling and a softmax class head. These settings are example values, not guaranteed optima.

import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers

num_classes = 2
inputs = keras.Input(shape=(x_train.shape[1], x_train.shape[2]))
x = inputs
for _ in range(3):
    x = layers.Conv1D(64, 3, padding="same")(x)
    x = layers.BatchNormalization()(x)
    x = layers.ReLU()(x)
x = layers.GlobalAveragePooling1D()(x)
outputs = layers.Dense(num_classes, activation="softmax")(x)
model = keras.Model(inputs, outputs)
model.compile(
    optimizer="adam",
    loss="sparse_categorical_crossentropy",
    metrics=["accuracy"],
)

history = model.fit(
    x_train, y_train,
    validation_data=(x_val, y_val),
    epochs=50,
    batch_size=32,
)

Use binary_crossentropy with a single sigmoid output if you deliberately configure a two-class binary head; a two-unit softmax with sparse integer labels is often simpler when the class count may change.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use a CNN or Transformer?

A Transformer classifier can combine attention blocks with feed-forward layers, Conv1D projections, global average pooling and a class-output head. Attention can be useful when relationships across distant time steps matter, but the existence of a Transformer example does not establish that it will outperform a CNN on your data. The Keras Transformer example also references older TensorFlow compatibility in its surrounding material, so verify the current notebook and installed TensorFlow/Keras versions before relying on code unchanged.

Compare them fairly

Decision factor Questions to answer
Held-out quality Which model performs better on the same untouched test partition and metric?
Sequence and data size Is the dataset large and the sequence long enough to justify attention complexity?
Cost What are measured training and inference times on the environment where the model will run?
Robustness Do results hold across classes, entities and later time periods?
Operations Can the team monitor, explain and update the architecture it selects?

Run a reproducible comparison rather than naming a universal winner. The supplied examples report no benchmark accuracy, training time or hardware requirement for your dataset.

Evaluate a classifier without leakage

  1. Train candidate models using only the training partition.
  2. Use validation data for architecture, preprocessing and hyperparameter decisions.
  3. After decisions are frozen, evaluate once on the reserved test partition.
  4. Report the class distribution and metrics that reflect the cost of errors.

Accuracy can hide poor minority-class performance. For imbalanced labels, add class-sensitive measures such as per-class precision and recall, a confusion matrix and a macro-averaged score. TensorFlow’s imbalanced-data tutorial explains why imbalance requires explicit treatment, although it is not a time-series classification example.

For a temporal application, inspect performance by time period or entity as well as in aggregate. A model that scores well on randomly mixed windows may fail when deployed on a later period or a previously unseen machine.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failure modes and fixes

Shape errors

If a layer expects three dimensions, add the feature axis for univariate data with x[..., np.newaxis]. Confirm that every feature channel has the same time indexing.

Validation results look unrealistically strong

Check for overlapping windows, duplicate entities, future-derived features, and scalers fitted on all data. Rebuild partitions according to the deployment timeline or entity boundary.

One class is rarely detected

Inspect class counts and the confusion matrix. Consider class weights or resampling within the training partition, and judge changes with minority-aware metrics rather than accuracy alone.

Sequences have unequal lengths

Decide explicitly whether to pad and mask, resample, or derive fixed-length windows. Record how padding and missingness are represented so training and inference use the same convention.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Code no longer runs unchanged

Keras examples have different update dates, and APIs evolve. Check the current notebook, TensorFlow version and Keras serialization behavior before promising exact compatibility.

Save and reuse the trained model

TensorFlow’s save/load guidance recommends the .keras format for Keras objects. Save the model together with the preprocessing contract: feature order, sequence length, missing-value handling, and the training-fitted normalization statistics.

model.save("timeseries_classifier.keras")
restored = keras.models.load_model("timeseries_classifier.keras")
probabilities = restored.predict(x_new)
labels = probabilities.argmax(axis=1)

If you use custom layers or metrics, verify the current custom-object requirements when loading. A saved classifier is not reproducible unless its preprocessing and label mapping are saved as well.

A focused workflow from raw signals to deployment

  1. Define the label, prediction unit and deployment timing.
  2. Audit timestamps, missing values, entity IDs, sequence lengths and class balance.
  3. Create leakage-safe train, validation and test partitions.
  4. Fit training-only normalization and apply it unchanged elsewhere.
  5. Reshape inputs to (batch, time steps, features).
  6. Train the fully convolutional baseline.
  7. Evaluate with the metrics and slices that match operational errors.
  8. Only then compare a Transformer or other architecture under the same protocol.
  9. Freeze preprocessing, label mapping and model serialization for inference.

TensorFlow tutorials are available as runnable Google Colab notebooks, which can be useful for experimentation without assuming a particular computer purchase. Resource limits still depend on sequence length, model size and dataset volume.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Next steps

Start with the Keras FordA classification example to understand loading, reshaping and the CNN structure, then adapt the split and normalization policy to your deployment. Use the forecasting tutorial only for transferable windowing and temporal-evaluation ideas. For broader Keras and TensorFlow coverage, TensorFlow lists Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow as optional further reading; verify the current edition and availability before purchasing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.