Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A deep forest is a layered machine-learning architecture built from decision-tree ensembles rather than neural-network layers. Zhou and Feng’s gcForest method was designed to retain layer-by-layer processing and feature transformation without training those layers through backpropagation. Its paper reports robust results across settings, but that is not proof that it generally beats convolutional neural networks (CNNs) or recurrent neural networks (RNNs). Which model is a better fit depends on the data, task, evaluation metric and experimental setup.

What is a deep forest?

A deep forest stacks processing stages made from tree ensembles. In Zhou and Feng’s gcForest, the ensembles act as the model’s modules: earlier stages process the input, and later stages work with transformed representations. The aim is to capture some of the layered processing associated with deep learning without building the model from differentiable neural-network layers.

The authors identify three intended characteristics: layer-by-layer processing, feature transformation within the model and sufficient model complexity. The depth and complexity can be determined in a data-dependent way, rather than fixed entirely by a manually specified neural-network architecture. That does not mean there are no modeling choices to make; it means the proposed method aims to reduce the amount of manual hyperparameter design.

Zhi-Hua Zhou and Ji Feng submitted Deep Forest to arXiv on 28 February 2017. The arXiv record lists a 6 July 2020 revision and a journal reference to National Science Review, 2019, volume 6, issue 1, pages 74–86.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How gcForest differs from CNNs and RNNs

CNNs and RNNs are neural-network families trained using differentiable operations and backpropagation. CNNs are commonly associated with spatial structure, such as images; RNNs are commonly associated with sequences, where order matters. These are broad associations, not strict limits on what the model families can be applied to.

gcForest takes a different route: it uses tree ensembles and does not rely on backpropagation to train a stack of neural layers. The paper presents this as a way to build a deep model using non-differentiable modules. It is therefore better understood as an alternative kind of layered model, not as a forest-shaped version of a CNN or RNN.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Comparison point Deep forest / gcForest CNN RNN
Core building blocks Layered decision-tree ensembles, in gcForest Neural-network layers commonly used to process spatial structure Neural-network layers commonly used to process ordered sequences
Training approach Tree-ensemble learning; the layered model does not use backpropagation Backpropagation through differentiable neural-network operations Backpropagation through differentiable neural-network operations
Commonly associated data Not confined to one modality; suitability depends on how the data are represented Images and other data with useful spatial structure Sequences and other data where order is important
Model complexity The paper says gcForest complexity can be determined in a data-dependent way Depends on the chosen architecture and training setup Depends on the chosen architecture and training setup
Evidence for a general performance winner Not established: the paper makes qualitative claims of robust performance, not a universal win Not established as a universal winner Not established as a universal winner

The table describes broad design differences, not a controlled head-to-head result. It does not imply that every CNN or RNN is appropriate for its associated data type, or that gcForest will perform well on every kind of input.

Does gcForest outperform CNNs or RNNs?

There is no basis here for a blanket claim that it does. Zhou and Feng report that gcForest was robust to hyperparameter settings and that, in most cases, it achieved excellent performance across data from different domains using the same default setting. Those are the authors’ qualitative claims about their method; they are not a guarantee for a new dataset, nor do they establish that gcForest beats CNNs or RNNs on every task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The exact benchmark details needed to substantiate a numerical comparison are not available here. Without the datasets, splits, preprocessing, competing model configurations, compute budgets and evaluation metrics, a claimed margin would be misleading. A model can appear to win because its input representation or tuning effort is better matched to the task, not because its model family is universally superior.

When should you try a deep forest?

Consider gcForest when you want to test a layered tree-ensemble approach, especially if reducing manual neural-architecture design or avoiding backpropagation is valuable for your project. Its reported robustness to hyperparameter settings makes it worth evaluating when extensive tuning is impractical, but the paper does not establish a particular dataset size or modality where it will reliably win.

A CNN is a natural candidate when spatial relationships in the input are central to the task, such as patterns across image regions. An RNN is a natural candidate when the ordering of sequence elements matters. These are starting points, not rules: the right comparison depends on the actual representation and objective.

  • Try gcForest when a tree-ensemble approach is a plausible fit and you can evaluate it against the models already used for your task.
  • Try a CNN when preserving and learning spatial structure is important.
  • Try an RNN when the task depends on ordered, sequential inputs.
  • Compare more than one family when the decision has material consequences and the evidence on your dataset is uncertain.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to make a fair comparison

Choose the models and evaluation plan before interpreting results. A useful comparison controls the factors that can otherwise make a benchmark misleading:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Define the task and metric. Choose the metric that reflects the real cost of errors for your use case. Do not treat scores measured with different metrics as interchangeable.
  2. Use the same data partitions. Keep training, validation and test data consistent across models, and avoid letting information from the test set influence model selection.
  3. Make input handling explicit. Record preprocessing and feature representation for each model. A fair comparison does not require identical transformations when models need different inputs, but it does require transparent, defensible choices.
  4. Give each model a comparable tuning opportunity. Record the settings and effort used. gcForest’s reported robustness to settings does not justify tuning competing models carefully while leaving one model at an unsuitable default.
  5. Compare resources as well as scores. Account for the compute and memory used in training and evaluation. The paper’s claim of fewer hyperparameters than deep neural networks is not, by itself, evidence that gcForest is always cheaper or faster.
  6. Check stability and practical fit. Look at variation across appropriate runs or data splits, then consider whether the model’s performance and operational demands suit the intended use.

Interpretability and small-dataset performance should also be assessed on the actual task rather than assumed from a model’s name. A tree-based design does not automatically make a complete layered ensemble easy to explain, and no general small-data advantage over CNNs or RNNs is established by the evidence described above.

Do you need backpropagation for deep learning?

No. Backpropagation is central to training many neural-network models, but Zhou and Feng’s proposal illustrates a different meaning of deep: multiple processing layers, feature transformation and sufficient model complexity can be built from non-differentiable decision-tree ensembles. The distinction is about how the model is constructed and trained; it does not make tree ensembles neural networks.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.