Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Yes. In Ertuğrul Mutlu’s 2026 preprint, small convolutional neural networks trained in opposite task orders later reached similar predictive performance after shared training, while selected internal layers still differed under the study’s representation-similarity measure. The finding is about specific MNIST-based experiments—not proof that networks generally retain permanent memories of every training experience.

What “behavioral convergence” means in this study

Mutlu’s paper uses “behavioral convergence” operationally: two networks meet a predeclared criterion for matching predictive performance. That does not mean they produce identical outputs for every possible input, or that their complete functions are equivalent.

Likewise, “representational convergence” refers to similarity in selected internal layers, measured with centered kernel alignment (CKA). CKA compares patterns of activity across examples. It is a useful way to compare representations, but it does not establish complete model identity or functional equivalence.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The paper’s primary representation-history score is H_repr = 1 - mean(CKA_conv2, CKA_fc1). It averages the CKA comparisons for the Conv2 and FC1 layers, then subtracts that mean from one. A larger score therefore indicates lower similarity for those selected comparisons; it is not a general-purpose measure of how different two networks are in every respect. The repository says logits and Conv1 are excluded from this primary score.

How the training-order experiment worked

The main protocol used a small convolutional network on MNIST-derived tasks. The two models began with identical weights. One learned digits 0–4 (task A) and then 5–9 (task B); the other learned B and then A. Both then received the same balanced 0–9 training distribution (task C), using the same deterministic batch sequence and checkpoint schedule. This setup tests whether a difference associated with task order remains measurable after both models receive a common later training experience.

The repository describes a SimpleCNN and reports paired-run validation, representation analyses, a long-horizon common-relaxation test, fresh linear probes, activation-function controls, and same-label rotated-MNIST controls. These details matter because the conclusion belongs to these protocols rather than to neural networks as a whole.

What the authors measured

Similar predictive performance, with residual representation differences

In the abstract, Mutlu reports that 16 of 20 paired runs met the predeclared behavioral-matching criterion. Across the reported comparison, the mean representation-history score was 0.139, with a 95% bootstrap confidence interval of 0.127–0.153, and prediction disagreement was about 3.1%. These are results from the study’s tested models and evaluation, not population-wide estimates. Read the arXiv abstract.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A 50,000-update stress test

In a long-horizon test of five paired seeds, the authors report that after 50,000 common optimizer updates the mean representation-history score was 0.190 (95% bootstrap CI 0.161–0.219), while the mean accuracy gap was 0.18 percentage points. This shows a measurable difference under the selected layer comparisons over that training horizon; it does not show that the difference persists forever.

Controls and a possible influence of activation choice

A same-label rotated-MNIST control reached behavioral matching across five paired seeds while retaining a mean representation-history score of 0.162. In a matched-learning-rate ReLU/LeakyReLU control, the 50,000-update representation residue was reduced by about 0.040 across five paired seeds. That directional result is consistent with activation-mediated plasticity contributing to the outcome, but it does not establish a causal mechanism.

Do different representations mean worse downstream performance?

Not necessarily. Mutlu reports that fresh linear probes with sufficient labeled data found practically equivalent linearly accessible class information in the two histories. The repository specifies a ±0.5-percentage-point equivalence margin for its endpoint using 500 examples per class. This result is limited to that probe setup: it does not show identical representations, and it does not rule out differences when a readout has less labeled data.

The distinction is important: accuracy asks how well a model predicts on an evaluation task; CKA asks how similar selected internal activity patterns are; a linear probe asks what a particular simple readout can extract from those representations. Agreement on one question does not settle the others.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the result does—and does not—establish

  • It establishes a protocol-specific result: in the tested small-CNN, MNIST-derived experiments, similar predictive performance could coexist with measurable differences in selected layer representations after shared training.
  • It does not establish a universal law: the evidence does not show that transformers, large models, or networks trained on other data behave the same way. The sources do not establish independent replication.
  • It does not prove permanent memory: the long-horizon observation covers 50,000 common updates in five paired seeds, not unlimited training.
  • It does not isolate the mechanism: the repository cautions that AB-versus-BA differences can overlap with catastrophic forgetting and ordinary last-task effects. The results alone do not prove universal hysteresis or causality.
  • It does not prove separate optimization basins: the repository notes that its weight interpolation is raw and not permutation-aligned, so a linear barrier in that analysis cannot demonstrate full basin disconnection.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What would make the evidence more general?

Useful follow-up comparisons would vary architecture and scale; dataset and whether task histories differ by labels or input domain; the duration and schedule of shared training; the representation metric and layers selected; the downstream readout and amount of labeled data; and the number of seeds and uncertainty estimates. These are open comparison axes, not findings already demonstrated by this study.

Reproducing the experiments

The author’s public repository provides code, configurations, result manifests, paper artifacts, and reproduction commands. Its README describes setting up a Python virtual environment, installing dependencies from requirements.txt, and running paired training configurations and validation; training downloads MNIST if it is not already available. Consult the repository for the exact commands and configuration details: Mutlu’s code and reproducibility repository.

The repository warns that differences in hardware, PyTorch, and CUDA can affect reproducibility, and says environment metadata is recorded when available. It identifies generated experiment outputs as the underlying source of truth.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.