Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Jean-Luc Martel’s experiment did not rebuild Encarta: it reconstructed an LHA -lh5- archiver from observed behavior, with the original implementation withheld as a grading key. Its results show a precise divide: the decoder reproduced tested round trips, but the encoder almost never matched the original compressed bytes on cases that exercised its compression decisions. That is a useful case study in what a test suite measures—not a general failure rate for AI-written code.

What the Encarta title refers to

The exact title, “Rebuilding Encarta showed me exactly where AI-written code breaks,” appears on a DEV Community tag page. The detailed experiment, however, is part of Martel’s broader series on reconstructing legacy systems with AI. The software target was LHA’s -lh5- compression method, not Encarta itself. The title should therefore be read as framing for the author’s investigation, not a description of the code he reconstructed. DEV Community’s Encarta tag listing

How Martel set up the test

Martel says the reconstruction started without a specification or source code. The models could query an oracle that returned outputs for inputs, while the original implementation was kept private as a grading key until the reconstruction was frozen. He describes -lh5- as LZSS compression with an 8 KB window followed by static Huffman coding. He chose it in part because the original could serve as an oracle, the algorithm had a public answer key, and compressed output was not uniquely determined by the decompressed bytes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The work was split across models: Martel names Gemini 3.1 Pro for the decoder, Codex/GPT-5 for the encoder and a cold-recall baseline, and Claude in the design discussion. Those model assignments and all results below are the author’s account, not independently verified or reproduced measurements. Martel’s article on DEV Community

#1 Best Overall
Encarta 97 Encyclopedia (Windows 95)
  • Used Book in Good Condition

What the reported scores mean

Evaluation Author-reported result What it tests
Decoder round trips 19 of 19 exact Whether decoding and re-encoding the tested cases recovered the expected data, as reported by Martel in 2026.
Trained encoder cases 1 of 12 byte-for-byte matches (8.3%) Whether the reconstructed encoder emitted the same compressed bytes as the original for cases used during reconstruction, as reported by Martel in 2026.
Held-out encoder cases 5 of 7 byte-for-byte matches (71.4%) Whether the output matched exactly on cases held out from reconstruction, as reported by Martel in 2026.

These measures answer different questions. A decoder can produce the right original data even when its implementation or encoding choices differ. Byte-for-byte comparison is stricter: it also tests whether the encoder makes the same choices as the reference. The higher held-out score does not show that the encoder generalized better. Martel says those cases were mostly random, incompressible, or trivial inputs, where stored-mode or other simple paths bypassed the harder compression heuristics. The trained cases included text, source code, and structured data that exercised those decisions.

Martel also reports identical rates for trained-seed and fresh-seed corpora, which he interprets as systematic divergence rather than overfitting to individual examples. The evidence supports a narrower reading: matching was concentrated in cases that did not put the encoder’s difficult choices to the test.

Where the encoder diverged

Huffman code lengths

The clearest mismatch involved assigning Huffman code lengths. The reconstruction used canonical assignment; Martel says the original assigned lengths in heap-extraction order, with ties determined by its exact sift-down comparisons. When symbols have equal frequencies, those choices can change code lengths and then cascade into a different bitstream. The reconstruction identified this source of divergence but did not reproduce the original sift order.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Matching behavior and untested implementation details

Martel says the reconstruction correctly inferred nearest-offset tie-breaking and one-step lazy matching when compared with the unsealed source and the recorded prior baseline. It did not model a match-finder chain cap, but that hidden detail did not affect the tested corpus. The corpus topped out at 8 KB, so it also did not trigger block splitting at a 32 KB buffer threshold. That path remains untested—not a demonstrated failure or success.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What this experiment can—and cannot—tell you

The experiment does not establish how often AI-written code fails in general. It is one author-reported reconstruction, using one algorithm, one oracle setup, and a limited corpus. Its value is in showing how the apparent answer changes with the evaluation criterion and input selection.

  • Separate data correctness from byte identity. A decoder round trip can pass while an encoder still differs from the reference.
  • Make tests reach difficult behavior. Random or incompressible data may take a simple path and say little about compression heuristics.
  • Treat untriggered paths as unknown. Passing tests establish behavior only for the inputs and boundaries actually exercised.
  • Freeze the reconstruction before comparing with the answer key. Martel describes a tagged commit, sealed source, and manifest check intended to reduce post-hoc contamination.
  • Record prior knowledge before querying. A cold-recall record can help distinguish remembered behavior from behavior inferred through experiments. Martel notes a limitation: the encoder and the recall record came from the same model, leaving a theoretical shared-prior concern.

As Martel puts it: “A strong oracle over a narrow corpus hides exactly the mechanisms your corpus never triggers, and it hides them silently, because everything it can see is green.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.