DEV Community author Fortaki says they assembled a multilingual and programming-code dataset for Central Asian AI that was nearly 150 GB uncompressed and about 27.7 GB as a compressed archive. The maximum-compression run took ten hours and involved repeated out-of-memory errors and system freezes. The account describes the scale and the trouble, but does not provide the hardware or a successful, reproducible configuration—so it is a useful project report, not a step-by-step recipe for avoiding memory failures.
What the dataset is—and what the size figures mean
Fortaki describes a collection of technical text and source code intended for multilingual AI work involving Central Asian languages. The human-language content is said to cover Russian, Kyrgyz, Kazakh, Uzbek, Tajik, and English; the programming languages named are Python, C++, Rust, and Go. Fortaki reports nearly 150 GB before compression and approximately 27.7 GB afterward. The linked Hugging Face card lists the same total-size figures.
The author’s itemized figures do not reconcile with the headline size: Russian 60 GB, Kyrgyz 23 GB, Kazakh 22 GB, Uzbek 7 GB, Tajik 7 GB, English 1.5 GB, and code 20 GB add up to about 120.5 GB. Neither page explains the roughly 29.5 GB difference. The figures should therefore be treated as reported estimates, not a verified inventory.
| Category | Fortaki’s reported size |
|---|---|
| Russian | 60 GB |
| Kyrgyz | 23 GB |
| Kazakh | 22 GB |
| Uzbek | 7 GB |
| Tajik | 7 GB |
| English | 1.5 GB |
| Programming code | 20 GB |
| Sum of itemized figures | About 120.5 GB |
| Reported total before compression | Nearly 150 GB |
| Reported compressed archive | Approximately 27.7 GB |
These are figures reported by Fortaki in the 2026 DEV Community account; the Hugging Face card also lists 150 GB total size and 27.7 GB total file size. The itemized language-and-code breakdown is not reconciled with the overall total in either page.
#1 Best Overall
Why the compression run became a memory problem
Fortaki says the archive was created with maximum-profile .7z compression and took ten hours. The author describes repeated out-of-memory errors and operating-system freezes during the effort. This establishes what happened on that machine, but not why a particular failure occurred or what setting would prevent it for another user: the account does not state the computer’s RAM, CPU, operating system, compressor version, exact settings, or a confirmed successful workaround.
Maximum compression can involve a more demanding workload than simply storing files, but the account does not include enough configuration detail to attribute the freezes to a specific setting or hardware bottleneck. It would be misleading to infer a minimum RAM requirement or prescribe a particular compression command from this report alone.
Rank #2
What is documented about the files and layout
The DEV account says the extracted files are .txt and .jsonl, and names Python, C++, Rust, and Go code alongside the six human languages. The Hugging Face card describes folders named /python, /rust, /cpp, /go, and /docs, with train and validation splits. Those are descriptions from the author and card, rather than a confirmed inspection of the archive: the card’s dataset viewer says it could not detect supported data files. Readers should check the actual downloadable files before relying on the stated formats, paths, or splits.
What the dataset might be used for
Fortaki proposes fine-tuning code models to work with comments and technical documentation in Central Asian languages, specialized technical translation, and continued pretraining or domain adaptation. The Hugging Face card also names multilingual model training, evaluation, cross-lingual code understanding, technical documentation, NLP, and cross-lingual information retrieval.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Rank #3
These are intended applications, not demonstrated outcomes. The reviewed pages report no evaluation results showing that the collection improves a model, translation quality, code understanding, or retrieval performance.
License and provenance: what readers should verify
Fortaki says the project is open source and distributed under CC BY 4.0; the Hugging Face card also states CC BY 4.0 and characterizes the contents as open-source components, public-domain texts, and synthetic benchmarks. These are the author and card’s descriptions, not an item-by-item rights audit. Neither page provides a complete source inventory, detailed cleaning or deduplication methods, or evidence that every upstream source permits redistribution under the stated arrangement.
Rank #4
Before using or redistributing the files, inspect their source and licensing information and assess whether the terms fit your intended use. The stated dataset license does not by itself establish the rights or provenance of every component.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What this account does—and does not—teach about avoiding OOM errors
The account is valuable as a warning about the operational cost of compressing a large collection at maximum settings: the author reports ten hours, repeated memory failures, and system freezes. It is not a reproducible troubleshooting guide. It does not identify the successful configuration, explain whether the archive was produced in one pass or through another workflow, or document filtering and parsing steps.
That distinction matters if your goal is to build a comparable dataset. Fortaki’s reported outcome gives a rough sense of the data and archive scale, but the missing machine and process details prevent a reliable estimate of the memory, time, or settings another builder will need. The account invites questions about filtering and parsing but does not describe those methods.
Quick Recap
Sources
- Fortaki’s DEV Community account (2026), the source for the reported assembly, size estimates, compression effort, contents, applications, and license statement.
- The linked Hugging Face dataset card, which repeats size and license claims, describes a proposed layout and uses, and reports that its viewer could not detect supported data files.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

