iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
An Android phone can serve as the local terminal for a machine-learning workflow, while Kaggle provides competition data and a place to submit predictions. In a September 22, 2026 case study, Malcolm Low reports using Termux, a Debian/Ubuntu userspace, Antigravity CLI, and the Kaggle CLI to train a 10-fold Titanic ensemble. The reported clean submission scored 0.80382 on Kaggle’s public leaderboard, but the setup and results are the author’s account—not an independently reproduced benchmark.
What the Android workflow involved
Low describes running the workflow on an Android 14 phone. Termux provided the Android terminal environment, with a Debian/Ubuntu userspace running through PRoot. Antigravity CLI orchestrated the coding-agent work, Python handled the machine-learning code, and the Kaggle CLI was used to interact with the competition.
The author reports using Python 3.14 and ARM64 builds of CatBoost, scikit-learn, pandas, and NumPy. Those are reported implementation details, not confirmation that the same package combination installs or runs reliably on other Android devices. The available account does not independently verify installation under PRoot.
Recommended Free Tools
What each part did
- Termux and PRoot: supplied the phone-based terminal and Linux userspace described in the case study. The Termux project calls its app an Android terminal emulator and Linux environment that works without rooting.
- Antigravity CLI: served as the coding-agent interface for shell work and Python development in the author’s setup. Google describes Antigravity CLI as a terminal-first interface for working with agents, but its official product materials do not document Android or Termux PRoot support. Treat this configuration as the author’s reported use, not a Google-supported Android setup.
- Kaggle CLI and competition: connected the local workflow to Kaggle for competition interaction. Kaggle’s documented cloud notebooks are a separate option, with versioned notebook environments, attached data sources, and configurable accelerators.
A Bluetooth keyboard is an optional comfort accessory for longer terminal sessions; Termux documents keyboard and external-display use, but neither is required for the workflow.
#1 Best Overall
How the 10-fold validation worked
In k-fold cross-validation, the training data is divided into k parts. For each round, a model trains on k−1 parts and validates on the remaining part; the held-out part changes until each part has served as validation data. The reported out-of-fold (OOF) metric summarizes predictions made on those held-out portions. Scikit-learn describes this procedure and notes that cross-validation can be computationally expensive.
For 10 folds, the training data is split into ten portions and the validation cycle is repeated across them. This uses training data efficiently for validation, but it is not a guarantee of generalization: the result depends on the data, split design, features, and modeling choices. It also does not replace a genuinely untouched final test set.
Rank #2
What “blend” means here
The case study describes a fixed weighted blend: component models produce probabilities, which are combined using predetermined weights. That differs from stacking, where a separate model is trained to combine the base models’ outputs. Scikit-learn documents stacking as an ensemble approach; the reported Titanic system is described as a weighted blend, not a learned stack.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsThe reported model and results
Low reports a blend of CatBoost, Random Forest, and Extra Trees. The weights and selected configuration details below are the author’s reported choices, not independently validated recommendations.
| Component | Blend weight | Reported settings |
|---|---|---|
| CatBoost | 60% | Depth 4; L2 regularization 4.0 |
| Random Forest | 25% | Depth 5; minimum leaf size 2 |
| Extra Trees | 15% | Depth 5 |
The article reports these results for its clean 10-fold blend:
| Measure | Author-reported result | What it represents |
|---|---|---|
| 10-fold OOF accuracy | 0.8698 (86.98%) | Validation predictions across held-out training folds |
| 10-fold OOF ROC-AUC | 0.9029 | Validation ranking metric across held-out training folds |
| Kaggle public score | 0.80382 | Public leaderboard score; the account does not establish the snapshot date |
| Kaggle public rank | 291 of 10,058 competitors (reported as top 2.89%) | Leaderboard position stated by the author |
The public score and rank should not be read as the same measurement as OOF accuracy or ROC-AUC. The account does not independently establish the leaderboard snapshot or verify the Kaggle submission.
The author’s 5-fold comparison
The article also compares the 10-fold blend with a 5-fold CatBoost model that included a group-survival feature. In the author’s experiment table, both have 0.8698 OOF accuracy; the 5-fold model has a 0.79904 public score and rank 492, while the 10-fold blend has a 0.80382 public score and rank 291. This is one reported comparison. It does not establish that ten folds or blending generally improves performance.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Why the higher historical-override score is not the clean result
Low reports a separate historical passenger-override experiment with a 0.83253 public score and rank 98, stated as top 0.96%. The author labels that experiment disqualified because the overrides used information in a way that constituted leakage. It is not the clean model’s result and should not be treated as a recommended shortcut.
The underlying issue is evaluation contamination: repeatedly changing a model or its predictions based on test-set performance lets information from that test set influence the system. Scikit-learn warns that this makes the evaluation cease to measure generalization. In this case, the audit and the label of leakage are the author’s account; they are not an external adjudication of the submission.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Phone-based execution or Kaggle notebooks?
These are different execution choices, not competing model types. In the case study, the phone is the local host; Kaggle supplies competition interaction and the submission destination. Kaggle-hosted notebooks instead provide a cloud environment with versioned notebooks, data attachments, and configurable accelerators.
| Consideration | Local Android setup in the case study | Kaggle-hosted notebooks |
|---|---|---|
| Compute and hardware | Runs on the phone’s reported ARM64 environment; the account does not establish runtime or resource limits. | Kaggle documents configurable notebook accelerators; availability and configuration depend on the notebook environment. |
| Packages and compatibility | Author reports Python 3.14 and ARM64 scientific packages; independent installation verification is not established. | Kaggle documents versioned notebook environments and Python or R notebooks and scripts. |
| Competition connection | Author reports using Kaggle CLI for competition interaction. | Notebook environment is hosted by Kaggle and can use attached data sources. |
| Reproducibility | The report gives a platform and selected model settings, but does not independently verify a complete reproducible environment. | Kaggle describes versioned notebook environments for reproducible and collaborative data science. |
Choose local execution if using a phone terminal is itself useful and you are prepared to manage platform-specific package compatibility. A hosted notebook is the more direct option when you want Kaggle’s documented notebook environment or need its configurable accelerators. Either way, compare models using held-out data and keep the final evaluation genuinely separate from tuning.
Quick Recap
What this case study does—and does not—show
- It demonstrates a reported workflow in which an Android phone hosted local command-line work while Kaggle handled competition interaction.
- It gives one author-reported set of OOF and public leaderboard results, plus model weights and selected settings.
- It does not independently verify the phone installation, reproduce the experiment, or establish a general performance advantage for Android, ten-fold validation, or blending.
- Its higher score from historical passenger overrides is explicitly labeled leakage by the author, not a clean performance result.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

