PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
The bug was not fixed. What GeneLab_999 did, in a write-up published on DEV Community on September 24, 2026, was narrow a surprising model behavior to a measurable cause and give the laya maintainers a check they could run. The laya multilingual checkpoint selected the first-listed option zero times across 300 Japanese and 290 English score items. The author’s tests indicate that the pattern follows the option’s slot position rather than a Japanese label or a flaw in the test harness. The change that was merged, PR #259, adds an offline regression check for that narrow failure. It does not retrain the model, and it does not show that the model’s urgency predictions are accurate. Retraining was still pending in the September 2026 records.
What laya is and what was tested
laya is described in the project as a non-autoregressive “System 1” decision model. Rather than generating a text response, it takes a passage plus typed questions and returns answers and probabilities in one forward pass. The question types that matter here are choice (pick one of several labels), score (an ordinal scale, such as urgency), and yes/no. The write-up refers to English and multilingual checkpoints.
The experiment began as a Japanese-language baseline for a separate project. The author wrote 300 Japanese business emails and 290 English ones. Labels were fixed before a local language model generated an email to match each one, and any email containing a label word was rejected and regenerated. Every email carried three questions: a department choice, an ordinal urgency score, and a cancellation-intent yes/no.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
These are author-built synthetic benchmark items, not a published or representative dataset. The figures below describe this benchmark only and should not be read as general accuracy for laya or for any other application.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
The symptom: a score task that never picked the lowest option
The baseline results differed sharply by question type.
| Question type | Metric (Japanese benchmark) | Model result | Baseline |
|---|---|---|---|
choice |
Accuracy | 0.747 | 0.380, majority-class |
score |
RPS (lower is better) | 0.232 | 0.197, majority-class |
| yes/no | Accuracy | 0.543 | 0.703, majority-class |
| yes/no | AUROC | 0.523 | not stated |
The write-up labels the 0.197 score baseline majority-class; the GitHub issue reports the same figure as a random baseline. Each row comes from the author’s own benchmark, reported by GeneLab_999 in September 2026.
The score result prompted the investigation. In the Japanese score task, the lowest option, “not urgent,” was the correct label for 77 of 300 emails, yet the model never predicted it. The question was whether the cause was the wording of the options, their order, or the number of levels.
Position, label, or harness?
The author tested three competing explanations in sequence: the option wording and order, the language, and the measurement harness itself.
Changing order, wording, and number of levels
The author ran five schema conditions on the Japanese items: original, reversed, reworded, reworded and reversed, and a four-level version. Across all five, the first-listed option was selected 0 or 1 times out of 300. In the original and reversed schemas, “not urgent” was chosen 0 times when it was listed first and 250 times when it was listed last. The label was not failing on its wording; the model was ignoring whichever option sat in the first slot.
Rank #2
Replicating on English emails
To test whether the problem was specific to Japanese, the author ran the same five conditions on the 290 English emails. The multilingual checkpoint chose the first-listed option zero times in every condition. The English checkpoint did not show this pattern in every condition, as the table below shows.
| Checkpoint | Japanese, n=300 (five conditions) | English, n=290 (five conditions) |
|---|---|---|
laya-multilingual |
0, 0, 1, 1, 0 | 0, 0, 0, 0, 0 |
English laya |
13, 56, 8, 1, 110 | 65, 74, 0, 5, 4 |
Each count is the number of times the first-listed option was selected, in the order original, reversed, reworded, reworded and reversed, and four-level. The English replication was run on laya 0.3.4, loaded through the README’s laya.load() and called with agent.predict(). In that setup, the multilingual checkpoint’s English score RPS was 0.340 against a 0.197 baseline, and its English yes/no AUROC was 0.355. These numbers are reported in GitHub issue #131, which the author opened on September 22, 2026.
Shuffling each item separately
A global reordering could still leave the label and the slot entangled. The author therefore shuffled the option order independently for each item. In the Japanese run, the first slot was selected 0 times out of 300, while slots two and three received 149 and 151 choices. Each label occupied the first slot on 90, 109, or 101 items, depending on the label. The author reads this as showing that the lost selections followed slot 1 rather than any particular label, with no comparable preference for the last slot.
Controls contributed by AlKor13
The author credits AlKor13 with examining the raw marker logits and testing three identical options. As the author reports it, changing only the checkpoint made the position effect appear or disappear, and the multilingual checkpoint showed a strong position effect in the identical-option control. AlKor13 also found that removing the level N: prefix eliminated the slot-0 suppression in the raw logits. The author notes that this prefix-free format is off-distribution for the model, so it cannot serve as a remedy. The write-up does not cite an independent reproduction of these controls, so they should be read as the author’s account of AlKor13’s work.
Why rewording was not a fix
The author compared three renderings of the score options: the shipped level N: format, a version without that prefix, and word ordinals. The paired comparisons on identical examples gave mixed results. Two of the four language-and-rendering comparisons were statistically significant (paired McNemar p-values of 0.145, 0.0007, 0.0003, and 0.350).
The changes were large at the item level. Under the prefix-free rendering, 56.7% of Japanese items and 56.9% of English items changed correctness. Headline accuracy moved by 6.6 points for Japanese and 16.2 points for English. A control also showed that removing the prefix lowered English-checkpoint accuracy from 0.583 to 0.500. The author’s conclusion is that the rendering effect was unstable and specific to each checkpoint. The write-up does not support the claim that dropping the prefix fixes the problem.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsThis is the main methodological lesson of the case. A change that moves the aggregate number a little can still reshuffle a large share of individual predictions, so a single accuracy figure is not enough to judge it.
The regression check in PR #259
PR #259 adds research/eval/presentation_checks.py and a set of offline regression tests. The script runs two checks, and each has its own threshold. It also verifies the harness before it judges the model.
Check 1: slot-0 marker logit
When options have identical text, this check compares the raw slot-0 marker logit with the mean across slots. It asks whether slot zero is unusually suppressed. The pass threshold is at least −0.20.
Check 2: first-slot selection rate
This check presents three real levels in all six permutations for each of ten fixed English support messages, then measures the proportion of decisions that select the first slot. The pass threshold is at least 0.15.
Rank #4
Harness verification
Before running the gates, the script compares its own inference path with Agent.system_one. A failed gate and a harness mismatch return separate exit codes, so a broken harness cannot be mistaken for a model failure, or the reverse.
Results on the documented setup
The results below come from a CPU, fp32 run with laya 0.3.7, the setup the PR documents. They are not a claim about all runtimes.
| Checkpoint | Slot-0 logit metric (threshold at least −0.20) | First-slot rate (threshold at least 0.15) | Outcome |
|---|---|---|---|
| English checkpoint | +0.664 | 0.217 | Passes both gates |
| Multilingual checkpoint | −0.492 | 0.017 | Fails both gates |
The PR also reports maximum probability parity differences of about 4.98e-5 and 4.92e-5 against the package inference path, which indicates that the harness reproduces the package’s outputs on this setup.
The maintainer’s response in the PR discussion was direct: “Thank you, a label-free check that answers one question (did the retrain remove the score slot prior?) is exactly what #131 needs, and exit codes that separate a failed check from a harness disagreement make it easy to trust. Merging.” The quote is attributed to NandhaKishorM, a laya maintainer.
What the check does not establish
The PR itself is explicit about its limits, and they matter as much as the result.
Best Value
- The check uses ten short English messages, covers
scorequestions only, and runs on CPU in fp32. It says nothing about Japanese, other question types, or other runtimes. - Passing is not accuracy. The check tests only the targeted positional behavior and does not measure whether any answer is correct.
- It does not train a replacement. Issue #131 remained open in the PR discussion pending a position-balanced multilingual checkpoint, and the write-up likewise describes retraining as still pending.
- It does not establish that the model’s urgency predictions are accurate, even on the multilingual checkpoint after a retrain. A retrained model would still need task-level evaluation.
- The tests are standalone scripts rather than pytest tests. Model checkpoints are not loaded in CI, and the offline tests are not registered in CI because a discussion about wiring research tests into CI was unresolved at the time.
The repository’s status may have changed since these September 2026 records. Check issue #131 and PR #259 directly for the current state of the retrain and the CI wiring.
Working inside a repository you do not maintain
The author describes a set of scope decisions that kept the contribution acceptable. PR #259 changes nothing under laya/ and adds no dependencies. Another contributor was already building a broader option-permutation framework, so the author kept this check focused on issue #131 rather than duplicating that work.
The author’s own summary of the approach is worth keeping in view: “The fastest way I’ve found to contribute to an ML repo you don’t maintain: measure it, make the measurement impossible to explain away, and then hand the maintainer a tool.”
The case also shows how to handle the predictable objection. The author anticipated a reader’s “Your harness is wrong,” and answered it with the parity check and the separate exit codes rather than with an argument.
Quick Recap
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

