Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
The Laya approach adds domain-specialized decision heads while keeping the model’s encoder frozen and retaining its original head as a fallback. Its author reports an accuracy increase from 59.6% to 67.3% on a small, author-labeled evaluation, but the result is preliminary: the test set was limited, the router was designed after it had been inspected, and independent validation and real-data calibration were not reported.
What is the frozen-encoder MoE on Laya?
In this design, “mixture of experts” means selecting among complete decision heads for a request. It is not the token-level sparse mixture-of-experts architecture often used inside large language models. The encoder is shared and frozen; a router selects a specialized head for some inputs, while the original head remains available for fallback. The proposal is described in an open preprint by Vishal Mysore that explicitly says it has not been peer reviewed. Read the preprint.
The underlying laya-typed-decisions model is described as having a 421-million-parameter ModernBERT-large encoder and a two-layer, 26.5-million-parameter decision head. Each expert is a copy of that decision head trained for a domain group. The encoder is not retrained for the experts.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →How does routing and expert training work?
Routing at inference
The system batches a router question and the user’s questions through the shared encoder. The original decision head answers the router question about the input kind; a fixed mapping sends the request to a specialized head where one is assigned. Unmapped or low-confidence cases use the original head. The author says those fallback inputs preserve the base model’s outputs.
#1 Best Overall
Training the expert heads
Expert heads are trained on synthetic, rule-labeled examples for their domain groups. The training process caches features from the frozen encoder and updates only the decision head. The author reports CPU-only training and use of cross-entropy plus a ranked probability score for ordinal questions. These are the author’s described methods, not independently reproduced findings.
What results does the author report?
The reported evaluation compares the general laya-typed-decisions model with the routed mixture. The author reports 312 questions across nine domains, grouped into 108 hand-written cases:
Rank #2
| Measure | General model | Routed mixture |
|---|---|---|
| Overall accuracy on the author’s 312-question evaluation (Vishal Mysore, 2026) | 59.6% | 67.3% |
| Accuracy on domains assigned an expert (Vishal Mysore, 2026) | 55.2% | 67.7% |
| Accuracy on uncovered domains (Vishal Mysore, 2026) | 66.7% | 66.7% |
The unchanged score on uncovered domains is consistent with the stated fallback design, but it does not establish that behavior will hold for other data. The author also reports 80.6% routing accuracy for a fine-grained kind-level router and 49.1% for an expert-level router. An int8 ONNX browser build is reported at 67.9% accuracy versus 67.3% for PyTorch on the same evaluation.
How strong is the evidence?
The results are preliminary experimental findings, not a general performance guarantee. The author wrote and labeled the 108 cases and notes judgment calls; the 312 questions are clustered within those cases, so they are not 312 independent observations. The fine-grained router was designed after the evaluation set had been seen, a threat to validity acknowledged by the article.
Rank #3
- Independent labels, evaluation across multiple random seeds, and calibration tests on real data were not reported.
- Improvements were concentrated in score and yes/no questions; choice-question accuracy declined.
- Because experts learn from synthetic rules, errors or assumptions in those rules can be carried into the trained heads.
- The reported browser and PyTorch scores come from the same small evaluation, not a separate validation study.
The article identifies larger independently labeled tests, a held-out router evaluation, real-data calibration, multiple seeds, comparisons with full fine-tuning and LoRA, and checks for synthetic-data artifacts as needed next steps. Those are proposed experiments, not completed results.
What would matter when assessing this design?
The approach’s appeal is that it adds domain-specific decision heads without duplicating or retraining the encoder, while leaving the base head in place for inputs without a suitable expert. Whether that trade-off is useful depends on more than aggregate accuracy:
Rank #4
- Router reliability: A wrong route can negate the benefit of a good expert. Test routing on data not used to design the router, including ambiguous and uncovered inputs.
- Per-domain and per-question-type behavior: Review scores for each domain and question type; aggregate gains can conceal declines such as the reported reduction for choice questions.
- Calibration: Accuracy alone does not show whether confidence scores are trustworthy. Real-data calibration remains unreported.
- Resource and training trade-offs: Shared frozen features avoid retraining the encoder, but each expert adds a decision head. The article does not establish comparative memory or cost against full fine-tuning, LoRA, or separate per-domain models.
- Data dependence: Synthetic rule-labeled examples make it important to check whether apparent gains reflect useful specialization or artifacts of the rules.
Can the work be reproduced?
The author provides a code repository, expert weights and browser build, and a live demo, along with instructions for creating an environment, evaluating the baseline, generating synthetic training data, caching encoder features, training heads, evaluating the mixture, and exporting and evaluating the browser build. Artifact and demo availability may change; consult the linked project materials for current access and commands.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThe author invites scrutiny and replication: “Negative results and failed replications are as welcome as confirmations, and every replication will be linked from the repository.”
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

