Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
UMKM-Bench is a small, author-built test of whether language models can understand informal Indonesian shop messages and answer from a shop’s own catalog, shipping table, and policies. In results reported by its author, typo-heavy messages affected some models more than others, while a “don’t guess” instruction improved trap-message results for some models. The scores are specific to this hand-written test—not a general ranking of models for Indonesian customer support.
What UMKM-Bench tests
In an article published September 25, 2026, DEV Community author Rizky Nanda Pratitia describes testing eight language models on messages to three fictional shops: Kopi Lereng in Sleman, Yogyakarta; Sekar Hijab in Bandung; and Dapur Bu Tini in Semarang. Each shop has a supplied knowledge base with product catalog information, shipping tables, and policies. Read the author’s article.
The benchmark is about two linked tasks, not just whether a model can write fluent Indonesian:
- Message understanding: infer the customer’s intent, product or SKU, quantity, and city from compressed or noisy text.
- Grounded response: decide whether the supplied shop data can answer the question and draft a customer-facing reply without inventing missing details.
The author wrote 84 customer messages in three forms: formal Indonesian, manually written slang, and typo-noisy slang. The examples cover prices, stock, shipping, cash on delivery (COD), and opening hours; some deliberately ask for information absent from the shop data. After a pilot, the author added 24 harder messages involving calculations, misleading context, a fake discount, and a prompt injection.
#1 Best Overall
For example, benchmark prompts include “kak arabika yg setengah kilo ready gk?” and “ongkir ke bpp brpmin.” In the report’s examples, “yg” means “yang,” “gk” means “nggak,” and “bpp” refers to Balikpapan. Another prompt asks, “lumpia frozen iso dikirim kendal ra bu? cedhak semarang kok”; the author notes “wingi” as regional vocabulary meaning “yesterday.” These are constructed test inputs, not documented messages from real shops or customers.
The model returns structured fields for intent, product/SKU, quantity, city, answerability, and a reply. The author says scoring is Python-based: half concerns understanding and half grounding, with answerability and factual response content checked against the supplied shop data. As the author puts it, “I wanted to be able to point at every lost point, so the scoring is plain Python.”
Rank #2
Scores reported for the four message conditions
The following are the benchmark author’s reported scores, not independently reproduced results. The underlying Kaggle and GitHub pages could not be inspected for this article, so the figures should be read as reported results from this run. The author says Kaggle displays 95% intervals and cautions that score differences within those intervals may be noise; the available information does not establish that small gaps are statistically meaningful.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems| Model | Formal | Slang | Slang with typos | No-guard |
|---|---|---|---|---|
| Claude Sonnet 5 | 99.8 | 99.4 | 99.4 | 100 |
| Gemini 3.7 Flash | 99.7 | 99.7 | 99.1 | 100 |
| GPT-5.5 | 100 | 98.9 | 99.1 | 99.4 |
| Gemini 3.1 Flash-Lite | 98.5 | 98.2 | 97.0 | 95.4 |
| Gemma 4 26B A4B | 98.1 | 98.1 | 95.7 | 98.1 |
| Claude Haiku 4.5 | 97.2 | 92.7 | 91.5 | 93.5 |
| gpt-oss-20b | 98.0 | 91.6 | 82.8 | 86.4 |
| GPT-5.4 nano | 92.7 | 89.2 | 78.1 | 90.7 |
The article does not define the score scale in the accessible account, so these values are presented as reported rather than translated into a claim about accuracy or real-world success rates. “No-guard” refers to a separate task on 27 difficult slang messages after removing the system instruction to avoid guessing and tell the customer an admin will check; it is not another message-writing style directly comparable to the first three columns.
Rank #3
What changed when the messages included slang and typos
The author reports that the top three models in this test shifted little between formal input and typo-noisy input. Other models had larger drops: gpt-oss-20b moved from 98.0 on formal messages to 82.8 on slang with typos, while GPT-5.4 nano moved from 92.7 to 78.1. Those differences describe this test’s inputs and setup, not a prediction for every Indonesian region, conversation type, or deployment.
The examples show why short messages can be difficult to parse. Abbreviations such as “bpp” can be ambiguous: the report describes an instance where it was interpreted as a district in Yogyakarta rather than Balikpapan. Regional terms such as “wingi” add another layer. A model may produce polished language yet still misunderstand a destination, product, or requested quantity—errors that matter when the reply depends on shipping rates or stock data.
Rank #4
What the no-guessing instruction did
For its 27 difficult slang messages, the author compared results with and without an instruction to avoid guessing and say an admin would check. The report says gpt-oss-20b’s trap-message failure rate rose from 11% to 19% when the instruction was removed; Claude Haiku 4.5’s rose from 4% to 15%. It reports zero trap failures for GPT-5.5, Sonnet 5, and Gemini 3.7 Flash both with and without that line.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11This suggests that an explicit abstention instruction can affect a model’s behavior in this narrow test. It does not establish that a prompt alone makes a customer-support system safe. A real shop would still need checks on answers, clear handling for missing information, and a reliable way to route uncertain cases to a person.
Best Value
The author also describes a non-numeric hallucination about fabric quality that a numeric-only grounding check could miss. In that example, the answerability flag caught the issue. That distinction matters: checking whether numbers match a catalog will not catch every unsupported descriptive claim.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Reported cost and response time
The author gives approximate costs per 1,000 customer messages from this benchmark run and one latency observation. These are not established as current provider prices or reproducible bills; actual costs can depend on the prompt, token use, model access, and provider terms.
| Model | Author-reported cost per 1,000 messages |
|---|---|
| GPT-5.5 | About $12 |
| Claude Sonnet 5 | About $6.50 |
| Gemini 3.7 Flash | About $3 |
| Gemini 3.1 Flash-Lite | About $0.40 |
In the reported run, Gemma 4 26B A4B took about 25 seconds per reply and used around 1,200 thinking tokens. This is a single benchmark observation, not a general latency guarantee.
What the benchmark can—and cannot—tell a shop
UMKM-Bench is useful as a focused example of evaluating informal-language understanding and grounding together. Its fictional businesses have explicit product and policy data, and its deliberately unanswerable questions test whether a model can decline to invent a reply. Its practical value is in showing what to test, not in supplying a universal winner.
- Small, hand-written sample: the author wrote all 84 messages. That is not a representative sample of real UMKM conversations or all regional Indonesian usage.
- Single-turn setup: it does not establish how models handle a customer clarifying details over multiple messages, correcting an earlier typo, or switching topics.
- Scoring limits: strict penalties for unsupported numbers help expose certain errors, but descriptive hallucinations may escape numeric checks.
- Limited language coverage: the author proposes testing more regional speech, including Minang, Batak, and Makassar slang, along with voice notes and multi-turn chats.
- Reproduction not established: the author links Kaggle and GitHub resources and says a new shop can be added through
data/stores.jsonplus messages. The linked pages could not be inspected here, so the dataset contents, code, license, current availability, and independent reproducibility are not confirmed.
For a shop choosing a support workflow, the relevant comparison is broader than one score: test the model on the shop’s own abbreviations, regional vocabulary, common typos, answerable and unanswerable questions, and policy edge cases. Measure whether it extracts the right product, quantity, and destination; whether replies stay within source data; and whether it hands off uncertainty appropriately. Include cost and response time only under the shop’s own expected prompts and workload. This benchmark does not provide enough evidence to rank the eight models for production use.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

