Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

UMKM-Bench is a small, author-built test of whether language models can understand informal Indonesian shop messages and answer from a shop’s own catalog, shipping table, and policies. In results reported by its author, typo-heavy messages affected some models more than others, while a “don’t guess” instruction improved trap-message results for some models. The scores are specific to this hand-written test—not a general ranking of models for Indonesian customer support.

What UMKM-Bench tests

In an article published September 25, 2026, DEV Community author Rizky Nanda Pratitia describes testing eight language models on messages to three fictional shops: Kopi Lereng in Sleman, Yogyakarta; Sekar Hijab in Bandung; and Dapur Bu Tini in Semarang. Each shop has a supplied knowledge base with product catalog information, shipping tables, and policies. Read the author’s article.

The benchmark is about two linked tasks, not just whether a model can write fluent Indonesian:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Message understanding: infer the customer’s intent, product or SKU, quantity, and city from compressed or noisy text.
  • Grounded response: decide whether the supplied shop data can answer the question and draft a customer-facing reply without inventing missing details.

The author wrote 84 customer messages in three forms: formal Indonesian, manually written slang, and typo-noisy slang. The examples cover prices, stock, shipping, cash on delivery (COD), and opening hours; some deliberately ask for information absent from the shop data. After a pilot, the author added 24 harder messages involving calculations, misleading context, a fake discount, and a prompt injection.

For example, benchmark prompts include “kak arabika yg setengah kilo ready gk?” and “ongkir ke bpp brpmin.” In the report’s examples, “yg” means “yang,” “gk” means “nggak,” and “bpp” refers to Balikpapan. Another prompt asks, “lumpia frozen iso dikirim kendal ra bu? cedhak semarang kok”; the author notes “wingi” as regional vocabulary meaning “yesterday.” These are constructed test inputs, not documented messages from real shops or customers.

The model returns structured fields for intent, product/SKU, quantity, city, answerability, and a reply. The author says scoring is Python-based: half concerns understanding and half grounding, with answerability and factual response content checked against the supplied shop data. As the author puts it, “I wanted to be able to point at every lost point, so the scoring is plain Python.”

Scores reported for the four message conditions

The following are the benchmark author’s reported scores, not independently reproduced results. The underlying Kaggle and GitHub pages could not be inspected for this article, so the figures should be read as reported results from this run. The author says Kaggle displays 95% intervals and cautions that score differences within those intervals may be noise; the available information does not establish that small gaps are statistically meaningful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Model Formal Slang Slang with typos No-guard
Claude Sonnet 5 99.8 99.4 99.4 100
Gemini 3.7 Flash 99.7 99.7 99.1 100
GPT-5.5 100 98.9 99.1 99.4
Gemini 3.1 Flash-Lite 98.5 98.2 97.0 95.4
Gemma 4 26B A4B 98.1 98.1 95.7 98.1
Claude Haiku 4.5 97.2 92.7 91.5 93.5
gpt-oss-20b 98.0 91.6 82.8 86.4
GPT-5.4 nano 92.7 89.2 78.1 90.7

The article does not define the score scale in the accessible account, so these values are presented as reported rather than translated into a claim about accuracy or real-world success rates. “No-guard” refers to a separate task on 27 difficult slang messages after removing the system instruction to avoid guessing and tell the customer an admin will check; it is not another message-writing style directly comparable to the first three columns.

What changed when the messages included slang and typos

The author reports that the top three models in this test shifted little between formal input and typo-noisy input. Other models had larger drops: gpt-oss-20b moved from 98.0 on formal messages to 82.8 on slang with typos, while GPT-5.4 nano moved from 92.7 to 78.1. Those differences describe this test’s inputs and setup, not a prediction for every Indonesian region, conversation type, or deployment.

The examples show why short messages can be difficult to parse. Abbreviations such as “bpp” can be ambiguous: the report describes an instance where it was interpreted as a district in Yogyakarta rather than Balikpapan. Regional terms such as “wingi” add another layer. A model may produce polished language yet still misunderstand a destination, product, or requested quantity—errors that matter when the reply depends on shipping rates or stock data.

What the no-guessing instruction did

For its 27 difficult slang messages, the author compared results with and without an instruction to avoid guessing and say an admin would check. The report says gpt-oss-20b’s trap-message failure rate rose from 11% to 19% when the instruction was removed; Claude Haiku 4.5’s rose from 4% to 15%. It reports zero trap failures for GPT-5.5, Sonnet 5, and Gemini 3.7 Flash both with and without that line.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This suggests that an explicit abstention instruction can affect a model’s behavior in this narrow test. It does not establish that a prompt alone makes a customer-support system safe. A real shop would still need checks on answers, clear handling for missing information, and a reliable way to route uncertain cases to a person.

The author also describes a non-numeric hallucination about fabric quality that a numeric-only grounding check could miss. In that example, the answerability flag caught the issue. That distinction matters: checking whether numbers match a catalog will not catch every unsupported descriptive claim.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reported cost and response time

The author gives approximate costs per 1,000 customer messages from this benchmark run and one latency observation. These are not established as current provider prices or reproducible bills; actual costs can depend on the prompt, token use, model access, and provider terms.

Model Author-reported cost per 1,000 messages
GPT-5.5 About $12
Claude Sonnet 5 About $6.50
Gemini 3.7 Flash About $3
Gemini 3.1 Flash-Lite About $0.40

In the reported run, Gemma 4 26B A4B took about 25 seconds per reply and used around 1,200 thinking tokens. This is a single benchmark observation, not a general latency guarantee.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the benchmark can—and cannot—tell a shop

UMKM-Bench is useful as a focused example of evaluating informal-language understanding and grounding together. Its fictional businesses have explicit product and policy data, and its deliberately unanswerable questions test whether a model can decline to invent a reply. Its practical value is in showing what to test, not in supplying a universal winner.

  • Small, hand-written sample: the author wrote all 84 messages. That is not a representative sample of real UMKM conversations or all regional Indonesian usage.
  • Single-turn setup: it does not establish how models handle a customer clarifying details over multiple messages, correcting an earlier typo, or switching topics.
  • Scoring limits: strict penalties for unsupported numbers help expose certain errors, but descriptive hallucinations may escape numeric checks.
  • Limited language coverage: the author proposes testing more regional speech, including Minang, Batak, and Makassar slang, along with voice notes and multi-turn chats.
  • Reproduction not established: the author links Kaggle and GitHub resources and says a new shop can be added through data/stores.json plus messages. The linked pages could not be inspected here, so the dataset contents, code, license, current availability, and independent reproducibility are not confirmed.

For a shop choosing a support workflow, the relevant comparison is broader than one score: test the model on the shop’s own abbreviations, regional vocabulary, common typos, answerable and unanswerable questions, and policy edge cases. Measure whether it extracts the right product, quantity, and destination; whether replies stay within source data; and whether it hands off uncertainty appropriately. Include cost and response time only under the shop’s own expected prompts and workload. This benchmark does not provide enough evidence to rank the eight models for production use.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.