iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
When several WhatsApp alerts are waiting, linking a customer’s reply to the newest alert can be wrong: a message asking “when will the parcel arrive?” may refer to an older order alert, not a newer viewing alert. ChatRail’s Danish Javed added a classifier to choose the matching alert—or say none fits—before the existing newest-alert fallback. In a small author-run test, the 12 Jev calls in a simulation cost $0.00027, but the result is not a production benchmark or a general cost guarantee.
Why recency caused the wrong reply match
ChatRail sends WhatsApp alerts about events such as dispatched orders, booked viewings, and invoices due. When a customer replies, the system needs to associate that message with the alert it concerns.
Its earlier matching logic used the WhatsApp reply button when available. Otherwise, it matched a single waiting alert; if several were waiting, it chose the newest one. With no waiting alerts, it linked nothing. That newest-alert rule ignored the message itself. A customer asking when a parcel would arrive could therefore have a reply attached to a more recent viewing alert.
Recommended Free Tools
The important distinction is between temporal proximity and conversational relevance: the latest alert is a useful fallback, but it is not proof of what the customer meant.
#1 Best Overall
How the revised matching flow works
Javed inserted Jev after the deterministic reply-button and single-alert cases, but before the newest-alert guess. When multiple alerts are open, the model is asked: “Which of these alerts is this message about?” It can select one candidate or answer “none of these.”
- Use the reply button when present. An explicit reply to a WhatsApp alert remains the first match path.
- Use the sole waiting alert when there is only one. This avoids a model call when there is no competing candidate.
- Ask Jev when multiple alerts are open. It chooses a candidate or abstains by selecting none.
- Accept a model choice only at or above the chosen 0.8 confidence threshold. Accepted selections are recorded as
model_choicewith the probability. - Fall back to the newest alert if Jev is uncertain, takes longer than two seconds, or returns an error. If no alert is waiting, the earlier logic links nothing.
This design preserves the existing fast, deterministic paths and limits the model’s role to the ambiguous multiple-alert case. Javed said the feature was live behind a flag and ran only for connections already using managed AI; messages from customers who had not enabled AI were not sent to Jev. He also reported that an outage could not prevent an inbound message from being received because the old matching fallback remained available.
Rank #2
What the reported comparison found
Javed compared Jev 1.13, Gemini 2.5 Flash Lite, and GPT-5.6 Luna on 14 cases he wrote himself. The figures below are his results, not independently verified performance guarantees.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →| Measure | Jev | Gemini 2.5 Flash Lite | GPT-5.6 Luna |
|---|---|---|---|
| Alert-selection accuracy | 93% — Danish Javed, 2026 | 93% — Danish Javed, 2026 | 93% — Danish Javed, 2026 |
| Median end-to-end latency | About 350 ms — Danish Javed, 2026 | About 640–1,000 ms — Danish Javed, 2026 | About 2.1–2.6 seconds — Danish Javed, 2026 |
| Estimated cost per 10,000 messages | About $0.25 — Danish Javed, 2026 | About $0.41 — Danish Javed, 2026 | About $1.20–$1.30 — Danish Javed, 2026 |
The latency measurements were end to end from Javed’s machine through OpenRouter, including network time, and were taken over two runs. The cost figures were charges reported by OpenRouter for the stated 10,000-message workload. They depend on the test setup and should not be treated as universal rates.
Rank #3
Each model selected the right alert in 93% of the 14 cases. In the five cases where the newest alert was wrong, all three models selected correctly, while the old newest-alert rule did not. That comparison suggests the model-based decision can address the specific failure mode; it does not show that Jev was more accurate than the two LLMs.
What the $0.00027 simulation does—and does not—show
In a separate simulation involving 12 contacts, the old rule linked 3 correctly and Jev linked all 12. The simulation reused alert texts from the earlier test, so it was not an independent sample. The Jev calls cost $0.00027 across 12 Jev calls — Danish Javed, 2026. That is the cost reported for those calls, not a general per-issue price or a promise that another workload will cost the same.
Rank #4
The evidence is particularly limited for production use. Javed reported that no production messages had yet reached the multiple-open-alert rule. He designed the 14 test cases himself, many to expose the recency rule’s weakness. The selected 0.8 confidence threshold had not been tested: every Jev selection in the reported runs was at least 0.97. The test therefore does not establish how often the system would abstain or fall back in real traffic.
Alert matching is different from answering the customer
Choosing which alert a reply concerns is not the same task as deciding whether the alert contains enough information to answer the customer’s question. Javed reported a distinct answerability test in which Jev scored 71%, while GPT scored 93–100% across two runs. He kept answerability with the reply model rather than using Jev for that task.
Best Value
This separation matters in system design: a tool that is economical for selecting among a fixed set of alerts is not necessarily the best choice for generating or grounding a reply. Evaluate those jobs independently, with separate test cases and success criteria.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to evaluate a similar reply-matching change
A useful evaluation should cover the specific ambiguity and the costs of a wrong association, rather than relying only on an aggregate accuracy score.
- Test the competing paths: explicit reply-button matches, one waiting alert, multiple waiting alerts, and no suitable alert.
- Include abstention cases: check whether the classifier can choose “none” when every candidate is wrong or incomplete.
- Compare on identical cases: measure the old rule and each candidate model on the same customer messages and alert sets.
- Measure the full path: include network and application latency, not just model response time.
- State the workload behind cost estimates: message volume, model, routing, and call pattern can change the result.
- Exercise fallback behavior: deliberately test low confidence, timeout, and errors so an unavailable model cannot block message handling.
- Validate against production traffic: a hand-built set can reveal a known defect, but real conversations are needed to assess performance beyond those examples.
Javed’s account is a useful illustration of a narrow model-assisted decision: preserve clear deterministic matches, ask a fixed-choice question only when ambiguity exists, and retain a fallback. Its results support investigating that design, not assuming that any model or confidence threshold will work equally well elsewhere.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

