Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A fraud score can look impressively predictive while measuring the bank’s alerting process rather than fraud itself. In a project reported by Syed Darain Qamar, the bank’s score had a ROC-AUC of 0.053 across 5,565 closed investigations—an apparent inversion that Qamar argues should not be “fixed” by simply reversing the score. The more useful lesson is to examine how cases were selected, then make decisions from inspectable behavioral evidence and explicit policy.

Why did the bank’s score appear to rank fraud backwards?

Qamar’s project began with six months of card transactions, 5,565 closed investigations, a fraud policy, and twenty alerts. The transaction data did not contain fraud labels, according to the author. Instead, the team compared the bank’s detection score with the outcomes of cases that had actually been investigated.

Across those 5,565 closed cases, Qamar reports that the bank score’s ROC-AUC was 0.053. ROC-AUC measures how well a score ranks positive cases above negative ones; a result this low appears close to the reverse of the desired ranking. But the observed cases were not a neutral sample of transactions: the score helped determine which alerts were opened. Qamar says high-scoring alerts often proved to be legitimate purchases, while confirmed fraud could arrive through customer reports and appear at low scores.

That selection process changes what the metric means. The score may have been useful for generating or routing alerts, yet perform poorly when evaluated only among already-investigated cases. The cases reveal the outcomes of an alerting workflow, not necessarily the distribution of fraud across all transactions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why not just reverse the score?

Qamar reports that reversing the score reached 93% on a balanced October holdout. He rejects that result as a dependable solution: the balanced benchmark and its trigger-score range created a particular sampling frame, so the reversal could fit how that benchmark was constructed rather than identify fraud in a broader population. A striking metric is not evidence of generalization when the evaluation sample does not match the intended use.

What remained after the score was set aside?

The project shifted attention from a bank-wide risk score to behavior that could be checked against a cardholder’s own history and connected transaction context. Within high-score alerts, Qamar reports a 93.4% fraud rate when the device was already familiar to the account, compared with 12.3% when the device was marked new. He also reports a 23% fraud rate for a purchase far above the customer’s median when nothing else had changed.

Those figures are associations in the project’s high-score alert sample, not universal rules. The author’s interpretation was that an unusual purchase amount alone was less informative than velocity relative to that card’s own rhythm, or concurrent activity in the cardholder’s home region. In practice, a potentially useful finding is not simply “large purchase”; it is a specific, time-bounded departure from a person’s usual pattern that an analyst can verify.

How did the behavior-based model perform?

Qamar reports fitting log-odds weights from investigations opened before October 2016, then evaluating the model on 278 October alerts of the same kind. The reported holdout results were ROC-AUC 0.849, accuracy 0.791, and Brier score 0.157. The author says the model used eleven named findings intended to be individually checkable, with the interface showing the arithmetic behind each probability.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These results apply to that 278-alert holdout and its sampling frame. They do not establish performance on all transactions, at other banks, or in later periods. The comparison also illustrates why a model should not be judged on a single number:

Approach What the article reports What the result can show
Bank detection score ROC-AUC 0.053 across 5,565 closed investigations, as reported by Qamar in 2026 How the score ranked outcomes within those selected closed cases; not its performance on the general transaction population.
Naive score inversion 93% on a balanced October holdout, as reported by Qamar in 2026 A result on that benchmark’s construction, which Qamar says does not demonstrate broader fraud detection.
Prior hand-tuned heuristic 0 out of 40 on the same holdout; it abstained on 31 cases, as reported by Qamar in 2026 Both its poor result on those cases and its limited coverage matter; abstentions cannot be treated as successful classifications.
Behavior-based log-odds model ROC-AUC 0.849, accuracy 0.791, and Brier score 0.157 on 278 October alerts, as reported by Qamar in 2026 Ranking, classification, and probability error on that alert sample; calibration for a wider population remains unestablished.

ROC-AUC describes ranking, accuracy depends on a decision threshold and class mix, and Brier score evaluates probability error. None alone establishes whether the probability estimates are calibrated for the cases where a bank will use them. A fair comparison also asks how many cases each method can assess, what false positives and missed fraud cost, and whether analysts can reproduce the stated evidence.

What did TigerGraph contribute?

TigerGraph served as evidence storage and case memory, rather than as proof that a prediction was correct. The project represented customers, cards, transactions, device profiles, billing regions, email domains, and closed cases as connected entities. This let an investigator examine relationships around a transaction and retain the reasoning attached to earlier cases.

Time boundaries were central to the design. Queries were cutoff-bounded so an investigation could not use information recorded later—including cases that closed after the investigation date. The agent re-derived claims with GSQL and compared aggregate results, sampled transaction fields, and the flagged transaction itself. The exporter blocked cases that failed those parity checks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The project also wrote investigations back into the graph, linked to their findings, transactions, implicated cards, device profiles, and cited prior cases. Its vector corpus contained 503 documents: 37 policy chunks and 466 similar analyst notes. Qamar says the two kinds of material were ranked separately because near-identical case notes could otherwise bury the policy guidance. Separating policy from precedent helps make a retrieved example informative without letting repeated historical patterns silently override a rule.

How did the investigator handle uncertainty?

The described policy called for verification before blocking when there was a single weak signal below 0.70. Rather than present an uncertain recommendation as settled, the agent recorded its initial recommendation, requested evidence, simulated a cardholder response, documented that assumption, and then revised its assessment while preserving both recommendations and the reason for the change.

The author describes two situations where the ordinary verification path did not fit:

  • Customer-reported transaction: A customer report already constitutes a denial, so the agent does not ask that cardholder to validate the reported transaction.
  • Shared-origin cluster: Activity connected across several customers cannot be resolved by asking only one cardholder; the described response is reporting and monitoring connected cards.

These are policy behaviors in the described project, not a general substitute for a bank’s own escalation and regulatory procedures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What do the twenty benchmark cases establish—and what do they not?

Across twenty benchmark cases, Qamar reports eleven fraud assessments, six legitimate assessments, and three uncertain assessments. The agent made seven evidence requests, changed four recommendations, and produced two suspicious activity reports. These counts describe the project’s small benchmark, not a broad estimate of how often such actions would occur in operational investigations.

The author notes that four of the twenty cases did not match the five documented typologies. The stated next challenges were to evaluate cases beyond those supplied, account for undocumented patterns, and connect thresholds to a real cost model. No independent replication, broad population study, or evidence of deployment outcomes is provided in the project account.

What should a team take from this case?

  • Audit the label and sampling process. Establish who received an alert, why it was investigated, and how an outcome became a label before interpreting a metric.
  • Match evaluation data to intended use. A balanced or trigger-selected benchmark can answer a narrower question than the full transaction population.
  • Keep evidence reproducible and time-bounded. Prevent future information from leaking into a case and verify graph-derived findings against the underlying records.
  • Separate risk ranking from action. Probability quality, coverage, false-positive and false-negative costs, and policy thresholds all matter to a blocking decision.
  • Make uncertainty visible. Preserve initial and revised recommendations, the evidence requested, and the assumptions that changed the conclusion.

As Qamar puts it in the DEV Community article, “An interface that renders beautifully and passes every schema check tells you nothing about whether the investigation is any good.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.