Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

SentinelGraph treats a fraud risk score as a reason to investigate, not a verdict. In a prototype described by its builder, Nirmal Joseph Ukken, the system combines TigerGraph graph queries, prior-case memory and policy rules, then either reaches a threshold-supported conclusion, asks for more evidence or routes an action to a human. Its reported benchmark is promising, but it does not establish performance in a live banking service.

What SentinelGraph is designed to do

Ukken described SentinelGraph in a September 24, 2026 post associated with Task 4 of the TigerGraph problem statement for Hacker House Goa 2026. The project frames an alert as an investigation trigger: “A risk score is a reason to look. Never a verdict.” The aim is an auditable recommendation and next step, rather than an unconstrained language-model decision. Read Ukken’s project account on DEV Community.

The described evidence graph links customers, cards, transactions, devices, email domains and billing regions. Other graph layers hold active cases, closed cases and policy knowledge. This lets the system examine relationships among entities and retrieve relevant history alongside a transaction’s risk signals.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the investigation loop works

  1. Open a case. The system creates a case and audit trail for the alert.
  2. Query connected evidence. TigerGraph queries can look for transactions on other cards sharing a device or email within a time window, build cardholder behavior profiles, and identify connected card components.
  3. Retrieve relevant memory. The system searches earlier cases by graph adjacency and uses TigerGraph native vector search to retrieve case and policy information.
  4. Assess evidence without double-counting. The design combines evidence while limiting the influence of correlated signals, so multiple indicators that reflect the same underlying fact are not automatically treated as independent confirmation.
  5. Stop, or ask for more. It applies an explicit probability threshold and requests potentially clarifying evidence when the evidence is not decisive.
  6. Act or route the case. Depending on policy, an action may execute automatically or wait for human approval; the case and outcome are written back to graph memory.

The author reports 16 installed GSQL queries exposed through TigerGraph MCP. These are implementation details reported by the project builder, not a TigerGraph product specification.

When the agent stops—and when it should not

The stated stop rule requires two independent evidence families to agree and the posterior probability to reach at least 0.85 or at most 0.15. The upper threshold supports a high-confidence fraud conclusion; the lower threshold supports a low-confidence one. If neither threshold is met with the required agreement, the system is meant to seek evidence that could resolve the uncertainty rather than force a binary answer.

Examples of requested evidence include step-up authentication or customer verification. The project author says some actions are routed to human reviewers: auto actions may execute, while L1 or L2 actions wait for human approval. “Uncertain” is an honest answer, Ukken says. The described thresholds and routing are features of this prototype’s policy, not general banking or regulatory standards.

What the language model does—and does not do

In the described architecture, deterministic evidence calculations and policy rules govern the decision logic. The language model is used for bounded additional tool calls and for drafting narrative or suspicious activity report (SAR) text, with output validation. That separation is intended to make the recommendation reviewable: the model can help gather or explain information without serving as the sole authority on whether a case is fraud.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the reported benchmark shows

Ukken says the project used IEEE-CIS card data with the fraud label removed, historical closed investigations, a policy, documented patterns and a set of 20 benchmark alerts. The post reports these figures:

Rank #3
Graphic Image Sports Illustrated Tiger Woods 25 Year Special Edition Leather Book
  • Commemorate Tiger Woods' 25-year journey with a billiant, fully illustrated table book from Sports Illustrated
  • Sturdy build and construction. The hand bounded green leather hardcover gives it the perfect vintage look and durability
  • Its polished aesthetic perfectly aligns with the golf theme of this book, lending an elegant touch to your bookshelf or coffee table.
  • 232 pages full of iconic vibrant photos and some of the best written coverage of Woods’s career
  • Beautiful Stories, a good read, and great photographies, the ideal gift book for any Tiger fan
Reported measure Project-reported result
Transactions 590,742
Closed cases 5,565
Installed GSQL queries 16
Largest detected ring 28 cards
Memory model AUC 0.914
Bank score AUC 0.866
Benchmark alerts 20: 10 legitimate, 9 fraud and 1 uncertain
SARs in benchmark 6

All values in the table are reported by Ukken in the September 24, 2026 post; the source is the project builder’s account, not an independent evaluation. The post says the memory model was trained on July through September data and tested on October. It also reports nearly double the bank score’s average precision, but gives no exact average-precision values.

The author says rerunning the same 20 alerts produced the same decisions. That supports repeatability on this particular benchmark as reported; it does not establish independent replication, accuracy on other populations, live-bank performance or operational savings. An AUC comparison alone also does not show how the system would perform at a bank’s chosen review thresholds or under its actual costs for missed fraud and false alarms.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What remains unproven or unfinished

The benchmark simulated evidence replies because real reply channels were not implemented. The post identifies real SMS or app replies, streaming ingestion, likelihood ratios learned from resolved agent cases rather than set by hand, and external enrichment as future improvements. The described work therefore does not demonstrate a fully live, end-to-end banking service.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Evidence traceability: the graph and case trail are intended to connect recommendations to evidence and prior cases.
  • Correlated signals: the design explicitly aims to limit treating related indicators as separate confirmations.
  • Uncertainty handling: it specifies stop thresholds and a path to request more evidence instead of always issuing a verdict.
  • Human oversight: the described policy routes some actions for L1 or L2 approval.
  • Deployment and validation: the post reports a project benchmark, not a live-service evaluation or a comparison against other vendors.

Those distinctions matter when assessing agentic fraud systems: a compelling architecture and a repeatable small benchmark are useful evidence about a prototype, but are not substitutes for validation on representative live workloads, reliable integrations and independently reviewed outcomes.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.