Skip to main content
Back to timeline
arXivSource publication:

Hermes dual-scale graph transformer cuts customer friction 44.44% and lifts fraud recall 24.66% on 130M+ high-risk sessions

Related research and updates

Synopsis

The work presents Hermes, a dual-scale architecture in which Micro-GT applies relation-aware attention over a temporally safe heterogeneous session neighborhood and Macro-GT attends over non-anticipative ecosystem climate tokens plus adaptive class prototypes; on more than 130 million high-risk transaction sessions it achieves a 44.44% relative reduction in customer friction (FR@50R) and a 24.66% relative improvement in fraud recall (R@5FR) over the production XGBoost system, with 85.12% overall ROC-AUC.

Interpretation

Hermes unifies local relational structure and global ecosystem context within a single session-level fraud prediction framework, fusing Micro-GT and Macro-GT representations through a learned projection and end-to-end training. Prior production tabular models score sessions in isolation and graph models mainly reason over local neighborhoods, lacking a unified mechanism that expresses both local relations and the evolving ecosystem state. Evaluated on a proprietary dataset of more than 130 million high-risk sessions with a chronological split (final four months test, preceding two months validation), with all models obeying serve-time and label-maturity constraints.

Micro-GT applies four structured attention masks in sequence (identity-self, relation-group, chain, full-context) to preserve account, device, and network relation semantics and model multi-hop relational chains. Relative to homogeneous SAGE-style aggregation, it introduces node heterogeneity and direct all-pair interactions, capturing higher-order dependencies beyond conventional message passing. Ablation shows Micro-GT (full attention) raises overall AUC from ATLAS's 84.02% to 84.58% and P2 AUC from 78.75% to 80.47%; adding relational masks brings overall AUC to 84.78% and P1 AUC to 83.10%.

Macro-GT attends over non-anticipative daily climate tokens (global fraud activity, platform behavior, cross-platform shifts, velocity and anomaly, ecosystem topology, risk regime) together with adaptive class prototypes, supplying ecosystem-level information beyond the local neighborhood. It lets the model sense ecosystem-level state changes such as attack volume, platform concentration, and infrastructure reuse rather than relying only on local graph patterns. Adding Macro-GT raises overall AUC to 84.93% and P2 AUC to 81.06%, lowers FR@50R to 5.519%, and raises R@5FR to 48.44%; class prototypes further push overall AUC to 85.12%, FR@50R to 5.373%, and R@5FR to 48.91%.

Temporal-stability analysis across four consecutive held-out test months shows Hermes consistently outperforms baselines, with no degradation relative to graph-based models in later months. It indicates that combining local relational modeling with global ecosystem context remains robust across periods of changing fraud activity. The test set is chronologically split and evaluated month by month, reporting that Hermes leads throughout the evaluation window and that gains remain stable across changing fraud activity.

Perspective

The result targets digital-banking high-risk session settings that have a heterogeneous session graph of shared account, device, and network identifiers and delayed adjudicated fraud labels; the model runs non-anticipatively under serve-time and label-maturity constraints, suited to deployment environments that must jointly control customer friction and improve fraud recall, and its dual-scale design offers a reusable paradigm for relational risk modeling that stays temporally stable amid an evolving fraud ecosystem.

Climate-token and class-prototype feature definitions are spread across the appendix, and the main text does not give all numerical details; data distributions, identifier coverage, and label delays at other institutions may affect transfer performance; monthly stability is presented as a figure without per-month numbers in the main text; and hyperparameters such as prototype counts, EMA momentum, and patience thresholds are not given in the main text, so reproduction requires the appendix and implementation details.

Sources