CyTReX turns DNP3 anomaly alerts into traceable ranked threat hypotheses, with a five-configuration comparison showing the complete pipeline gives the richest evidence but weaker label anchoring
Synopsis
The paper presents CyTReX, an explainable threat reasoning framework for distributed energy resource (DER) networks that consolidates edge anomaly detection evidence, cloud attack interpretation, SHAP feature attributions, surrogate decision rules, and MCP-enabled cyber threat intelligence into a unified evidence packet that constrains large language model reasoning to produce ranked threat hypotheses and attack trees; across five configurations on a DNP3 case study, adding reasoning components improved hypothesis specificity, evidence traceability, and analytical grounding, with the complete pipeline providing the richest evidence-grounded reasoning context.
Fig. 1: Representative DER-integrated grid communication architecture and attack surfaces [ 9 ] .
arXivInterpretation
An edge-cloud reasoning workflow that transforms DER network anomalies into ranked threat hypotheses to support Security Operations Center (SOC) analyst triage. Unlike anomaly detection that emits a single class label, the workflow treats detection output as intermediate evidence rather than a final decision and explicitly surfaces secondary hypotheses and evidence gaps. Evaluated on a DNP3 communication dataset as the case-study environment, with the same representative edge alert batch used across all five configurations so differences are attributable to the active evidence components.
A unified evidence packet that consolidates edge anomaly evidence, cloud attack interpretation, SHAP network-feature attributions, surrogate decision rules, evidence-quality signals, and provenance metadata as structured LLM input. Unlike feature-level explanation alone or free-form LLM inference, the packet requires every ranked hypothesis and attack-tree branch to be traceable to explicit evidence and communicates incomplete or conflicting evidence rather than suppressing it. The model explanation layer combines SHAP attributions with surrogate decision rules and computes a Rule-SHAP consistency score and a surrogate rule-quality score to generate evidence-quality signals; CTI is retrieved behavior-oriented through an MCP interface from MITRE ATT&CK for ICS and CAPEC with provenance recorded.
A five-configuration DNP3 case study evaluating each reasoning component's contribution to hypothesis specificity, evidence traceability, and unsupported-claim reduction. Unlike an aggregate evaluation, the design isolates the effects of edge evidence, CTI, cloud attack interpretation, model explanation, and the complete pipeline. The edge layer selected an Isolation Forest plus nearest-neighbor ensemble (accuracy 98.43, F1 99.21) and the cloud layer selected LightGBM (accuracy 99.58, F1 99.58); REPLAY appeared as H1 under all five configurations, C3 gave the strongest exact label anchoring, and C5 provided richer semantic interpretation while shifting protocol-specific labels toward broader threat effects.
Out-of-distribution (OOD) testing on held-out DNP3_ENUM and DNP3_INFO reconnaissance-like behaviors to examine whether the framework expresses unseen behavior as uncertain rather than forcing it into known classes. Unlike a closed-set classifier that can only emit known classes, the test explicitly examines reasoning behavior outside the trained attack-label space. C2 produced broad hypotheses such as parameter manipulation, data collection, initial access, and lateral movement from edge evidence and CTI; C5 enriched the output with cloud interpretation, SHAP attribution, and surrogate rule evidence, but the closed-set classifier could also pull unseen behavior toward known classes such as replay.
Perspective
The framework targets security operations for DER-integrated grids, converting a flagged network alert into a structured investigation object so analysts can see the most supported threat interpretation, related secondary possibilities, relevant CTI and mitigation context, and remaining evidence gaps before deciding whether to escalate, investigate further, or approve a response action; the output is intended to support analyst interpretation and decision readiness, not automated response execution. Evaluation runs in an emulated DER network environment with three virtual DER devices generating DNP3 traffic in real time, a semi-supervised anomaly detector at the edge, a supervised classifier in the cloud, and llama3.1:8b used uniformly for all attack-tree generation, so output differences across configurations are attributed to the active evidence components. CTI retrieval is behavior-oriented through an MCP interface, filters generic domain vocabulary such as DNP3, ICS, and SCADA, and records source references and retrieval status.
A careful reader would still watch: the five-configuration comparison rests on a single representative edge alert batch and one DNP3 case-study environment, so generalization to larger datasets and broader DER communication scenarios remains to be evaluated; C5 is weaker than C3 under exact label matching, indicating that combining CTI and explanation evidence shifts protocol-specific labels toward broader threat effects, so the trade-off between label anchoring and semantic richness needs to be weighed against operational needs; the OOD cases show the closed-set classifier can pull unseen reconnaissance behavior toward known classes, leaving explicit open-set detection an open problem; latency is not monotonic across C1-C5, reflecting the combined cost of LLM generation, CTI retrieval, evidence-packet size, and attack-tree complexity rather than a direct function of the number of active components; and the weights for Rule-SHAP consistency, surrogate rule quality, and verified confidence are design parameters whose alternatives still need study for their effect on hypothesis ranking and evidence-quality labels.
