SERA-IDS uses structured experience retrieval to lift 7B–14B small models from as low as 1.94% to 87.66% macro F1 on NetFlow intrusion detection
Synopsis
SERA-IDS converts training-time classification errors into structured experience rules carrying feature conditions, target classes, confidence, and provenance, admits them through a confidence gate into a frozen Experience Library, and at inference retrieves up to five class-diverse rules whose conditions are matched against the query flow as evidence for a locally deployable small language model; on NF-BoT-IoT and NF-ToN-IoT, macro F1 for Llama 3.1, Phi-4:14B, and Qwen2.5:7B rises from 1.94%–11.66% zero-shot to 69.50%–87.66% without fine-tuning or paid hosted inference.
Fig. 1: Overview of the proposed SERA-IDS framework. (a) Offline rule construction from misclassified training flows, including error analysis, confidence-based admission, and structured rule storage. (b) Online classification using feature representation, diversity-aware top-five retrieval, condition verification, and SLM-based decision reasoning.
arXivInterpretation
An error-driven, confidence-gated structured experience-rule framework: an Error Analysis Agent generates rules only for misclassified flows, each rule containing a behavioral description, numerical and categorical conditions, class-confusion information, tool-derived confidence, and provenance; confidence is computed from the average strength and consistency of signals such as an abuse-reputation-style score, a malicious/unknown/benign assessment, a threat-pulse-style activity score, and similarity to reference attack-technique descriptions, with candidates above threshold entering the active library and the rest going to an administrative queue. Compared with free-form textual experience, the rules preserve measurable feature conditions and confused-class information, making experience a hybrid object that can be retrieved and condition-verified rather than a prose explanation matched only by semantic similarity. At the end of construction the active libraries held 7,326 rules for NF-BoT-IoT and 10,664 for NF-ToN-IoT, with 56 and 3,043 lower-confidence candidates in the administrative queue; rule-generation rates generally declined at later stages, suggesting growing coverage of recurring errors.
A diversity-aware top-five retrieval and condition-verification mechanism: a FAISS candidate pool is retrieved, the highest-similarity rule per class is kept and the classes with the highest best-rule similarity are selected, and the query flow is compared against the retrieved numerical and categorical conditions to produce feature-level evidence passed to the decision SLM alongside the flow and rules. The retrieval unit moves from a single free-form experience to five structured experiences with explicit conditions, and condition evidence participates as contextual support rather than a deterministic vote. Holding the decision model fixed, SERA-IDS adds 16.86, 39.70, and 20.11 points over MA-IDS on NF-BoT-IoT and 19.86, 72.36, and 4.77 points on NF-ToN-IoT; macro F1 rises monotonically from zero-shot through MA-IDS to SERA-IDS for every model on both datasets.
The Experience Library transfers across models: it is built by Llama 3.1, then frozen and reused by Llama 3.1, Phi-4:14B, and Qwen2.5:7B with no fine-tuning, no parameter updates, and temperature zero. This indicates the benefit of structured external memory does not depend on the model that built it; Phi-4 reaches 87.66% macro F1 on NF-ToN-IoT, above the library-building Llama 3.1 and above the earlier GPT-4o result of 85.22% (though the present evaluation uses ten classes versus nine previously). Macro F1 for the three models is 86.46%, 69.50%, and 83.51% on NF-BoT-IoT and 84.16%, 87.66%, and 83.37% on NF-ToN-IoT; under SERA-IDS macro recall stays between 94.78% and 97.14% while macro precision ranges from 54.36% to 79.87%.
A characterization of the precision–recall and latency–performance trade-offs: relative to MA-IDS, SERA-IDS lowers macro precision for Llama 3.1 and Qwen2.5 by 1.12 to 6.62 points but raises their macro recall by 17.22 to 34.87 points; on latency, Llama 3.1 and Qwen2.5 stay below 0.4 s/flow while Phi-4 requires 3.05 to 3.45 s/flow. It quantifies both the gains and the costs of structured experience, noting the recall gain may come from five class-diverse rules exposing the model to more candidate classes, while the study does not isolate retrieval width from rule representation. Latency is measured as end-to-end local inference rather than line-rate IDS throughput; Phi-4 achieves the best macro F1 on NF-ToN-IoT but the lowest on NF-BoT-IoT, so its higher latency does not consistently buy better detection.
Perspective
The result targets multiclass intrusion detection over NetFlow flow records with a fixed class set, suited to a deployment where a lightweight detector pre-filters flows and an SLM performs analytical adjudication rather than line-rate packet-level detection. The library is frozen before testing and test labels are used only for metric computation, so the method addresses known attack classes within the training distribution; construction and inference share the same representation of 14 numerical features plus categorical and protocol information, and rule conditions are based on that representation. Because the library built by Llama 3.1 is reused by Llama 3.1, Phi-4:14B, and Qwen2.5:7B, structured rules can serve as cross-model external memory for security teams that want local execution, controlled data handling, and predictable cost.
A careful reader would still watch several open questions: the comparison jointly varies rule representation and retrieval width, so the separate contributions of structured conditions and top-five retrieval cannot be read off the present results; the confidence signals come from an offline synthetic testbed rather than a live threat-intelligence feed, leaving external validity to be examined; latency is end-to-end local inference and does not cover line-rate throughput; and the stated future work on component ablations, unseen attacks, and library robustness has not yet been carried out. In addition, some equations and configuration values are not fully rendered in the text, so reproducing exact thresholds and retrieval parameters would require consulting the original figures and tables.
