Skip to main content
Back to timeline
arXivSource publication:

Sigma-Hunter fine-tunes a 7B model on 3,635 Sigma rules, scoring 8.17 overall against 7.88 for the strongest general-purpose baseline

Synopsis

The work builds Sigma-Hunter by expanding 3,635 validated open-source Sigma rules into 7,663 question-answer and analyst-reasoning instruction examples, partitioned at the source-rule level to prevent leakage; after LoRA fine-tuning of a 7B Mistral and a Phi-4 model, evaluation on 188 held-out rule-generation examples across syntax, approximate field consistency, and semantic judging gives Sigma-Hunter-Mistral an overall 8.17, above the strongest general-purpose baseline at 7.88 and untuned Mistral at 4.61.

Source-provided article image: Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering
Fig. 2 ·

Fig. 2 : Different initial learning rates are evaluated for the Sigma-Hunter-Mistral model before the fine-tuning stage.

arXiv

Interpretation

The paper presents a reproducible pipeline that programmatically converts validated Sigma rules into instruction examples for rule generation, rule explanation, ATT&CK mapping, false-positive analysis, threat-hunting guidance, log-source reasoning, and rule refinement, using source-rule-level partitioning so that all examples derived from one rule stay within a single split. Relative to example-level random splitting, source-rule-level partitioning prevents different tasks derived from the same rule from appearing in both training and evaluation and inflating performance; the paper notes this prevents exact rule leakage while closely related rule variants may remain across partitions, leaving similarity-aware splitting to future work. Dataset counts are explicit: 3,635 source rules (433 emerging-threat, 130 threat-hunting, 3,072 generic detection) expanded into 7,663 examples, partitioned 80/10/10 into 6,130 training, 766 validation, and 767 test examples; the corpus skews toward generic detection rules at 84.5% of source rules.

The paper trains two domain-adapted models, Sigma-Hunter-Mistral (7B) and Sigma-Hunter-Phi4, using LoRA parameter-efficient fine-tuning and exports them as GGUF/Ollama packages for local inference. Unlike reliance on hosted model services, the capability is packaged for operator-controlled environments, targeting disconnected or air-gapped settings where analysts cannot reach external model services; training two model families tests whether domain adaptation helps across architectures rather than only one. Fine-tuning configuration is specific: base model mistral-7b-instruct-v0.3 in 4-bit quantization with roughly 7.29B total parameters, LoRA adapters across all 32 transformer blocks targeting attention and MLP projections, 41,943,040 trainable parameters (0.575% of the model), 2,048-token sequence length, three epochs, per-device batch size 4, gradient accumulation 4, 8-bit AdamW with cosine schedule, and checkpoint selection by lowest validation loss.

The paper adapts multi-dimensional structured-output evaluation to Sigma generation, measuring syntax validity, approximate field consistency, and model-judged detection logic, completeness, rule selectivity, and log-source match, and reports that fine-tuning improves semantic quality rather than only formatting. The paper shows syntax validity is a weak proxy for semantic rule quality: several baselines emit structurally valid rules with near-ceiling log-source match yet differ materially in detection logic, completeness, and selectivity; matched-model comparisons show Sigma-Hunter-Mistral rising from 4.61 to 8.17 and Sigma-Hunter-Phi4 from 6.82 to 8.03. The quantitative benchmark contains 188 held-out rule-generation examples with identical prompts across models; semantic scores are averaged over successfully judged outputs, with valid-judgment counts varying by model. Table V illustrates the gap: Sigma-Hunter-Mistral preserves both the curl.exe executable constraint and the file:/// command-line pattern, whereas gpt-oss:20b retains only the protocol-handler condition; both are syntactically valid with field-consistency scores of 10, but the omission reduces the latter's selectivity and overall scores to 4 and 5.

The paper positions Sigma-Hunter as an analyst-in-the-loop capability rather than an autonomous deployment system: the model returns a candidate rule with a short explanation, likely log-source requirements, and false-positive considerations, and the analyst checks it against local telemetry and SIEM field mappings, runs it on historical data, tunes filters or allowlists, and only then promotes it through detection-as-code review. This framing keeps generated output inside a human accountability chain, emphasizing a shorter path from threat description to a reviewable detection artifact rather than replacing the detection engineer's judgment. The paper states the intended workflow remains analyst-directed and notes that model-based judging is not independent expert validation, with human detection-engineer validation left as future work.

Perspective

The result is aimed at detection engineers and defensive cyber operations teams working in disconnected or air-gapped environments, and applies to guided generation where the analyst supplies a threat description plus partial detection cues (condition expression, selection names, field names with modifiers, representative field values, and the intended log-source product, service, and category), rather than unconstrained generation from threat reports alone. The capability is packaged for local inference via GGUF/Ollama so prompts, threat reports, logs, and detection artifacts stay within the operator-controlled environment; the intended workflow is analyst-in-the-loop, checking against local telemetry and SIEM field mappings, running on historical data, tuning filters or allowlists, and promoting through detection-as-code review. Next steps include human analyst validation of generated rules and extension to backend-specific detection languages and operational schemas.

Semantic scores come from gpt-oss:20b as a judge model, which the paper describes as comparative semantic assessment rather than independent expert validation, with human detection-engineer validation left as future work; rule selectivity is a semantic judgment, not empirical alert selectivity measured on labeled telemetry; field consistency uses a curated telemetry-category dictionary and is an approximate signal rather than an authoritative SIEM schema validator; log-source match approaches the ceiling for most models and provides little separation; the corpus skews toward generic detection rules (84.5% of source rules), which the paper notes may limit relative gains on emerging-threat and threat-hunting tasks; source-rule splitting prevents exact rule leakage but closely related rule variants may remain across partitions; and semantic metrics are averaged over successfully judged outputs, with valid-judgment counts varying by model.

Sources