Skip to main content
Back to timeline
arXivSource publication:

MARS compares direct classification with claim scoring on 1,195 PE/ELF binaries: direct classification is more accurate across all six models, while retained claims enable policy replay

Synopsis

MARS is a malware triage framework for PE and ELF binaries that, with a shared evidence collector and identical static evidence bundles, compares direct LLM classification against structured behavioral claims scored by a fixed policy; across 1,195 binaries, 1,001 near-duplicate clusters, and six models, direct classification is more accurate for all six models with paired advantages of 3.7 to 20.9 percentage points, while the claim path retains the policy inputs and permits replay and policy revision without another model call.

Source-provided article image: MARS: Malware Analysis with Rule-Based Scoring of LLM Claims
Figure 1 ·

Figure 1: MARS evaluation architecture. The full-corpus comparison uses a shared static evidence bundle for rules-only classification, direct LLM classification, and single-pass claim scoring. The two LLM paths are evaluated across six models. Additional evidence and support assessment are examined in subset studies.

arXiv

Interpretation

Under matched evidence collector, evidence bundles, and models, direct classification achieves higher overall accuracy for all six models and higher malicious alert recall in ten of twelve platform and model combinations. Prior comparisons of malware analysis systems were confounded by differences in evidence collectors, models, and inputs; holding the evidence bundle and model fixed restricts the comparison to the complete decision paths, including their prompts and output schemas. 1,195 binaries, 1,001 near-duplicate clusters, six models; paired accuracy advantage of 3.7 to 20.9 percentage points with all six cluster-bootstrap 95% intervals above zero; the ordering persists when errors remain in the denominator and in architecture strata, a metadata-hard subset, and a diagnostic that selects the claim threshold on evaluation labels.

Claim mediation provides no consistent reduction in performance variation across models, and repeated execution shows that a fixed scoring policy does not ensure end-to-end stability. The result separates the reproducibility of a stored policy calculation from the stability of the complete pipeline: a fixed scoring function is reproducible only for a fixed claim set, since claim generation and severity assignment vary across calls. Across ten runs, exact verdict consistency ranges from 36.7% to 93.1%, with direct classification having the highest alert consistency in every platform and model combination; exact consistency is lower for malicious samples in 17 of 18 model, platform, and path combinations; on PE, changing samples sit closer to the benign threshold (e.g., 0.28 versus 3.40).

Retained claims make verdict computation inspectable and revisable: recomputing the policy from archived records reproduces all 9,456 recorded labels, including 2,366 from two models withdrawn by their provider. A stored direct label can only be retrieved unchanged, whereas the claim path additionally supports recomputing a verdict from the same retained inputs under a different scoring policy without renewed model access. All 9,456 records reproduced; applying the evaluated policy to those records agrees with the deployed default on 4,682 labels (49.5%), a disagreement reflecting a policy change rather than a replay failure.

Under predefined evidence perturbations, removing indicator fields reduces malicious alert recall by 29.2 points for the claim path versus 10.0 for direct classification; in a family-identification probe, claims beat verdict labels but score below evidence text. The perturbation study quantifies how sensitivity to evidence modification differs across decision architectures, and shows that the structured interface retains policy inputs rather than stronger family-discriminative information. Perturbation subset of 182 files, 728 bundles, 4,368 calls, three models; under S1 recall falls 34.1 points for rules only, 29.2 for single-pass, and 10.0 for direct classification, with the larger single-pass loss in all four architecture strata; in the family probe, evidence text reaches 61.1% (PE) and 83.1% (ELF) point accuracy, above every claim representation.

Perspective

The work targets malware triage settings that require a sample-level benign/malicious verdict, covering PE and ELF binaries, basic static evidence (optionally decompiler and runtime evidence), eleven claim categories, and an additive scoring policy. For operational settings that need explicit, revisable rules over retained model outputs, the claim path offers policy inspection and revision: category weights and thresholds can be changed independently of claim generation, and archived records can be replayed and rescored after the generating model becomes unavailable. The family-identification probe and the dynamic-evidence follow-up are auxiliary analyses for a particular classifier and 20 selected cases, respectively.

The corpus is not family- or time-balanced, its class proportions do not represent deployment prevalence, family annotations were not independently verified, and difficult benign categories such as potentially unwanted programs and administrative tools remain incompletely covered; strong shallow-metadata baselines and unmasked identity cues limit attribution of absolute verdict accuracy to behavioral interpretation. Matching evidence and models does not isolate representation format from other path differences such as prompts, response schemas, output budgets, parsing, and threshold calibration, and only the claim path uses a development-calibrated threshold. A collector defect causes 11 malicious samples to fail in both LLM paths, with evidence bundles up to 25.2 MB. The perturbation study covers four fixed modifications to stored bundles and does not establish that these modifications can be realized through binary changes or characterize adaptive manipulation. Provider withdrawal prevents renewed claim extraction for two models, and floating model tags limit exact re-execution. Claims were not expert-validated for semantic accuracy, and effects on analyst decisions were not measured.

Sources