ArGuard Shared Task: 58 teams registered and 35 competed on Arabic meme and LLM-prompt harm detection, with best systems reaching macro-F1 of 0.823/0.419/0.984/0.790
Synopsis
ArGuard is a shared task on harmful content detection in Arabic, with two tracks: Track A on multimodal hate detection in Arabic memes and Track B on harmful prompt detection for Arabic LLM safety evaluation; 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers, with teams exploring models such as AraBERT, Jais, and Qwen3-VL, and the best systems achieving macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, where fine-grained meme classification in A2 was the most challenging setting due to sparse labels and train-test distribution shifts.
Figure 1: Overview of the ArGuard tracks and subtasks. Top: Track A covers binary (A1) and fine-grained (A2) Arabic hateful meme detection. Bottom: Track B covers binary safety (B1) and fine-grained harm-category (B2) classification of LLM prompts.
arXivInterpretation
The work proposes and organizes the ArGuard shared task, splitting Arabic harmful content detection into two tracks: Track A for multimodal hate detection in Arabic memes and Track B for harmful prompt detection in Arabic LLM safety evaluation. Relative to single-modality or single-task evaluation setups, this covers both image-text memes and LLM prompts, and explicitly separates multimodal hate detection from LLM safety evaluation. Based on the task design described in the abstract: two tracks and four sub-settings (A1, A2, B1, B2), with 58 registered teams, 35 final-evaluation participants, and 27 system-description papers.
The best systems achieved macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, showing clear differences in difficulty across sub-settings. These numbers provide a current performance reference for each sub-setting of Arabic harmful content detection, with B1 near ceiling and A2 markedly lower. Taken from the best-system macro-F1 values reported in the abstract as final shared-task evaluation results.
Fine-grained meme classification in A2 is identified as the most challenging setting, which the abstract attributes partly to sparse labels and train-test distribution shifts. This shifts the explanation of difficulty from model capability toward data-level sparsity and distribution differences, giving concrete direction for future data construction and evaluation design. The abstract states that A2 was the most challenging setting and cites 'sparse labels and train-test distribution shifts' as a partial cause; no further quantitative breakdown is provided.
Participating teams explored models such as AraBERT, Jais, and Qwen3-VL, spanning Arabic pretrained language models and multimodal models. This indicates that both Arabic-text-oriented and multimodal model approaches were tried on the task, forming a comparable methodological picture. The abstract lists the model names explored by participating teams, without giving specific configurations or per-team scores.
Perspective
The work targets the specific setting of Arabic harmful content detection and serves two kinds of users: teams building Arabic content moderation or LLM safety filtering systems, who can use the best macro-F1 on A1, A2, B1, and B2 as a performance reference; and researchers studying multimodal hate detection and prompt safety evaluation, who can learn from the directions explored with models such as AraBERT, Jais, and Qwen3-VL. The results apply to the shared-task evaluation conditions; the abstract does not state whether they transfer directly to other languages, other platforms, or real production traffic.
The abstract does not give the data size, annotation process, or label taxonomy details for each sub-setting, nor the extent of label sparsity and distribution shift in A2, making it hard to judge how much of the 0.419 result is driven by data factors. It also does not list per-team scores or detailed model configurations, so the relative contribution of different approaches cannot be compared. In addition, the abstract does not state whether the data and system-description papers are publicly available, so readers planning to reproduce or reuse the work would need to consult the original text.
