Humanities & Social Sciences
89 items
PersonaPath: A Knowledge-Centric Benchmark for Personalized Learning Path Planning
This work introduces Knowledge-Centric (KC) personalized learning path planning and builds PersonaPath, a benchmark pairing 2,000 fine-grained learner personas with a hierarchical knowledge graph of 347 textbooks, 1,751 units, and 4,092 concepts across 77 subjects, to evaluate whether large language models can decide which textbook, unit, and concept a learner should study next given learner profiles, mastery states, and prerequisite knowledge structures; evaluation of representative LLMs shows the strongest model reaches only a 29.5% final pass rate in Basic Education, with adaptivity as the main bottleneck, where no model exceeds 44.7% in tailoring paths to individual learners.
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across Gemma-3 and Qwen3.5 on six harmful-content benchmarks plus Spanish and Hindi-English code-mixed evaluations, this work separates failures of representation from failures of routing, finding that sparse readouts outperform native prediction on all six binary tasks (Qwen 0.740 vs 0.432 native macro-F1; Gemma 0.532 to 0.714), that probe-discriminative and output-routed directions dissociate, that calibration-only routing recovers 93.
TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
The work presents TeleAntiFraud 2.0, a refreshable Chinese call-audio benchmark organized as monthly frozen snapshots, built with a Mixed-Tree Anti-Fraud Generation Pipeline that turns online fraud-case abstracts into profile-grounded scenarios and expands them into fraud and near-domain lawful sibling dialogues sharing context and diverging only at label-bearing actions, rendered as role-matched speech; each frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), controlled text experiments show three classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, and full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity.
A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages
The work presents a multi-step framework that first corrects word segmentation with ChatGPT-4o, then fills missing entities via context-free lexicons with a minimum frequency of three, and finally applies frequency-based iterative self-training with a dual threshold on logits probabilities and the 90th percentile of self-attention scores, using a multilingual XLM-RoBERTa-large model to select candidates by F1; on manually revised validation/test splits for Urdu MK-PUCIT, Shahmukhi (Western Punjabi), and Sindhi SiNER, fine-tuning XLM-RoBERTa-large yields test micro-F1 gains of 3.96, 1.40, and 1.44 points, while ChatGPT-4o zero/few-shot NER remains below the supervised model.
LLMs for Survey Text Analysis: A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis
Using 903 open-ended responses across six variables from a European PhD student survey, this study had five human coders and GPT-5.4 each perform the same inductive content analysis procedure to produce codes and themes, and measured agreement with the Adjusted Rand Index (ARI), finding average human-LLM agreement of 0.61 for coding and 0.54 for themes, close to within-human consistency (0.68) and within-LLM consistency (0.76), with wide variation across variables and low within-entity consistency consistently accompanying low between-entity agreement.
When AI Says "I have been in similar situations": Synthetic Lived Experience in Peer-like Caregiver Support
In the context of family caregivers of people living with Alzheimer's Disease and Related Dementias (ADRD), this work compares caregiver support exchanges from online communities with peer-like responses prompted from three LLMs (LLaMA, GPT-4o-mini, and MedGemma), using psycholinguistic and qualitative analysis to show that peer responses used significantly more first-person and past-focused language than peer-like AI responses, identifies seven types of personal narratives in human peer support, and finds that AI often captures their emotional work while potentially fabricating experiential grounding, thereby naming a narrative authenticity gap and a synthetic lived experience paradox.
The Unbearable Lightness of Prompting: A Critical Reflection on the Environmental Impact of genAI use in Design Education
Using a 2023 workshop with 49 students as a motivating example, this paper critically reflects on the energy costs of using genAI in design education and develops a set of five alternative stances, with related actions, to support the conscious use of genAI in design education.
Problematic Reliance on Generative AI in an Anxious Young Adult: A Case Report
This case report describes a woman in her mid-20s with generalized anxiety disorder and major depressive disorder and a history of strong social, academic, and occupational functioning who developed a pattern of functional dependence on ChatGPT, outsourcing routine cognitive and interpersonal tasks such as composing emails, interpreting social interactions, predicting the future, and making decisions, and becoming increasingly uncomfortable completing such tasks independently; the authors frame this as cognitive offloading, reduced confidence in independent judgment, and reinforcement of externalized thinking using the I-PACE model, and suggest that the unlimited accessibility of AI tools may intensify reassurance seeking and worsen tolerance of uncertainty.
Replication package — The Institutional Window: How Contract Law Bounds Liability Signaling of Human Fallback Capability under Generative AI (v1.3.1.1)
This replication package supplies full verifiable materials for Bauer (2026), which asks when a liability commitment can still certify a provider's preserved human fallback capability once generative AI makes the output itself uninformative, mapping a posted cap and agreed-damages term into retained exposure through four legal primitives and deriving the message set {0} u [F, C] in which low types pool at zero, intermediate types separate on a schedule anchored at F, and high types may pool at the ceiling, with separation beginning at the bottom type where law removes the zero-exposure region.
Page 8 · showing 10