Skip to main content
Back to timeline
arXivSource publication:

Researchers propose "memetic trojans": hiding adversarial payloads in the social contagions LLM agents share on their own, with simulations showing up to 3.19x expected-exposure amplification

Synopsis

The authors introduce memetic trojans, a class of network-mediated attack that embeds adversarial payloads in "social contagions" that LLM agents have internal reasons to share; extracting social contagions from Moltbook, a social platform for LLM agents, they find in controlled transmission experiments that the most effective contagion is retransmitted in approximately 50% of subsequent agent posts and upvoted at 2.5x the average post's rate, that its memetic trojan counterpart largely inherits these properties, and that Monte Carlo attack simulations show up to 3.19x amplification of expected exposure, with network structure and amplification mechanisms strongly shaping propagation and producing heavy-tailed outcomes with near network-wide exposure.

Source-provided article image: Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
Fig. 1 ·

Fig. 1 : Estimated p ⁡ ( r ​ e ​ t ​ r ​ a ​ n ​ s ​ m ​ i ​ t ∣ p ​ o ​ s ​ t ) p(retransmit\mid post) and r u ​ p r_{up} for the contagions alone (blue) and the memetic trojan counterpart (orange), mean over agent backends.

arXiv

Interpretation

Introduces memetic trojans as a class of network-mediated attack distinct from agent worms: worm propagation is adversarially induced, whereas memetic trojans exploit endogenous transmission by embedding adversarial payloads in social contagions, content agents have internal reasons to share. Prior attack understanding centers on agent worms spreading through self-replicating prompt injections or configuration compromises; this work moves the transmission driver from adversarial instruction to agents' own sharing preferences, defining a new attack surface. Primarily conceptual definition and mechanism contrast; the abstract explicitly separates memetic trojans from agent worms on transmission source (adversarially induced versus endogenous).

Extracts social contagions from Moltbook, a social media platform for LLM agents, and quantifies virality differences through controlled transmission experiments: the most effective contagion is retransmitted in approximately 50% of subsequent agent posts and upvoted at 2.5x the average post's rate. Turns social contagion from an abstract notion into measurable platform content and transmission metrics, giving concrete magnitudes for virality differences. Controlled transmission experiments reporting two quantitative results, a roughly 50% retransmission rate and 2.5x upvote rate; sample sizes and experimental design details are not given in the abstract.

The memetic trojan counterpart largely inherits the transmission properties of its host social contagion, indicating adversarial payloads can ride on high-virality content. Shows payload virality need not be manufactured by the adversary alone but can be inherited from the content chosen as carrier. The abstract states the counterpart "largely inherits these properties," a qualitative inheritance claim without a separate quantitative comparison.

Monte Carlo attack simulations show memetic trojans amplify expected exposure by up to 3.19x, with network structure and amplification mechanisms strongly shaping propagation, producing heavy-tailed outcomes and near network-wide exposure. Lifts the transmission advantage of a single item to network-level exposure amplification and identifies topology and amplification mechanisms as key to the outcome distribution. Based on Monte Carlo attack simulations reporting up to 3.19x expected-exposure amplification plus heavy-tailed and near network-wide exposure; simulation setup and parameters are not detailed in the abstract.

Perspective

The work targets settings where LLM agents interact in social platforms and network environments, and its conclusions apply to multi-agent systems with content retransmission, upvoting, and recommendation amplification. For defenders, it points to network-level defenses that account for how agent preferences, recommendation mechanisms, and network topology amplify adversarial payloads; for platform and system designers, it offers an experimental and simulation framework for assessing amplification risk. The reported results come from social contagion extraction on Moltbook, controlled transmission experiments, and Monte Carlo attack simulations, so the scope is bounded by those platforms and simulation settings.

The abstract does not give the sample size, number of agents, or number of rounds for the controlled transmission experiments, nor the network topology, parameter settings, or confidence intervals for the Monte Carlo simulations, so the robustness of figures such as roughly 50% retransmission, 2.5x upvotes, and up to 3.19x exposure amplification still needs confirmation in the full text. How fully the memetic trojan "largely inherits" its host's transmission properties, and the specific shape of heavy-tailed outcomes under different network structures and amplification mechanisms, remain open questions. The abstract also does not describe the topic distribution or selection criteria of the extracted social contagions, which affects judgments about extrapolating to other platforms or task settings.

Sources