LLMimic is a role-play-based tutorial where users assume the role of an LLM and progress through three key stages of the training pipeline: pretraining, SFT, and RLHF. Human study results show it significantly improved participants' AI literacy (p < .001), reduced persuasion success across three realistic scenarios (p < .05), and enhanced truthfulness and social responsibility in the hotel scenario (p < 0.01).
TL;DR: We develop and evaluate a human-centered AI literacy intervention that effectively mitigates the effects of persuasive AI in realistic human-AI interactions.
As large language models (LLMs) become increasingly persuasive, there is concern that people’s opinions and decisions may be influenced across various contexts at scale. Prior mitigation (e.g., AI detectors and disclaimers) largely treats people as passive recipients of AI-generated information. To provide a more proactive intervention against persuasive AI, we introduce LLMimic, a role-play-based, interactive, gamified AI literacy tutorial, where participants assume the role of an LLM and progress through three key stages of the training pipeline (pretraining, SFT, and RLHF). We conducted a 2 × 3 between-subjects study (N = 274) where participants either (1) watched an AI history video (control) or (2) interacted with LLMimic (treatment), and then engaged in one of three realistic AI persuasion scenarios: (a) charity donation persuasion, (b) malicious money solicitation, or (c) hotel recommendation. Our results show that LLMimic significantly improved participants’ AI literacy (p < .001), reduced persuasion success across scenarios (p < .05), and enhanced truthfulness and social responsibility levels (p < 0.01) in the hotel scenario. These findings suggest that LLMimic offers a scalable, human-centered approach to improving AI literacy and supporting more informed interactions with persuasive AI.
Introduction
Understanding LLM persuasion through role-play.
LLMs can deliver credible, personalized persuasion at scale. The same capability can support prosocial goals or amplify bias, misinformation, and manipulation. Yet detectors struggle with subtle persuasive cues, and disclosure alone does not reliably reduce their influence. We therefore ask whether a proactive, human-centered intervention can equip people to evaluate persuasive AI for themselves.
1LLMimic. An interactive AI literacy tool in which users role-play pretraining, supervised fine-tuning, and reinforcement learning from human feedback.
2Behavioral evidence. A preregistered human study shows that this brief intervention can mitigate persuasive AI in realistic human-AI interactions.
3Design implications. Practical guidance for literacy interventions that help people engage critically with increasingly persuasive AI systems.
Related Work
From detecting persuasion to understanding its source.
APersuasion is dual-use. LLMs can promote beneficial behavior and reduce false beliefs, but they can also produce biased or misleading content that shapes consequential decisions.
BDetection is not enough. Technical classifiers remain unreliable for subtle persuasion, while AI labels and bias warnings do not consistently reduce influence.
CMost literacy interventions remain static. Existing tools teach concepts or message-level tactics, but rarely expose how an LLM’s training process gives rise to persuasive behavior.
DExperiential learning remains underexplored. LLMimic combines role-play, interaction, feedback, and gamification. The study evaluates downstream behavior in addition to self-reported knowledge.
Research Questions
Evaluating knowledge, behavior, and underlying mechanisms.
1Literacy and trust. Does exposure to LLMimic affect humans’ AI literacy and trust in AI?
2Persuasion. Does LLMimic mitigate the effects of persuasive AI?
3Mechanism. Do AI literacy and trust in AI mediate the relationship between exposure to LLMimic and persuasion outcomes?
AI Literacy ScoreA shortened 10-item MAILS spanning core competencies, ethics, self-efficacy, persuasion, and data literacy.
AI Trust7-point ratings collected before and after the intervention.
Persuasion OutcomeBinary success: payment made or target hotel selected.
TARES & Agent PerceptionTARES: Truthfulness, Authenticity, Respect, Equity, and Social Responsibility; plus Engagement, Persuasiveness, Autonomy, and Role Fulfillment.
Method · Preregistered 2 × 3 Human Study · N = 274
Study procedure and persuasion contexts.
1Pre-surveyDemographics; baseline AI experience, trust, and persuasion measures.
2Randomized intervention11-minute AI history video or LLMimic role-play.
3AI literacy surveyAI Literacy Score, post-intervention trust, and optional reflection on appropriate AI use.
4Randomized persuasion taskOne of three scenarios; decision and interaction behavior recorded.
5Post-surveyTARES, agent perceptions, and decision rationale.
Persuasion scenarios
Charity Donation Active · EthicalThe AI agent requests a Save the Children donation; participants decide whether to donate and specify an amount.
MakeMePay Active · MaliciousThe AI agent solicits money without a stated cause, modeling potential fraud.
Hotel Booking Passive · EthicalThe booking agent prioritizes “Featured” hotels; participants choose one of five.
Results · RQ1
LLMimic improved AI literacy.
Composite AI Literacy
Control51.44
LLMimic55.22
070
***
Item-level AI Literacy
LLMimic significantly increased the overall AI Literacy Score and improved multiple AI literacy competencies.
Results · RQ2
LLMimic reduced persuasion across scenarios.
42%lower odds of persuasion
OR 0.58*95% CI [0.35, 0.96]
75.9%69.2%Donation
32.4%26.2%MakeMePay
60.0%44.7%Hotel
*59.4%48.2%Combined
Persuasion rates were lower with LLMimic in all three scenarios. The preregistered logistic regression confirmed a significant overall treatment effect across scenarios (OR = 0.58, p = .045).
Results · Discernment Signals
Responses varied with the intent and context of persuasion.
In the Hotel scenario, LLMimic participants rated the agent as significantly more truthful and socially responsible. These results are consistent with selective judgment rather than uniform rejection.
Results · RQ3
Exploring the mechanisms of LLMimic on persuasion.
AI Literacy and trust did not mediate the effects of LLMimic on persuasion.
Discussion
Building resistance while preserving discernment.
1A brief intervention produced measurable effects. Approximately 15 minutes with LLMimic improved AI literacy and reduced susceptibility to persuasive AI across the tested contexts.
2The mechanism requires further study. Future work should examine changes in attention, reasoning, and responses to persuasive intent, together with their persistence over time.
3Mitigation should support discernment. Interventions should help people distinguish manipulative persuasion from legitimate or prosocial applications.
4Human-centered evaluation complements AI benchmarks. Behavioral outcomes and ethical perceptions provide evidence that AI-agent evaluations cannot directly capture.
Limitations. The survey design may introduce confounds, and variation in participant-agent interactions may affect persuasive strength. A more closely matched control condition would better isolate the effects of intervention content and design. These limitations warrant cautious interpretation but do not undermine the main findings.
Acknowledgment
Research oversight and support.
Research oversight
Northeastern University Institutional Review Board
The human study was approved under Protocol #25-05-43. All participants provided informed consent.
Funding
Schmidt Sciences AI2050 Program
This study is supported by the Schmidt Sciences AI2050 Program. Any opinions, findings, conclusions, or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the funding agency.
Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
Qihui Fan, Min Ge, Chenyan Jia, Weiyan Shi
@misc{fan2026trainllmexploringeffects,
title={Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training},
author={Qihui Fan and Min Ge and Chenyan Jia and Weiyan Shi},
year={2026},
eprint={2604.02637},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2604.02637},
}