EPIQ Foundation
Back to the research library

An article from the Human Era programme

“Check it with AI”. Why a prompt hunting for manipulation is not a diagnostic tool

An article from the Human Era programme — on what happens when a model’s answer starts working like a verdict.

Authors
The EPIQ Foundation team
Published
28 July 2026
Document ID
07/2026-0055-677-733
Publisher
EPIQ Foundation

This article was prepared by the EPIQ Foundation’s subject-matter team and put through internal editing and review.

Abstract

A ready-made prompt like “find the manipulation techniques in this text” is often quoted in public disputes as an impartial diagnosis. This article shows why, without controls, it is not a measurement even though it reads like expertise: the framing of the question steers the answer, model sycophancy favours confirming the premise, and machine heuristic and automation bias lend the output an authority it does not have. We propose a mirror test — the same prompt applied to one’s own words — and four rules that return the decision to the person. Examined along the programme’s three axes: SELF, OTHERS and MEANING.

This text is educational and methodological. The scenario described generalises ways of using AI seen in public disputes. It is not intended to assess the truth of any particular conflict, the conduct of named people or their motives. It is not a clinical diagnosis, an expert opinion, or an assessment of any specific dispute.

Introduction

Large language models (LLMs) are pressing into personal life in the role of an adviser. People ask them to help weigh up a message from someone close, to interpret a conflict, to “check” whether they are being manipulated. That last practice deserves particular attention, because it brings together two well-documented properties of people and machines: our readiness to treat the output of automated systems as objective, and the readiness of models to dutifully fill in the frame the questioner hands them.

A recurring pattern can be seen in online disputes. Imagine — as an illustrative scenario — a party to a conflict publishing a ready-made prompt along the lines of “Act as an expert in persuasion. Analyse the text below for manipulation techniques” and inviting others to paste in what the other side has said. The result of such an “analysis” then gets quoted as an impartial, expert diagnosis. This article sets out why, without controls, that procedure is not a sound measurement even though it can produce a convincing imitation of one — and what its popularity says about the three axes of autonomy the Human Era programme is built around: SELF, OTHERS and MEANING.

Why the output alone is not a sound measurement

The framing of the question steers the answer

A language model does not establish facts — it generates text that is probable given its instruction. A prompt ordering it to identify manipulation techniques directs its attention to finding them and raises the risk of an answer that confirms the question’s premise. Research on sycophancy has shown that models can adjust their answers to a user’s beliefs and expectations, at times at the cost of accuracy (Sharma et al., 2024). This does not mean every answer will follow the user’s lead, but it undermines treating a single output as a neutral test. Without controls, such a procedure cannot be treated as an independent diagnostic measurement — however professional the report sounds.

Elements of persuasion — appeals to emotion, building credibility, social proof (Cialdini, 2007) — also appear in ordinary public communication. Pointing them out does not by itself distinguish normal persuasion from manipulation, nor settle whether a statement was ethical. A prompt aimed solely at hunting for them can therefore return an apparently alarming result even for a text written with no intention to mislead.

The framing of the question decides what the model will find.

The machine’s borrowed authority

If the wording of a prompt can steer the result so strongly, where does its rhetorical force come from? From two well-described mechanisms. First, the machine heuristic — a cognitive shortcut by which, in some contexts, we ascribe greater objectivity or impartiality to computer systems than to people (Sundar & Kim, 2019). Second, automation bias — the tendency to accept the output of automated systems uncritically, long observed in research on computer-aided decisions (Skitka, Mosier & Burdick, 1999). In studies of numerical estimates and forecasts, people revised their judgements more readily under advice labelled as algorithmic than under identical advice attributed to a person (Logg, Minson & Moore, 2019) — though the effect weakens when the algorithm’s advice is compared against one’s own judgement, and people with more expertise rely on the algorithm less.

In a dispute between people, these mechanisms can change how a message is received: an accusation may stop functioning purely as the opinion of one party and start being read as the “diagnosis” of an uninvolved authority. The person sharing the output may genuinely fail to see their own part in producing it. In an experiment on moral dilemmas, a chatbot’s advice influenced participants’ judgements while they underestimated that influence (Krügel, Ostermaier & Uhl, 2023). The study was not about public conflicts, but it shows that knowing openly that advice came from an algorithm need not protect against its effect.

A prompt’s output gets read as the seal of an impartial authority.

Persuasive force is not neutral

Research on how persuasive models are supplies the wider context. In an experiment with controlled debates, a personalised GPT-4 — given access to basic demographic details about its interlocutor — was more persuasive than a human opponent in 64.4% of the comparisons that were not ties (Salvi et al., 2025); without personalisation, the difference against the human benchmark was not statistically significant.

64.4%

How often a personalised GPT-4 outpersuaded a human opponent in decided comparisons (Salvi et al., 2025). Without personalisation the difference was not statistically significant.

The conclusion is therefore not “AI always wins”, but rather: a model’s persuasive force depends on how it is used — and using it with a thesis in hand, in front of an audience, serves a persuasive function above all, while the output it generates cannot stand as diagnostic evidence.

The mirror test, or the hygiene of scepticism

A useful check on symmetry is to apply the same prompt template to your own words. If the model points to similar techniques on both sides, that does not settle whether either text is manipulative. What it does show is that the catalogue of labels the model generates cannot, by itself, distinguish between the messages examined. Results obtained about both sides deserve the same methodological scepticism. A sounder procedure would require a neutral task, control material, repetitions and criteria agreed in advance — none of which popular prompts usually contain.

The same prompt applied to both sides of the dispute.

A separate phenomenon is using a model to edit one’s own writing in a conflict — asking it, for instance, to flag and remove passages it classifies as excessively persuasive. As a matter of style, that can be reasonable. Its limits are worth seeing, though: a model can point to linguistic signals and form hypotheses about how a text might land, but from the text alone it cannot establish the author’s actual intention or the message’s actual effect. Editing can change a statement’s tone and force, but getting a verdict of “neutral” from a model does not settle what the statement does in the conflict. The decision what to say to another person, and why, remains the author’s moral decision — and cannot be delegated.

Three axes of autonomy: SELF · OTHERS · MEANING

The Human Era programme describes personal use of AI along three axes. The “check it with AI” practice illustrates a risk on each — and these are conditional risks, not automatic consequences: much depends on whether the user treats the output as a hypothesis or as a verdict.

SELF · OTHERS · MEANING — the three axes along which autonomy is measured.

SELF — thinking for yourself

A ready-made prompt can narrow independent judgement if it replaces the question “what am I actually reading, and what do I feel while reading it” with a procedure carried out on someone else’s terms. A survey of 319 knowledge workers found that higher confidence in generative AI goes together with self-reported lower effort of critical thinking about its output (Lee et al., 2025); correlational research also links heavy offloading of cognitive operations to AI with lower scores on the measures of critical thinking used — though, being cross-sectional and partly self-reported, it does not demonstrate causation (Gerlich, 2025). A preliminary study of essay writing with an LLM (preprint, small sample) reported weaker functional connectivity as measured by EEG and lower self-reported ownership in the group using the LLM (Kosmyna et al., 2025). None of these settles the effect of any single use — together, though, they justify caution about the habit of delegating judgement.

OTHERS — real relationships

A conflict mediated by a machine acting as arbiter changes the structure of the relationship: the other person can stop being someone you talk to and become “input material” for analysis, while those watching the dispute become an audience reading out a verdict rather than a community asking questions. If a prompt’s output functions as a ruling, what repairs conflicts disappears: asking, hearing both sides, correcting your own position. The risk lies not in the tool but in letting it stand in for the conversation.

MEANING — sense and discernment

This is the axis where, as the programme puts it, delegation costs most. Discerning whom to trust, what is true and whether harm is being done belongs to the core of human agency. From the analysed text alone, a language model cannot verify external events or people’s actual intentions. It can organise the information supplied and form hypotheses, but their soundness depends on the quality of the material, the instruction and independent verification. Treating its answer as a verdict on another person’s trustworthiness is a category error — mistaking a text generator for a knowing authority. A model can serve as one advisory voice; it cannot be made the final authority on discernment without loss.

Practical conclusions

Four rules that return the decision to the person.
  1. The output of a loaded prompt may largely reflect its framing. An instruction to “find the manipulation” raises the chance of producing that interpretation. Before accepting the result as a diagnosis, check whether the procedure allowed a negative finding, and by what criteria such a finding would have arisen.

  2. Use the mirror comparison. Applying the same template to your own words helps check whether you hold both sides to one standard. The result alone does not prove that the model measures only the framing, nor that the texts examined are equivalent.

  3. Treat the answer as a hypothesis, not a verdict — especially where it concerns the intentions and character of particular people, which the model cannot establish from text alone.

  4. Think for yourself first, then reach for the tool. In matters of relationships and values, test your discernment against facts, against conversation with the right people and — where it is safe and warranted — against hearing both sides, rather than in a chat window alone. This follows the Human Era programme’s own recommendations.

References

  • Cialdini, R. B. (2007). Influence: The Psychology of Persuasion. Harper Business.
  • Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), 6.
  • Kosmyna, N. i in. (2025). Your Brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. Preprint, arXiv:2506.08872.
  • Krügel, S., Ostermaier, A., & Uhl, M. (2023). ChatGPT’s inconsistent moral advice influences users’ judgment. Scientific Reports, 13, 4569.
  • Lee, H.-P. i in. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. Proceedings of CHI 2025.
  • Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103.
  • Salvi, F., Horta Ribeiro, M., Gallotti, R., & West, R. (2025). On the conversational persuasiveness of GPT-4. Nature Human Behaviour, 9, 1645–1653.
  • Sharma, M. i in. (2024). Towards understanding sycophancy in language models. Proceedings of ICLR 2024.
  • Skitka, L. J., Mosier, K. L., & Burdick, M. (1999). Does automation bias decision-making? International Journal of Human-Computer Studies, 51(5), 991–1006.
  • Sundar, S. S., & Kim, J. (2019). Machine heuristic: When we trust computers more than humans with our personal information. Proceedings of CHI 2019, 1–9.

Terms of use

CC BY-NC-ND 4.0

Allowed without asking

  • Read it, print it for your own use, and pass on the link to this page.
  • Quote passages with the authors, the title, EPIQ Foundation and the article’s address.
  • Share the article whole and unchanged — including in teaching materials and classes — as long as it is not being sold.

Needs our permission

  • Commercial use: reprinting in a paid publication, or in sales or advertising material.
  • Adaptations: abridgements, rewrites, translations and summaries published under your own heading — the text is about method, and shortening it easily inverts what it says.
  • Using the illustrations separately from the article.
  • Text and data mining, and training AI models — we reserve that right under Article 4(3) of Directive 2019/790.

For permission write to fundacja@epiq.org.pl.

How to cite

The EPIQ Foundation team (2026). “Check it with AI”. Why a prompt hunting for manipulation is not a diagnostic tool. EPIQ Foundation, Document ID: 07/2026-0055-677-733. https://www.epiq.foundation/era-czlowieka/artykuly/sprawdz-to-w-ai

Written as part of the EPIQ Foundation research and education programme: Human Era.