Flawed AI Summaries Distort Eyewitness Memory β Deep Summary
Original URL: https://neurosciencenews.com/ai-false-memory-llm-31261/ Published: 27 September 2026 (Neuroscience News) Underlying study: "AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory" β Mattea Sim, Yael Eiger, Tadayoshi (Yoshi) Kohno (Georgetown University / University of Washington) Study link: arXiv, 23 Sept 2026 β DOI 10.48550/arXiv.2609.28820 (open access) Venue: Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES)
Scope note: this summary is based on the Neuroscience News article, which includes the paper's abstract. The full paper was not read.
1. Core claim
Misleading AI-generated summaries of an event can change what people remember of an event they saw first-hand. This held even when readers knew a machine wrote the summary. The authors argue that the common safeguard of keeping a "human in the loop" to catch AI mistakes is weaker than assumed, because the human's own memory can absorb the AI's errors.
2. Why it matters (the setting)
- Institutions are adopting large language models to condense long video and audio into short narrative digests: corporate meetings, clinical case notes, and police body-worn camera logs.
- Psychology has long held that human episodic memory is reconstructive rather than a fixed recording, so it is open to post-event misinformation.
- The paper frames AI as a new delivery channel for that misinformation, and one that can do harm without any adversarial intent.
3. Study part one: how good are AI video summaries?
- Models tested: consumer multimodal tools including ChatGPT (OpenAI) and Gemini (Google).
- Material: animated traffic-incident videos adapted from classic eyewitness-reliability research.
- Findings:
- On average the summaries omitted 51.6% of central scene events.
- In 95% of runs they left out the single most consequential event: a car hitting a pedestrian.
- Errors included omissions, recurring hallucinations and descriptive gaps. Omission of critical details was the most frequent error type.
- Co-author Yael Eiger said she was struck by how poor the summaries were, and worried that police departments may deploy summarisation tools without rigorous testing.
4. Study part two: does a bad summary change memory?
- Participants: 331.
- Stimulus: an animated video of a red car approaching an intersection (stop sign or yield sign), turning, and hitting a pedestrian.
- Delay: 24β48 hours later, participants read a narrative summary of the incident.
- Some summaries were accurate.
- Others contained altered, erroneous details.
- Provenance manipulation: participants were told the summary came either from an AI model or from a human transcriber.
- Outcome measure: a memory recognition test on the original event.
Results
| Group |
Memory accuracy for the original event |
| Read an accurate summary |
83.6% |
| Read a misleading summary |
44.8% |
- Telling readers the text was AI-written gave no protection. False-memory acceptance was about the same whether they thought a machine or a human wrote it.
- Participants' prior trust in AI and familiarity with AI also did not buffer the effect.
5. The "human-in-the-loop" argument
- The authors say people are often expected to correct AI mistakes, but their results suggest human memory can instead be distorted by those mistakes.
- Illustrative risk: an officer or witness reads an erroneous AI incident summary before writing a statement or deposition, and unknowingly adopts the AI's errors as their own memory.
6. Quotes and attributions from the article
- Mattea Sim (lead author, Georgetown's Massive Data Institute) says AI is a new way of delivering misinformation, and urges critical thought about whether and how AI should summarise information in high-stakes settings.
- Yael Eiger (UW PhD candidate) is concerned about police use without rigorous testing.
- Yoshi Kohno (Georgetown, co-author) places the work within a growing effort on humanβAI relationships, grounded in psychology and computer science.
7. Planned next steps
The team intends to move from synthetic animations to real police body-worn camera footage, to measure how commercial summarisation tools change official reports and civilian testimony in legal proceedings.
8. Key numbers at a glance
| Metric |
Value |
| Central events omitted by AI summaries (average) |
51.6% |
| AI summaries missing the car-hits-pedestrian event |
95% |
| Participants |
331 |
| Delay before reading summary |
24β48 hours |
| Memory accuracy, accurate summary |
83.6% |
| Memory accuracy, misleading summary |
44.8% |
| Effect of labelling the text "AI-written" |
None observed |
Source article: https://neurosciencenews.com/ai-false-memory-llm-31261/ β original research via arXiv DOI 10.48550/arXiv.2609.28820.