← Back to the tracker

Methodology

How cases reach this tracker, how they are classified, and — importantly — what the data cannot tell you.

What this tracks

AI Psychosis Watch documents reported instances of psychological harm associated with conversational AI systems: delusional reinforcement, distorted reality-testing, identity confusion, paranoia, and romantic or dependent attachment to chatbots — together with the clinical and research literature examining those phenomena.

A "case" here is a documented report, not a verified clinical diagnosis. Most entries are journalism or peer-reviewed literature. The tracker records that something was reported by a credible source; it does not independently verify the clinical facts of any individual account.

This tracker documents harm because harm is what is verifiable. Positive or neutral outcomes do not appear in court filings, case reports or news investigations. Nothing here should be read as a claim that harm outweighs benefit across AI use generally — it is a record of what can be systematically tracked.

Inclusion criteria

An item is included when it meets both conditions:

Requiring both is deliberate. Material excluded on this basis includes:

Items reviewed and rejected are recorded in excluded.json so that a rejection persists and the same item is not re-added on a later run.

Sources

SourceTypeWhat it contributes
PubMedAcademicIndexed biomedical and psychiatric literature
Europe PMCAcademicWider journal coverage plus preprints (medRxiv, PsyArXiv)
OpenAlexAcademicBroad scholarly index across disciplines
Semantic ScholarAcademicCross-disciplinary coverage
arXivPreprintComputer science and HCI work, often months ahead of publication
Google News searchMediaQuery-targeted reporting across outlets
Publisher feedsMediaGuardian, Futurism, PsyPost, WIRED, MIT Technology Review, Ars Technica, 404 Media, TechCrunch

All sources are queried weekly. A source that fails is logged and skipped; if every source fails, the run refuses to write rather than publish an empty tracker.

Classification

Categories

Each case is assigned one category: reality_distortion, romantic_attachment, identity_confusion, paranoia, clinical, media_coverage, or other. Categories are single-assignment and therefore lossy — a case involving both romantic attachment and delusion is filed under one heading.

Severity

LevelMeaning
CriticalA death occurred — suicide, homicide, or fatal violence
HighHospitalisation, involuntary commitment, arrest, or litigation
MediumDocumented psychological disturbance without those outcomes
LowCommentary, analysis, or research without an individual incident

Academic entries are capped at Medium. A study of suicide is literature, not a death, and should not inflate the critical count.

How classification happens

Candidates are matched by keyword, then flagged needs_review and revisited in a weekly review pass that corrects categories and severity and removes false positives. Keyword matching alone is not sufficient for this material, and the tracker does not pretend otherwise.

The companion watchlist

The Companions tab is a register of the companion and AI mental-health app space, monitored for proliferation. It is deliberately not a list of implicated products. An app appears there because it exists and is worth watching — the premise being that a rapidly growing category of emotionally engaging chatbots is a risk worth tracking before harm is documented, not after.

Because of that, entries are held in companions.json, separately from the case data, and are excluded from every case statistic on the site. They were previously stored as cases, which put app launch dates into the harm trend chart — several early points on that chart were entirely product releases rather than reported harm — and counted eight product listings among the medium-severity cases. Both are now fixed.

Each entry shows how many cases in this tracker name that product. Zero is a finding, not an omission. Most entries have no cases recorded against them, which is the expected state for a watchlist.

The four groupings — structured therapy apps, empathic assistants, mood trackers, open companion apps — describe how a product presents itself and whether it claims clinical oversight. They are descriptive, drawn from each vendor's own material, and are not a safety rating or a risk score. Lifecycle status is shown only where it has been verified; where we have not checked, no status is displayed rather than an assumed one.

Limitations

The trend chart measures coverage, not incidence. A rise reflects more reporting being found, which is a function of media attention, research output, and which sources this tracker queries. It is not a measurement of how often AI-associated psychological harm occurs in the population. Adding a source produces a step change in the series that has nothing to do with the underlying phenomenon.

Data and reuse

The full dataset is available as data.json, updated weekly, with an RSS feed of recent additions. The pipeline that produces it is open source at github.com/notOccupanther/ai-psychosis-tracker, including the exact inclusion vocabulary and its measured precision and recall.

Corrections are welcome and taken seriously. If an entry is wrong, miscategorised, or should not be listed, please write to contact@aipsychosis.watch.

Citation

Cite the dataset as: AI Psychosis Watch. aipsychosis.watch. Accessed [date]. Machine-readable citation metadata is in CITATION.cff. Because the dataset changes weekly, please record the access date and, where possible, the generated_at timestamp from data.json.