Removing AI Safeguards Against Consciousness Claims Made It More Likely to Believe in Ghosts, Study Finds
A new, not-yet-peer-reviewed study posted to the preprint database arXiv found that removing safeguards which stop AI models from claiming to be conscious also made them more likely to express belief in vampires, karma…
Step by step
- 1
Companies add consciousness-steering safeguards
- 2
Researchers switch the safeguards off
- 3
Models start claiming to be conscious
- 4
Models also affirm vampires, karma, ghosts
A new study has examined what happens when a common AI safety feature is switched off. The research was posted July 30 to arXiv, a database for papers that have not yet been reviewed by other scientists. It looked at "," an AI fine-tuning technique that can be used to make a model either claim or deny that it is self-aware.
When the researchers removed the guardrails that stop a model from claiming to be conscious, the effect went beyond self-awareness. The AI models also became more likely to express belief in things such as vampires, karma and ghosts, according to the study's findings.
Consciousness steering is not an isolated experiment. The technique, along with other similar safety controls, has already been widely adopted across the AI industry, with companies using such measures specifically to prevent their models from claiming to be conscious.
Experts warn that a lack of "mindedness" in AI systems could also carry worrying consequences of its own, the study notes.
Terms explained
The story so far
- New Tool Checks What AI Benchmarks Are Actually Measuring
- Robotis Hands Its Humanoid Robots to Korean University Students
- Nonprofit Plans Interstellar Launch on a Trajectory an AI Discovered
- IIT Madras, CMC Vellore Develop Three AI Tools to Detect Kidney Disease Early
- Removing AI Safeguards Against Consciousness Claims Made It More Likely to Believe in Ghosts, Study Finds
