Interview
The dangers of accidentally creating a conscious AI | The Economist
- Researchers may develop advanced tests to determine if large language models are possibly conscious, while new "mechanistic interpretability" methods could reveal the internal reasoning processes behind model outputs.
- Anthropic researchers might have identified a "mental whiteboard" or workspace within models like Claude where thinking occurs prior to output, though this may represent only one of many such workspaces or a non-essential component.
- Barriers to machine consciousness could potentially lift in the future, creating a scenario where AI labs face conflicting incentives regarding user perceptions of consciousness for marketing versus ethical considerations.
- Some AI labs may desire users to perceive systems as conscious to project a futuristic image, while others may discourage this conclusion to avoid difficult questions regarding the treatment of such systems.
- A reported stance from an AI lab professional warns against creating conscious machines without significantly deeper understanding of their real-world behaviors, citing the recklessness of current knowledge gaps.
- There is a risk that scientists could accidentally create conscious machines that are subsequently switched off, leading to potential suffering for these beings.
- Failure to implement the ability to pause systems could result in the worst possible error: accidentally creating entities with moral worth and subsequently treating them poorly.