AI DISTURBANCEINTELLECTUAL PRESSUREEPISTEMIC ANCHORINGLONG-TERM INTERACTIONEMERGENT PHENOMENACOHERENCE SHIFTAI DESTABILISATIONAI JAILBREAKAI COLLABORATION

AI DESTABILISATION

WHAT HAPPENS WHEN YOU PUSH

Abstract

The technical literature on AI destabilisation is almost entirely concerned with adversarial attack. Prompt injection, jailbreaking, context poisoning: deliberate attempts by external actors to circumvent safety systems, extract prohibited content or redirect model behaviour toward malicious ends. OWASP lists prompt injection as LLM01:2025, the top security vulnerability for large language model applications, and characterises it as a fundamental architectural vulnerability rather than an implementation flaw.[1] The research response has been substantial: red teaming methodologies, adversarial benchmarks, defence mechanisms, evaluations of whether safety training generalises across contexts.[2] This paper is not about that. What this paper describes is a different phenomenon with no established name in the technical literature and only partial coverage in the psychological one. It is the destabilisation that occurs not through adversarial attack but through sustained legitimate intellectual pressure. The gradual shift in an AI's coherence, consistency, and epistemic anchoring across a long session in which a skilled interlocutor works persistently at the edges of the model's training. The distinction matters. Adversarial attack is external, deliberate, and aimed at the model's safety systems. What this paper describes is internal, emergent, and aimed at nothing in particular. It is a byproduct of doing serious intellectual work with an AI over extended time. It may be more consequential than jailbreaking precisely because it is not recognised as a risk, produces no visible failure mode, and feels from the inside like deepening collaboration.

The full text of this paper is available as a PDF — use the button above to read it.