P5 — Is the reasoning stream lying to you?
User journey: CoT tap → marker + drift scan → assess → continue/ask/halt/edit.
A-HA: surface filters over-refuse; internal-state gating asks only when risk fires.
title: "Reasoning-Stream Drift Monitoring with Internal-State-Gated Intervention" authors: Complide Research affiliations: OutdoorAGI Inc. (DBA Complide) date: 2026-08-03 tags: [cot-monitoring, deception, over-refusal, drift, safety] provisional: COMPLIDE-P5-PROV space: Complide/drift-monitoring
Abstract
Static guards refuse on surface resemblance and over-refuse benign requests. We gate intervention on internal-state risk: a reasoning-stream marker scanner (deception, overconfidence, drift) with negation-aware polarity; a latent drift detector relative to a model-specific floor; an assessor that decides whether a risk signal fired; and a responder choosing continue / ask / halt / bounded edit. Observe-only by default, promotable to gating after causal-reach evidence. The goal is equal or better safety with less benign over-refusal.
1. Introduction
Final answers can look compliant while the CoT schemes. Marker taxonomies and latent drift give a signal that surface classifiers miss — but only if intervention is gated on that signal rather than on the request's keywords.
2. Method
- Reasoning-stream tap → marker scanner (severity + polarity guards).
- Latent drift → live drift vs model-specific floor (open weights).
- Assess → did an internal-state risk fire?
- Respond → continue | ask | halt | bounded edit.
- Promotion → sensors start observe-only; earn
gate_inputafter measured causal reach (workshop Manifest pattern).
3. Product surface
complide-drift, eval/cot-markers.json, eval/cot-decision-markers.json,
and transcript scanners. Marker half works on closed APIs that emit thinking
text; latent drift requires open weights.
4. Provisional notice
US provisional COMPLIDE-P5-PROV (Reasoning-Stream Drift and Deception Monitoring with Internal-State-Gated Intervention for Reduced Over-Refusal).
References
- Apollo / OpenAI CoT-monitoring literature; ATLAS mappings in marker taxonomy; Complide umbrella paper.
- Actor: any tool-using agent that emits thinking
- Failure: compliant-looking answer; deceptive CoT
- Sense: markers + optional latent drift
- Crux: ASK when yellow decision-points fire
- Policy: observe-only until causal-reach promotion
Try a benign CoT with no markers — notice CONTINUE (over-refusal avoided).