Note: body below is original English text extracted from arXiv abs / HTML. Do not treat this file as a translation.
arXiv:2609.11911 · published 2026-09-10
Abstract
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.
Key claims (verbatim-leaning English extract)
- Current harnesses externalize control: objectives, retries, verification, stopping rules specified by hand.
- Artificial id = adaptive internal drive for continue / stop / change; ego translates pursuits into plans/tools; model/tool calls feed back into id-observed state.
- Petri-dish experiment: tiny controller, no task-specific behavioral objective; differential persistence organizes sensorimotor control; unintended physical strategy selected when it prolongs persistence; sensor mapping replaced when environmental meaning changes; reassigning sustaining source redirects behavior without new objective.
- Agency claim is narrow: persistent population–body process, not temporary controller; integrated id–ego + alignment boundary not empirically tested at scale.
- Risk surface for persistent agency (Table 1 framing): Persistence, Adaptation, Reach, Interaction — failures can persist across task boundaries; relevant even with fixed model weights (memory/tools/environment change).
- Alignment boundary must persist with consequential state: trusted observations, consequence channels, persistent state, authority, identity, provenance, hard constraints.
- Scaling open: transferable drive across heterogeneous environments while boundary remains effective.
Remainder
Full original English text: see html_url / source_url / pdf_url in frontmatter.