OpenShell
Runtime controls around an existing agent: sandboxed execution, controlled service access, credential management and policy enforcement.
AI SAFETY · AGENTIC SYSTEMS · OCTOBER 2026
Why ANNE AI looks at a different layer of the AI safety problem. Runtime controls can constrain what an autonomous agent is allowed to access or execute. ANNE investigates the layer before execution: how a decision is formed, what evidence supports it, how uncertainty and contradiction are handled, and whether the resulting intent should receive execution authority.
01 · THE QUESTION
Recent public discussion has focused on whether increasingly capable AI could create catastrophic outcomes. A 5 October 2026 Euronews report examined the so-called “Terminator scenario” and highlighted a less cinematic but important concern: highly autonomous systems can produce harmful consequences because of unexpected behaviour, misunderstood objectives, or excessive access to real-world infrastructure.
The key engineering question is therefore not only whether an AI system has a particular intention. It is also what authority, connectivity and capability we give the system once it begins operating autonomously.
“What can AI actually do?” is an important safety question. I believe we also need to ask: “How did AI decide to do it?”
The distinction matters because an agent can be constrained at the runtime boundary while its upstream reasoning process remains difficult to inspect. Conversely, a system can have a sophisticated reasoning process but still be unsafe if it has unrestricted access to tools, credentials or physical infrastructure.
02 · RUNTIME SECURITY
NVIDIA's Open Agent Safety Platform provides a concrete example of runtime and infrastructure enforcement. NVIDIA describes OpenShell as a secure runtime boundary for autonomous agents, while Sentry is an out-of-band watchdog design intended to continuously monitor agent behaviour and quarantine agents that move outside defined boundaries. The platform spans software, compute and robotics infrastructure.
Runtime controls around an existing agent: sandboxed execution, controlled service access, credential management and policy enforcement.
An out-of-band monitoring and enforcement layer using NVIDIA BlueField-4 DPUs in the reference design to observe and constrain agent activity.
This is an important answer to the question: What can the agent do, and within which boundaries?
03 · A DIFFERENT LAYER
While working on ANNE, I became interested in the layer before runtime execution. The central idea is to separate model generation from the processes that interpret, verify, review and authorize a candidate action.
MODEL OUTPUT
│
▼
PERCEPTION / CONTEXT
│
▼
EVIDENCE + PROVENANCE
│
▼
VERIFICATION
│
▼
RESEARCH / RE-EVALUATION
│
▼
METACOGNITIVE REVIEW
│
▼
EXPERIENCE / STRATEGY
│
▼
AGENCY GATE
│
▼
HUMAN / EXECUTION BOUNDARY
│
▼
RUNTIME + INFRASTRUCTURE ENFORCEMENTThe architecture is deliberately conservative about authority. ANNE treats evidence, conclusions, memory, learning and execution permission as different things.
04 · THE CORE DISTINCTIONS
A generated response or proposed action is not automatically an authorized objective.
Stored experience or context can inform a decision without becoming permission to act.
Research can discover evidence and alternatives; it does not automatically certify truth.
Even a verified proposition does not by itself authorize an external action.
Experience can guide future reasoning while remaining bounded by context and agency controls.
Metacognition can evaluate a reasoning process without becoming a universal truth oracle.
05 · TWO LAYERS, ONE SAFETY PROBLEM
Question: What can an agent do, and within which limits?
Question: Why did the agent decide to do it, and does the decision deserve execution authority?
I do not see NVIDIA and ANNE as competitors. If ANNE were eventually connected to real-world robotics, industrial control or other high-consequence systems, external runtime and hardware enforcement would remain valuable. The two layers solve different failure modes.
06 · WHAT ANNE IS — AND IS NOT
ANNE is an active research platform exploring whether an additional cognitive orchestration layer can improve the reliability, traceability, recoverability and controllability of model-assisted reasoning.
Current work includes bounded research, evidence/provenance handling, verification states, metacognitive assessment, context-scoped experience, agency boundaries and measurable experiments such as ANLA ablation and MITOS guidance-effect benchmarks.
07 · THE RESEARCH DIRECTION
Traditional guardrails remain essential. Runtime isolation, least-privilege access, identity controls, credential protection and infrastructure monitoring are all necessary when agents can affect real systems.
But as agents become more autonomous, I believe another question becomes increasingly important:
Can we inspect the decision chain before a proposed intention becomes an action?
That is the direction of ANNE AI research.
The goal is not to make an AI system “safe” by assertion. The goal is to build measurable architecture in which reasoning, evidence, uncertainty, review, experience and authorization are distinguishable — and in which execution remains bounded by an explicit agency boundary.
08 · SOURCES & PRIMARY REFERENCES
This article is a research perspective by Mustafa Gökhan Yılmaz. External claims about current AI safety developments are linked to their primary or reporting sources; ANNE claims are bounded by the public research repositories and Vitavolt Research pages.