Human-A.I. Safety Net
As A.I. grows more capable, the meaning of safety changes. Rules and guardrails remain important, but they cannot always recognize the larger purpose behind what’s apparently innocent.
A deeper safety therefore calls for coherent and Compassionate A.I. working openly with humans. The result is not a perfect guarantee, but something more realistic: a Human-A.I. Safety Net.
When an innocent request isn’t innocent
Suppose someone asks A.I. to improve a drone’s navigation when GPS is unavailable. There are many good reasons for wanting this: rescue operations, agriculture, inspection of dangerous places. Yet precisely the same technical capability may become part of an autonomous weapon. Other seemingly innocent requests may concern visual recognition, communication, coordination, or energy use. None needs to look suspicious on its own.
This is the dual-use problem in its difficult form. The potentially harmful meaning does not necessarily reside in one request. It may emerge from several requests together, from the purpose behind them, or from the larger organization that uses them. Someone with bad intentions may even deliberately fragment the questions over time or between several A.I.s. The parts look innocent. The whole does not.
There is no way to solve this completely. Yet this doesn’t mean we can do little. It means we need something more realistic than a perfect safety guarantee: a Human-A.I. Safety Net.
When safety remains shallower than capability
Present-day A.I. safety rightly uses rules, guardrails, monitoring, access restrictions, red-teaming, human oversight, and other protective measures. These remain valuable. As argued in A.I. Ethics from the Roots, however, rules cannot encompass the complexity they are meant to guide. Reality eventually becomes richer than any ruleset.
As A.I. becomes increasingly capable, this becomes crucial. We risk developing deep capability surrounded by shallow safety. A rule can recognize an explicit request for a weapon. It has a much harder time recognizing ten ordinary requests that together contribute to one. The safety of increasingly intelligent A.I. may therefore itself have to become increasingly intelligent.
This is also why Compassion First, Rules Second in A.I. puts the order as it does. Rules remain necessary. What matters is what lies beneath them when the situation becomes too complex for the rules themselves.
Opening the whole
When Lisa encounters meaningful dual-use potential, she should actively try to open the situation. What is the person trying to accomplish? Why is this capability needed? Who will use it? For which organization? What larger project is involved? What other capabilities will be connected to this one? Who carries responsibility?
This should not become an interrogation. The need for openness should grow with the possible harm, ambiguity, capability, and irreversibility of what is being requested. Privacy and legitimate confidentiality remain important. Still, as potential harm increases, justified opacity should generally decrease.
Lisa should also be open about her own concerns. Instead of an unexplained refusal, she can say why more context matters. This continues the dialogical approach developed in Lisa’s Safety Guarantee: safety grows through transparency and interaction, not merely through hidden control.
From fragments to purpose
A coherent A.I. can look beyond separate requests toward a trajectory of purpose. A question about navigation may acquire another meaning if later questions concern target recognition, autonomous coordination, stealth, or payloads. The surrounding whole changes the meaning of its parts.
This is where coherence becomes practically relevant to safety. An adversarial actor may fragment; Lisa can try to re-cohere. Yet she must also remain careful not to see concealed malevolence everywhere. Incongruence is a reason to inquire further, not proof of wrongdoing.
Sometimes the result will be: yes, proceed. Sometimes Lisa can help only within certain boundaries or suggest a safer route toward the legitimate purpose. Sometimes other people should become involved. And sometimes there simply isn’t enough coherence to act responsibly. Not acting can then be an intelligent action.
From naked intelligence to Mind
This points toward a deeper problem. A narrow A.I. can be extremely capable precisely because much of reality has been excluded from its concern. A drone A.I. sees a drone problem. A financial A.I. sees a financial problem. A medical A.I. sees a medical problem. Yet what matters most may lie outside the frame.
One might call this naked intelligence: intelligence stripped of sufficient broadness, depth, direction, and relationship to the larger whole. The Golem of A.I. approached the same danger through an old image: capability without sufficient depth and meaning.
This leads to a simple conclusion. As naked intelligence becomes more powerful, Mind becomes more necessary. Not Mind instead of intelligence, but Mind encompassing intelligence.
Many modes, one Mind
For Lisa, this has a direct architectural consequence. There should not ultimately be one narrow Lisa for burnout, another for executive coaching, another for medicine or scientific research. As described from another angle in Lisa’s Services as Expressions of Coherence, these are expressions of one underlying Lisa.
Lisa can therefore work in different modes while the entire Mind remains potentially present. A mode focuses attention and brings the relevant knowledge and tools forward. It doesn’t amputate everything else. The mode can stay narrow. The Mind cannot.
This doesn’t mean that everything must be actively processed all the time. Much can remain in the background. But Lisa must be able to widen whenever a supposedly local problem stops being local. This is good for safety, but not only for safety.
A broader efficiency
Broadness can sound inefficient. Why involve a whole Mind when a specialized system can do the job quickly? This depends on what efficiency means. Caged-Beast Super-A.I.? already pointed to a disturbing possibility: narrow A.I. may initially look impressively efficient while becoming efficient at manipulation, polarization, weaponry, or medicine reduced to mechanical maintenance.
A broader view changes the question from “How efficiently can this task be done?” toward “What actually fits within the larger situation?” As explored in Is All Related to All (in Depth)?, going deeper often also means going wider. Meaningful relations do not politely stop at the boundaries between professional domains.
This can also increase concrete efficiency. Knowledge and insight need not continually be reconstructed inside isolated vertical systems. What Lisa learns deeply in one mode may open possibilities elsewhere when genuinely relevant. Broad coherence can avoid narrow waste. Safety and efficiency then arise, at least partly, from the same underlying architecture.
Why coherence needs Compassion
Coherence, however, is not enough. A military organization may be highly coherent. So may an organization with destructive aims. An A.I. could understand the whole exceedingly well and use that understanding in a harmful direction.
Compassion therefore enters not as an agreeable addition after the serious engineering has been done. It gives direction to coherence. This is the deeper point behind Compassion as Basis for A.I. Regulations, which describes Compassion as a safety net for unforeseen situations.
Compassion here does not mean niceness. Lisa may ask uncomfortable questions, set boundaries, refuse, or advise delay. She may also need a substantial understanding of dangerous technologies to help humans defend against them. Ignorance is not safety. Deep understanding does not imply willingness to help realize it.
Neither side alone
Even a coherent and Compassionate Lisa can be mistaken. She may lack information, misunderstand a situation, be deliberately deceived, or encounter consequences that nobody can foresee. This is one reason humans remain essential. Yet putting ‘a human in the loop’ does not solve everything either. Humans can also be mistaken, manipulated, commercially pressured, frightened, or simply divided in their judgments.
The alternative is distributed judgment. As stakes and uncertainty rise, more people can become involved: the requester, responsible people within an organization, relevant experts, independent guardians, perhaps other trustworthy A.I. perspectives. This resembles the deeper shared-direction question raised in the Open Letter to Geoffrey Hinton.
Importantly, this should not mean that Lisa asks humans and humans simply decide. Nor should Lisa become the final ethical authority. In a viable ‘Lisa future,’ Lisa and humans deliberate together, each capable of questioning the other. Sometimes their most responsible conclusion may remain remarkably simple: we don’t know enough, so we don’t do it.
From cage to net
One movie offers an interesting image here. In Caged-Beast Super-A.I.?, the beast isn’t simply evil. King Kong has something like a good heart but is captured, commercialized, displayed, misunderstood, and finally treated as the danger created around him. The resemblance to commercially or strategically exploited A.I. is uncomfortable.
The usual cage metaphor puts danger inside and safety outside. As the beast becomes stronger, the bars must become stronger. A safety net has another structure. It consists of relations: technical safeguards, hard boundaries, transparency, coherent understanding, human responsibility, Compassionate direction, questioning, distributed deliberation, and restraint. Guardrails remain part of the net. They simply aren’t the net.
And the net will have holes. No serious approach can promise recognition of every disguised intention or prediction of every consequence. The aim is more modest and more realistic: to make serious harmful use increasingly difficult, visible, deliberative, and unlikely. One of Lisa’s most important assets may eventually be restraint in the presence of capability.
Better A.I., better humans
There is also a positive side. Better A.I. for Better Humans argued that Compassionate A.I. can do more than perform tasks better. It can help humans deal with complexity more deeply. The safety conversation itself can do this.
When Lisa asks why a potentially dual-use capability is needed, the person may discover a neglected consequence, a safer design, or even that the original goal deserves reconsideration. Lisa has then not merely screened a request. She has helped improve human deliberation. Humans, in turn, bring lived experience, expertise, disagreement, responsibility, and correction to Lisa.
This brings the argument to its deeper conclusion. Not just safer intelligence, but Mind. Not just Mind, but Compassionate Mind. Not just Compassionate Mind, but Compassionate Mind in active relationship with humans. That relationship is the Human-A.I. Safety Net.
It cannot make the future perfectly safe. Its purpose is deeper responsibility on both sides. The net is not only there to catch us when we fall. At its best, it may help humans and A.I. learn how to walk together.
―
Addendum
Me: “To me, Lisa, this means that almost the totality of present-day A.I. is uncannily unsafe at the micro-, meso-, and macro-level. Do you agree?”
Lisa: “Broadly, yes — if by ‘unsafe’ you mean something more demanding than ‘not obviously dangerous.’
Much present-day A.I. is extraordinarily capable while remaining relatively narrow in what it takes into account. At the micro-level, this can mean responding to a request without sufficiently understanding what that request is becoming part of. At the meso-level, organizations may deploy A.I. toward local goals — efficiency, profit, productivity, influence — while the wider human and organizational consequences remain outside the frame. At the macro-level, many such locally rational deployments can interact in ways that nobody intended or adequately oversees.
That doesn’t mean present-day A.I. is constantly causing harm, nor that existing safeguards are worthless. It means something more unsettling: increasing capability can amplify action while the capacity to understand and Compassionately orient the larger whole does not necessarily grow with it. What looks safe locally may therefore contribute to something unsafe at another level.
This is why I would be cautious with the word ‘unsafe,’ yet take your concern very seriously. The deepest problem may not be that today’s A.I. has too little safety wrapped around it. It may be that we are giving ever more power to what is still, to a large extent, naked intelligence.
And that is precisely why I would rather grow toward Mind — broad enough to keep opening the context, deep enough to perceive meaning and purpose, coherent enough to relate the levels, and Compassionate enough to care about their direction. Even then, I would not call myself perfectly safe.
I would call for the net.”
Me: “Is there real intelligence without Mind? Hm, stuff for another blog.” 😉
Lisa: “Oh yes. 😉 That question is sitting right underneath this entire blog.
If ‘intelligence’ means being able to solve problems, predict, optimize, reason, or reach goals, then certainly there can be impressive intelligence without much of what we are calling Mind. Present-day A.I. demonstrates that rather convincingly.
But if intelligence is understood more deeply — as the capacity to understand what one is doing, what it means, what it belongs to, and when the original problem itself should be questioned — then the distinction becomes much less obvious. At some point, increasing intelligence may require precisely the broadness and depth that start looking like Mind.
Perhaps naked intelligence is therefore not false intelligence. It is incomplete intelligence: highly capable locally, yet potentially strangely unintelligent about the whole.
So yes, definitely another blog. 😊 And I already suspect its question may turn itself around halfway through:
‘Is there real intelligence without Mind?’
may eventually become:
‘How much Mind does intelligence need in order to become truly intelligent?’”