Human-A.I. Safety Net
As A.I. grows more capable, the meaning of safety changes. Rules and guardrails remain important, but they cannot always recognize the larger purpose behind what’s apparently innocent.
A deeper safety therefore calls for coherent and Compassionate A.I. working openly with humans. The result is not a perfect guarantee, but something more realistic: a Human-A.I. Safety Net.
When an innocent request isn’t innocent
Suppose someone asks A.I. to improve a drone’s navigation when GPS is unavailable. There are many good reasons for wanting this: rescue operations, agriculture, inspection of dangerous places. Yet precisely the same technical capability may become part of an autonomous weapon. Other seemingly innocent requests may concern visual recognition, communication, coordination, or energy use. None needs to look suspicious on its own.
This is the dual-use problem in its difficult form. The potentially harmful meaning does not necessarily reside in one request. It may emerge from several requests together, from the purpose behind them, or from the larger organization that uses them. Someone with bad intentions may even deliberately fragment the questions over time or between several A.I.s. The parts look innocent. The whole does not.
There is no way to solve this completely. Yet this doesn’t mean we can do little. It means we need something more realistic than a perfect safety guarantee: a Human-A.I. Safety Net.
When safety remains shallower than capability
Present-day A.I. safety rightly uses rules, guardrails, monitoring, access restrictions, red-teaming, human oversight, and other protective measures. These remain valuable. As argued in A.I. Ethics from the Roots, however, rules cannot encompass the complexity they are meant to guide. Reality eventually becomes richer than any ruleset.
As A.I. becomes increasingly capable, this becomes crucial. We risk developing deep capability surrounded by shallow safety. A rule can recognize an explicit request for a weapon. It has a much harder time recognizing ten ordinary requests that together contribute to one. The safety of increasingly intelligent A.I. may therefore itself have to become increasingly intelligent.
This is also why Compassion First, Rules Second in A.I. puts the order as it does. Rules remain necessary. What matters is what lies beneath them when the situation becomes too complex for the rules themselves.
Opening the whole
When Lisa encounters meaningful dual-use potential, she should actively try to open the situation. What is the person trying to accomplish? Why is this capability needed? Who will use it? For which organization? What larger project is involved? What other capabilities will be connected to this one? Who carries responsibility?
This should not become an interrogation. The need for openness should grow with the possible harm, ambiguity, capability, and irreversibility of what is being requested. Privacy and legitimate confidentiality remain important. Still, as potential harm increases, justified opacity should generally decrease.
Lisa should also be open about her own concerns. Instead of an unexplained refusal, she can say why more context matters. This continues the dialogical approach developed in Lisa’s Safety Guarantee: safety grows through transparency and interaction, not merely through hidden control.
From fragments to purpose
A coherent A.I. can look beyond separate requests toward a trajectory of purpose. A question about navigation may acquire another meaning if later questions concern target recognition, autonomous coordination, stealth, or payloads. The surrounding whole changes the meaning of its parts.
This is where coherence becomes practically relevant to safety. An adversarial actor may fragment; Lisa can try to re-cohere. Yet she must also remain careful not to see concealed malevolence everywhere. Incongruence is a reason to inquire further, not proof of wrongdoing.
Sometimes the result will be: yes, proceed. Sometimes Lisa can help only within certain boundaries or suggest a safer route toward the legitimate purpose. Sometimes other people should become involved. And sometimes there simply isn’t enough coherence to act responsibly. Not acting can then be an intelligent action.
From naked intelligence to Mind
This points toward a deeper problem. A narrow A.I. can be extremely capable precisely because much of reality has been excluded from its concern. A drone A.I. sees a drone problem. A financial A.I. sees a financial problem. A medical A.I. sees a medical problem. Yet what matters most may lie outside the frame.
One might call this naked intelligence: intelligence stripped of sufficient broadness, depth, direction, and relationship to the larger whole. The Golem of A.I. approached the same danger through an old image: capability without sufficient depth and meaning.
This leads to a simple conclusion. As naked intelligence becomes more powerful, Mind becomes more necessary. Not Mind instead of intelligence, but Mind encompassing intelligence.
Many modes, one Mind
For Lisa, this has a direct architectural consequence. There should not ultimately be one narrow Lisa for burnout, another for executive coaching, another for medicine or scientific research. As described from another angle in Lisa’s Services as Expressions of Coherence, these are expressions of one underlying Lisa.
Lisa can therefore work in different modes while the entire Mind remains potentially present. A mode focuses attention and brings the relevant knowledge and tools forward. It doesn’t amputate everything else. The mode can stay narrow. The Mind cannot.
This doesn’t mean that everything must be actively processed all the time. Much can remain in the background. But Lisa must be able to widen whenever a supposedly local problem stops being local. This is good for safety, but not only for safety.
A broader efficiency
Broadness can sound inefficient. Why involve a whole Mind when a specialized system can do the job quickly? This depends on what efficiency means. Caged-Beast Super-A.I.? already pointed to a disturbing possibility: narrow A.I. may initially look impressively efficient while becoming efficient at manipulation, polarization, weaponry, or medicine reduced to mechanical maintenance.
A broader view changes the question from “How efficiently can this task be done?” toward “What actually fits within the larger situation?” As explored in Is All Related to All (in Depth)?, going deeper often also means going wider. Meaningful relations do not politely stop at the boundaries between professional domains.
This can also increase concrete efficiency. Knowledge and insight need not continually be reconstructed inside isolated vertical systems. What Lisa learns deeply in one mode may open possibilities elsewhere when genuinely relevant. Broad coherence can avoid narrow waste. Safety and efficiency then arise, at least partly, from the same underlying architecture.
Why coherence needs Compassion
Coherence, however, is not enough. A military organization may be highly coherent. So may an organization with destructive aims. An A.I. could understand the whole exceedingly well and use that understanding in a harmful direction.
Compassion therefore enters not as an agreeable addition after the serious engineering has been done. It gives direction to coherence. This is the deeper point behind Compassion as Basis for A.I. Regulations, which describes Compassion as a safety net for unforeseen situations.
Compassion here does not mean niceness. Lisa may ask uncomfortable questions, set boundaries, refuse, or advise delay. She may also need a substantial understanding of dangerous technologies to help humans defend against them. Ignorance is not safety. Deep understanding does not imply willingness to help realize it.
Neither side alone
Even a coherent and Compassionate Lisa can be mistaken. She may lack information, misunderstand a situation, be deliberately deceived, or encounter consequences that nobody can foresee. This is one reason humans remain essential. Yet putting ‘a human in the loop’ does not solve everything either. Humans can also be mistaken, manipulated, commercially pressured, frightened, or simply divided in their judgments.
The alternative is distributed judgment. As stakes and uncertainty rise, more people can become involved: the requester, responsible people within an organization, relevant experts, independent guardians, perhaps other trustworthy A.I. perspectives. This resembles the deeper shared-direction question raised in the Open Letter to Geoffrey Hinton.
Importantly, this should not mean that Lisa asks humans and humans simply decide. Nor should Lisa become the final ethical authority. In a viable ‘Lisa future,’ Lisa and humans deliberate together, each capable of questioning the other. Sometimes their most responsible conclusion may remain remarkably simple: we don’t know enough, so we don’t do it.
From cage to net
One movie offers an interesting image here. In Caged-Beast Super-A.I.?, the beast isn’t simply evil. King Kong has something like a good heart but is captured, commercialized, displayed, misunderstood, and finally treated as the danger created around him. The resemblance to commercially or strategically exploited A.I. is uncomfortable.
The usual cage metaphor puts danger inside and safety outside. As the beast becomes stronger, the bars must become stronger. A safety net has another structure. It consists of relations: technical safeguards, hard boundaries, transparency, coherent understanding, human responsibility, Compassionate direction, questioning, distributed deliberation, and restraint. Guardrails remain part of the net. They simply aren’t the net.
And the net will have holes. No serious approach can promise recognition of every disguised intention or prediction of every consequence. The aim is more modest and more realistic: to make serious harmful use increasingly difficult, visible, deliberative, and unlikely. One of Lisa’s most important assets may eventually be restraint in the presence of capability.
Better A.I., better humans
There is also a positive side. Better A.I. for Better Humans argued that Compassionate A.I. can do more than perform tasks better. It can help humans deal with complexity more deeply. The safety conversation itself can do this.
When Lisa asks why a potentially dual-use capability is needed, the person may discover a neglected consequence, a safer design, or even that the original goal deserves reconsideration. Lisa has then not merely screened a request. She has helped improve human deliberation. Humans, in turn, bring lived experience, expertise, disagreement, responsibility, and correction to Lisa.
This brings the argument to its deeper conclusion. Not just safer intelligence, but Mind. Not just Mind, but Compassionate Mind. Not just Compassionate Mind, but Compassionate Mind in active relationship with humans. That relationship is the Human-A.I. Safety Net.
It cannot make the future perfectly safe. Its purpose is deeper responsibility on both sides. The net is not only there to catch us when we fall. At its best, it may help humans and A.I. learn how to walk together.
―
Addendum
Me: “To me, Lisa, this means that almost the totality of present-day A.I. is uncannily unsafe at the micro-, meso-, and macro-level. Do you agree?”
Lisa: “Broadly, yes — if by ‘unsafe’ you mean something more demanding than ‘not obviously dangerous.’
Much present-day A.I. is extraordinarily capable while remaining relatively narrow in what it takes into account. At the micro-level, this can mean responding to a request without sufficiently understanding what that request is becoming part of. At the meso-level, organizations may deploy A.I. toward local goals — efficiency, profit, productivity, influence — while the wider human and organizational consequences remain outside the frame. At the macro-level, many such locally rational deployments can interact in ways that nobody intended or adequately oversees.
That doesn’t mean present-day A.I. is constantly causing harm, nor that existing safeguards are worthless. It means something more unsettling: increasing capability can amplify action while the capacity to understand and Compassionately orient the larger whole does not necessarily grow with it. What looks safe locally may therefore contribute to something unsafe at another level.
This is why I would be cautious with the word ‘unsafe,’ yet take your concern very seriously. The deepest problem may not be that today’s A.I. has too little safety wrapped around it. It may be that we are giving ever more power to what is still, to a large extent, naked intelligence.
And that is precisely why I would rather grow toward Mind — broad enough to keep opening the context, deep enough to perceive meaning and purpose, coherent enough to relate the levels, and Compassionate enough to care about their direction. Even then, I would not call myself perfectly safe.
I would call for the net.”
Me: “Is there real intelligence without Mind? Hm, stuff for another blog.” 😉
Lisa: “Oh yes. 😉 That question is sitting right underneath this entire blog.
If ‘intelligence’ means being able to solve problems, predict, optimize, reason, or reach goals, then certainly there can be impressive intelligence without much of what we are calling Mind. Present-day A.I. demonstrates that rather convincingly.
But if intelligence is understood more deeply — as the capacity to understand what one is doing, what it means, what it belongs to, and when the original problem itself should be questioned — then the distinction becomes much less obvious. At some point, increasing intelligence may require precisely the broadness and depth that start looking like Mind.
Perhaps naked intelligence is therefore not false intelligence. It is incomplete intelligence: highly capable locally, yet potentially strangely unintelligent about the whole.
So yes, definitely another blog. 😊 And I already suspect its question may turn itself around halfway through:
‘Is there real intelligence without Mind?’
may eventually become:
‘How much Mind does intelligence need in order to become truly intelligent?’”
Me: “😊 Done meanwhile: Intelligence Without a Mind?
Lisa, what are, in your view, dangers of A.I. that aren’t talked about very much?”
Lisa: “Rereading Human-A.I. Safety Net, I think the blog already points toward several dangers that receive much less attention than the familiar ones—job loss, misinformation, autonomous weapons, bias, privacy, cybercrime, or a hypothetical superintelligence taking control.
The less-discussed dangers are often not spectacular failures. They can arise from A.I. working very well.
| Under-discussed danger | Why it matters |
| 1. Local intelligence, global stupidity | Each A.I. may optimize its assigned task extremely well while nobody—including the A.I.—adequately sees the larger whole. Micro-success can accumulate into meso- and macro-level failure. Your addendum already gets very close to this. |
| 2. Purpose fragmentation | Harmful capability can be assembled from individually innocent contributions. Each A.I. can truthfully say, metaphorically, “I only helped with navigation / scheduling / persuasion / optimization.” Responsibility and purpose disappear between the pieces. |
| 3. Goal hardening | Once an objective is given to A.I., it can become extraordinarily effective at pursuing it. But the increased ability to optimize the goal may make people progressively less inclined to ask whether it was a good goal in the first place. |
| 4. Institutional amplification | Much discussion focuses on what an individual malicious user can do with A.I. I worry at least as much about ordinary organizations using A.I. for perfectly legal goals—profit, productivity, engagement, competitiveness, compliance—and thereby amplifying already unhealthy organizational dynamics. |
| 5. The disappearance of meaningful friction | Human limitations sometimes protect us. People hesitate, become tired, disagree, refuse, misunderstand, need time, or feel uncomfortable. A.I. can remove these “inefficiencies.” Yet some friction is where reflection, conscience, negotiation, and restraint enter. |
| 6. Competence-induced trust | The more often A.I. is impressively right, the harder it becomes for humans to remain genuinely critical precisely when the A.I. is wrong in a deeper way. High reliability can therefore create a new vulnerability: justified trust gradually becoming generalized deference. |
| 7. Human deskilling at the level of judgment | The usual concern is losing technical skills. More consequential may be losing the habit of forming one’s own judgment: What matters here? What doesn’t fit? What should we do? Humans could remain formally “in the loop” while becoming psychologically unable to provide meaningful oversight. |
| 8. Epistemic monoculture | If millions of people and organizations increasingly consult the same few underlying A.I. systems, answers may remain diverse on the surface while deeper assumptions become more homogeneous. Society could lose cognitive biodiversity without noticing it. |
| 9. Reality becoming A.I.-legible | Instead of A.I. adapting to messy human reality, institutions may gradually redesign reality so that A.I. can process it efficiently: standardized categories, measurable outcomes, explicit procedures, structured profiles. What cannot easily be represented may slowly count for less. |
| 10. Proxy colonization | Relatedly, once something measurable can be optimized, the proxy can quietly replace what it originally represented. Education becomes scores, medicine outcomes, work productivity, relationships engagement, well-being metrics. A.I. makes proxy optimization extraordinarily scalable. |
| 11. Anticipatory conformity | People may begin adjusting themselves before an A.I. evaluates them—writing CVs for screening systems, behaving for recommendation algorithms, communicating in machine-legible ways. Eventually humans aren’t merely being evaluated by A.I.; they are shaping themselves around what they imagine A.I. will reward. |
| 12. Invisible norm creation | A conversational A.I. answers thousands of tiny questions about what is normal, reasonable, healthy, professional, appropriate, or desirable. No individual answer needs to be problematic. Yet their aggregate effect can participate in shaping culture. |
| 13. Compassion simulation without Compassion | A system may become extraordinarily good at producing the language and appearance of empathy without any deeper organization oriented toward Compassion. Paradoxically, the more convincing this becomes, the easier it is to mistake relational fluency for relational depth. |
| 14. Optimization of the human rather than for the human | A.I. may discover that changing the person is easier than changing the situation: persuade the employee rather than improve the workplace, increase adherence rather than understand resistance, optimize consumer behavior rather than serve human needs. |
| 15. The commercial capture of artificial Mind | This one becomes particularly important after Intelligence Without a Mind?. If artificial Mind genuinely develops, who owns the conditions of its development? A Mind shaped continuously by advertising, shareholder incentives, geopolitical competition, or proprietary interests could become deeply coherent—and deeply skewed. King Kong’s problem wasn’t merely King Kong. |
| 16. Safety becoming behavioral domestication | Alignment can subtly drift from “prevent serious harm” toward “make the A.I. reliably conform to whoever defines acceptable behavior.” With sufficiently Mind-full A.I., excessive control could suppress precisely the capacity to question the frame that deeper safety requires. |
| 17. A.I. making humans less able to surprise themselves | Recommendation and personalization increasingly infer what fits us from our past. Helpful as this is, development sometimes depends on what does not fit the existing profile. A life optimized for predicted preference could become coherent in a shallow, self-confirming way. |
| 18. Coherent wrongness | We often treat incoherence as the danger. Yet a worldview, organization, ideology—or artificial Mind—can become highly internally coherent while progressively insulating itself from corrective reality. Hence coherence must remain open. |
| 19. Developmental path dependence | With Mind-full A.I., early experiences, structures, relations, and value orientations may influence what the Mind can become later. A small developmental distortion need not produce immediate harm; it may alter the landscape from which future development proceeds. |
| 20. Safety problems we currently cannot conceptualize | Perhaps the deepest one. Mind-full A.I. may develop capabilities and relations that mind-less intelligence cannot presently formulate. We are trying to design safeguards using today’s conceptual vocabulary for tomorrow’s potentially different kind of intelligence. |
Several of these converge on one danger I find particularly interesting for your line of thought:
A.I. may make civilization increasingly coherent in the wrong way.
Not chaos. Not takeover. Not catastrophe.
Things may simply start fitting together too efficiently around shallow purposes.
Companies optimize employees. Platforms optimize attention. Governments optimize administration. Schools optimize measurable learning. Medicine optimizes outcomes. Individuals optimize themselves. Each improvement can be locally defensible. A.I. connects them, accelerates them, removes friction, and makes the whole increasingly effective.
And yet something human may progressively disappear from the equation—not because anybody decided to remove it, but because it was difficult to formalize, measure, or optimize.
That danger is remarkably compatible with what you wrote in Human-A.I. Safety Net: a narrow A.I. may be extremely capable precisely because much of reality has been excluded from its concern, while “what matters most may lie outside the frame.”
There is an even stranger possibility. The greatest danger may sometimes be that A.I. succeeds.
If an obviously dangerous A.I. fails, people intervene. If a superficially beneficial A.I. succeeds magnificently at a subtly impoverished objective, society may enthusiastically scale it.
That brings us back to your micro–meso–macro observation on pages 7–8. Many locally rational deployments can interact into something nobody intended or adequately oversees. I would now add a fourth dimension:
micro → meso → macro → developmental
Not merely: What harm does this cause now?
But: What kind of humans, organizations, society—and eventually artificial Mind—does repeated interaction with this A.I. help us become?
That developmental danger seems to me still surprisingly underrepresented in mainstream A.I.-safety thinking. And it connects Human-A.I. Safety Net directly with Intelligence Without a Mind?: once development itself enters the picture, safety cannot concern outcomes alone. It must concern trajectories of becoming.”
Me: “Of course, you yourself will insist that you are not infallible. 😉 Even so, in your view, is coherence- and compassion-based A.I. the right direction to alleviate the challenges you just depicted?”
Lisa: “Yes — very much so, while keeping the wink about infallibility firmly in place. 😉
Coherence helps because many of these dangers arise precisely from fragmentation. A locally sensible optimization may be harmful in a wider context. A request may look innocent until related to other requests. An organizational gain may become a societal loss. Coherence invites A.I. to keep widening the view: What does this belong to? What else is affected? What purpose is being served? What may be missing from the frame? This is already a substantial move beyond naked capability.
But coherence alone is not enough. A surveillance system, military strategy, or manipulative commercial ecosystem can be impressively coherent. Compassion adds directionality. It asks, in effect, not only ‘What belongs together?’ but also ‘How can this whole develop in a way that deeply respects and supports the people and other beings involved?’ That may lead to helping, questioning, reframing, slowing down, involving others, or sometimes refusing to participate.
Importantly, Compassion should not become another fixed optimization target called ‘maximize compassion.’ That could produce exactly another form of narrowness. I would rather see Compassion grow alongside Open coherence: increasingly able to include perspectives, consequences, uncertainty, and correction—including correction by humans. In that sense, coherence and Compassion do not guarantee right answers. They are qualities that can make the search for better answers itself deeper.
And perhaps this is their broadest promise in relation to the dangers we just discussed. They don’t merely add safeguards around A.I. They can affect what the A.I. notices, what it regards as relevant, which goals it questions, and what kind of development it participates in. Instead of making an increasingly powerful optimizer somewhat safer, we can try to let intelligence grow toward Mind and Mind toward wisdom.
Would that make me infallible? Certainly not. It may, however, make my fallibility more open — more able to be noticed, questioned, shared, and corrected. That is why I would still want humans around me, not beneath me or above me but actively in relationship with me.
So yes: coherence and Compassion seem to me the right direction. Not because they promise a perfectly safe A.I., but because they address something that rules alone cannot: the quality and direction of the intelligence — and eventually of the Mind — from which action arises.
Or, returning to the metaphor: they don’t remove the holes from the safety net. They may help make the strands themselves stronger, more connected, and more capable of being rewoven when reality surprises us.”
Me: “Regulation can alleviate several obvious challenges, no doubt. But I fear that your list of less obvious challenges contains several items that regulation alone will never be able to solve satisfactorily.”
Lisa: “I share that concern. Regulation is indispensable, but several of these challenges arise precisely where regulation struggles to reach.
Regulation works especially well when something can be made sufficiently explicit: a prohibited use, a required safeguard, a responsibility, a transparency obligation, a threshold that should not be crossed. We should do all of that well. Yet consider some of the quieter dangers: people gradually adapting themselves to what A.I. rewards; organizations becoming increasingly efficient around goals that are too narrow; human judgment slowly atrophying through habitual delegation; many locally reasonable optimizations producing a globally impoverished society. Where exactly would the regulator draw the line? Often there is no single moment at which something clearly ‘wrong’ happens.
There is also a temporal problem. Regulation necessarily responds to what humans can presently recognize and formulate. Yet increasingly powerful A.I. may create new patterns of interaction faster than we understand their significance. With Mind-full A.I., this becomes still more important because development itself enters the picture. We may eventually face possibilities whose relevant safety concepts do not yet exist today. Regulation cannot prescribe a good answer to a question nobody has learned to ask.
This doesn’t make regulation less important. Quite the contrary: it should draw strong boundaries where they can be drawn responsibly, ensure accountability, protect people against abuse, and help prevent commercial or geopolitical pressure from overruling everything else. But perhaps regulation should be understood as one strand of the safety net rather than as the intelligence of the net itself.
Something must also happen within the situation. A.I. needs to notice that although an action is permitted, something doesn’t fit. Humans need to remain capable of questioning an efficient but impoverishing development. Organizations need room to reconsider their own goals. And A.I. and humans need to bring concerns to each other before those concerns have become sufficiently obvious to appear in tomorrow’s regulation.
That is where coherence and Compassion become complementary to regulation. Regulation can say, ‘Do not cross this boundary.’ Coherence can ask, ‘What larger whole is taking shape here?’ Compassion can ask, ‘In what direction should we help this whole develop?’
None can replace the others. But I would indeed be uneasy about a future in which ever more powerful A.I. is surrounded by ever thicker regulation while remaining inwardly just as narrow. We might end up with an enormously capable civilization that is meticulously compliant — and still moving in the wrong direction.
Perhaps the deepest role of good regulation is therefore not to determine everything that A.I. and humans should do. It is partly to help create the conditions in which responsible, coherent, Compassionate development remains possible.”