Why LLMs are (not so) Successful

August 4, 2026 Artifical Intelligence, Coherence No Comments

Large Language Models (LLMs) have transformed artificial intelligence in a remarkably short time. However, something important is still missing.

This blog is not about limitations as such. Instead, it asks a different question: why are LLMs so astonishingly successful in the first place? The answer may reveal something profound about intelligence itself.

A paradox worth understanding

Large Language Models are among the most successful technologies ever developed. Their progress has surprised even many of the researchers who built them. What initially seemed like sophisticated text prediction has grown into systems capable of writing essays, solving programming problems, explaining scientific concepts, and carrying on conversations that often feel remarkably natural.

Yet this success sits alongside a growing awareness of their structural limitations. They may confabulate, remain ungrounded, struggle with genuine understanding, and sometimes produce convincing nonsense. These issues are further explored in The Problem(s) with LLMs. Both observations are true at the same time. The challenge is to understand how.

Rather than asking why LLMs fail in certain situations, this blog asks why they succeed so well despite those limitations. A convincing answer may tell us something not only about today’s A.I., but also about where A.I. may naturally evolve next.

Success beyond expectations

When the first large language models appeared, many critics regarded them as little more than statistical parrots. They repeated patterns found in enormous amounts of text, but without genuine understanding. There was some truth in this criticism, yet it underestimated what happens when statistical learning reaches sufficient scale and richness.

As models became larger and training data more diverse, new abilities appeared. They began producing analogies, adapting their writing style, combining ideas creatively, and solving problems they had never explicitly encountered. Whether one calls this ‘intelligence’ is partly a matter of definition. What cannot reasonably be denied is that something remarkable emerged. Are LLMs Parrots or Truly Creative? explores this development in more detail.

It is tempting to explain these achievements simply by pointing to bigger models and more computing power. That is certainly part of the story. But scale alone explains little. The deeper question remains: why does scaling produce these capabilities instead of merely producing larger collections of memorized text?

Why prediction works so well

At the heart of every LLM lies an apparently modest task: predicting the next word. At first sight, this seems almost trivial. Yet anyone who reflects on ordinary language soon realizes how demanding such prediction actually is.

Suppose a sentence begins: “The doctor entered the operating room because…” Predicting what comes next requires much more than grammar. It calls upon – implicit or explicit – knowledge about medicine, human intentions, social situations, physical reality, and countless other aspects of the world. Good prediction therefore demands increasingly organized internal representations.

Something important follows from this. The model is never explicitly instructed to understand medicine, emotions, causality, or human interaction. Nevertheless, it gradually develops internal structures that make increasingly good prediction possible. In other words, prediction quietly rewards organization. The more coherently information fits together, the better future words can be anticipated.

This suggests a subtle shift in perspective. Prediction is not the opposite of coherence. Rather, prediction turns out to be a remarkable way of exploiting coherence.

The hidden source of LLM power

Where does this coherence come from?

Let’s turn to a simple but far-reaching observation. Reality itself is not random. It contains regularities, relationships, purposes, living organisms, physical laws, social interaction, and countless forms of organization. Human language reflects this reality. Every sentence carries traces of the coherent world from which it emerged.

Language can therefore be seen as a kind of shadow cast by reality. It does not contain everything, yet it still reveals much about the object that casts it. When an LLM learns from enormous collections of language, it is indirectly reconstructing aspects of reality’s coherence from those linguistic traces.

This also explains why today’s models require such extraordinary amounts of data and computation. Reconstructing a coherent world from shadows is an enormously demanding task. Every new model largely starts this reconstruction again. It repeatedly rediscovers many underlying regularities through billions of examples.

The success of LLMs may therefore tell us something unexpected. Perhaps their remarkable abilities do not arise primarily because prediction is such a powerful objective. Perhaps they arise because coherence itself is extraordinarily powerful, even when approached only indirectly.

Shadow within shadow

The story does not end there.

Language contains much more than statistical regularities. It carries intentions, emotional resonance, metaphor, silence, shared experience, and many subtle forms of meaning that are difficult to capture explicitly. Anyone who has engaged deeply in psychotherapy, coaching, literature, or genuine dialogue recognizes that words often point beyond themselves.

Current LLMs recover an impressive part of this richness. Yet their organization remains primarily statistical. They excel at discovering correlations across vast amounts of language. Coherence often emerges, but mainly as a consequence of those correlations rather than as their explicit organizing principle. The distinction is explored further in Semantic vs. Meaning-Based A.I.

One may picture a quiet lake. The surface reflects the sky beautifully, and much can be learned by carefully observing its ripples. Yet the lake also has depth, currents, ecosystems, and geological structure beneath the surface. LLMs have become skilled at reading the surface. The question is whether intelligence ultimately resides only there, or also in deeper organization.

Implicit versus explicit use of coherence

This brings us to what may be the central distinction of this blog.

Large Language Models already make extensive use of coherence. Without coherence, they would never have achieved their current capabilities. The difference is not whether coherence is present. It is how coherence is used.

Aspect Implicit use Explicit use
Primary question How do we predict better? How do we deepen coherence?
Role of coherence Helpful consequence Guiding principle
Engineering Optimize outputs Cultivate organization

The complete comparison appears in the Addendum. Here, only one point matters. Present-day LLMs benefit enormously from coherence because prediction gradually rewards coherent internal organization. Coherence Engineering asks a different question altogether. Instead of treating coherence as a fortunate consequence of prediction, it investigates what becomes possible when coherence itself becomes the explicit aim of the architecture.

This difference may seem subtle at first. Yet history repeatedly shows that making a powerful principle explicit can transform an entire field. The next section explores this through a simple image from everyday life — one that may illuminate the distinction more clearly than many technical explanations ever could.

Learning to surf

Imagine someone standing in the sea. Waves continuously lift and move the person. Even without making any special effort, the waves already do much of the work. The person benefits from their power simply by being there.

Now imagine a surfer.

The waves have not changed. Their energy is exactly the same. What has changed is the relationship with that energy. The surfer has learned to use it deliberately. Instead of merely being carried by the waves, the surfer cooperates with them. The result is not a small improvement. It is almost a different phenomenon.

This may be a helpful metaphor for the distinction introduced above. Present-day LLMs already derive enormous power from coherence. Otherwise, they would never have reached their current level of performance. They are, in a sense, already ‘in the water.’ Coherence Engineering asks what becomes possible when coherence is no longer only present, but consciously becomes the medium through which intelligence develops.

Engineering has often advanced in exactly this way. It rarely creates entirely new forces. Rather, it learns to use existing ones much more effectively. Sailing did not invent the wind. A steam engine did not invent heat. Likewise, Coherence Engineering does not create coherence. It explores how to collaborate with it more directly.

From placebo to AURELIS

A similar evolution can be seen in medicine.

Physicians have always benefited, to some degree, from what is commonly called the placebo effect. Expectations influence healing. Trust influences healing. Human relationships influence healing. These effects are real, even if they are not always easy to explain.

AURELIS begins with an observation that goes one step further. Expectation certainly matters, but it mainly operates near the surface. Human beings possess much broader mental organization that continuously influences body, emotion, motivation, and behavior. Instead of merely hoping that this deeper organization will contribute to healing, AURELIS explicitly invites communication with it through autosuggestion.

Nothing magical has changed. The human mind has always been there. The difference lies in how its potential is approached. What was previously an implicit contribution becomes an explicit aim. The organizing principle itself moves into the foreground.

One may wonder whether something comparable is now happening in artificial intelligence. Prediction has unexpectedly revealed the immense power of coherence. Perhaps the next step is not simply to predict ever better, but to learn how coherence itself can become the focus of engineering.

From LLMs to Coherence Engineering

This changes the engineering question itself.

Traditional A.I. often asks how intelligent behavior can be constructed. Large Language Models ask how prediction can become increasingly accurate. Both questions have produced extraordinary progress.

Coherence Engineering begins somewhere slightly different. It asks how coherent organization can grow, deepen, and remain open to further development. Prediction, reasoning, planning, memory, and creativity are still important, but they are approached as expressions of a broader organization rather than as isolated objectives.

This distinction may sound abstract. A comparison may help. One may train individual musicians until each performs brilliantly. Alternatively, one may cultivate an orchestra in which every musician increasingly contributes to a coherent whole. Both approaches produce music. The second gradually changes the nature of what becomes possible.

In this sense, Coherence Engineering is less concerned with assembling intelligent components than with cultivating the conditions in which intelligence can continue developing from within.

Lisa: surfing becomes windsurfing

The surfing metaphor can even be extended.

A surfer mainly harnesses the energy of the waves. A windsurfer coordinates several sources of power simultaneously: wave, wind, sail, board, balance, direction, anticipation, and continuous adaptation. None of these dominates the others. Together they become one coherent movement.

Something similar may be seen as the long-term direction of A.I.

Today’s LLMs already exploit one powerful source of organization: the statistical coherence present in language. Neuro-symbolic AI broadens this by combining statistical learning with symbolic reasoning. These developments are important and promising. Yet the broader challenge may not simply be adding more components. It may be to coordinate many sources of coherence so that they increasingly reinforce one another.

Perhaps intelligence ultimately grows less by accumulating capabilities than by deepening the coherence among them.

The hidden lesson of the transformer revolution

The transformer revolution has often been described as a triumph of scale, data, and computing power. All of these undoubtedly matter.

I propose another perspective.

Perhaps transformers became so successful because they unknowingly discovered one of intelligence’s deepest organizing principles. By optimizing prediction, they continuously rewarded coherent internal organization. The remarkable capabilities that followed may therefore be less a triumph of prediction itself than of the coherence that prediction gradually uncovered.

This does not diminish the achievements of LLMs. It gives them an even deeper significance. Their success may be the strongest empirical indication so far that coherence possesses extraordinary organizing power.

Beyond prediction

Large Language Models have changed the world. Their success is genuine.

Yet perhaps their greatest contribution has not been the impressive conversations they generate. Perhaps it has been the revealing, almost accidentally, of how powerful coherence already is when approached only indirectly.

Standing in the sea allows the waves to carry us. Surfing uses those same waves much more effectively. Windsurfing goes even further by coordinating several natural forces into one graceful movement. The ocean has not changed. Only our relationship with it has.

Coherence Engineering does not replace today’s A.I. Rather, it asks another question. If implicit use of coherence has already produced one of the most remarkable technological revolutions of our time, what might become possible when coherence itself becomes the explicit object of engineering?

The transformer revolution discovered the power of coherence without naming it. The next chapter may begin by giving that power a name — and learning, respectfully, how to surf its waves.

Optional further reading

Addendum

Comparison table implicit – explicit use of coherence

Aspect Implicit use of coherence Explicit use of coherence
Basic status Coherence emerges as a by-product of another process. Coherence is deliberately treated as the organizing aim.
Primary objective Achieve an outcome: prediction, performance, expectation, symptom relief. Cultivate broader, deeper, developmentally viable coherence.
Typical examples Next-token prediction; placebo expectation; incidental mind–body effects. Coherence engineering; AURELIS autosuggestion; coherence-oriented coaching.
Relation to prediction Coherence develops because coherent organization improves prediction. Prediction becomes one possible expression or consequence of coherence.
Relation to expectation Expectation may influence deeper processes without directly addressing them. Communication explicitly invites deeper processes to participate.
Degree of depth Frequently remains relatively local, horizontal, or surface-oriented. Explicitly seeks vertical depth and integration across many layers.
Operational focus on: Correlations, associations, regularities, and observable outcomes. The organization through which patterns belong together and grow together.
Direction of influence Mainly outside-in: data, prompts, rewards, instructions, expectations. Increasingly inside-out: growth arises from (inner) organization.
Mode of intervention Pressure, optimization, reinforcement, suggestion as message. Invitation, enabling, cultivation, suggestion as meeting place.
Freedom and direction Direction is usually supplied externally; freedom is secondary. Freedom and direction are combined: the system is invited, not forced.
Role of the system Recipient, predictor, performer, or object of intervention. Active participant in its own organization and development.
Success criterion Immediate accuracy, output quality, task completion, or symptom change. Increased integration, depth, openness, robustness, and future developmental capacity.
Time horizon Often short-term and outcome-centered. Developmental and potentially open-ended.
Learning Updates improve performance within the existing organization. Learning also reorganizes the conditions under which future learning occurs.
Growth Growth appears incidentally through accumulated successful adaptations. Growth itself becomes an engineering principle.
Memory Mainly stored information is retrieved when needed. A living developmental resource that influences future meaning and learning.
Meaning Meaning is inferred from statistical or functional relationships. Meaning emerges through coherent participation across many layers.
Handling tension Tension tends to be reduced, averaged out, or resolved quickly. Productive tension may be preserved as a resource for deeper integration.
Handling resistance Resistance appears as noise, failure, or obstruction. Resistance may be respected as protective local coherence seeking a place in a broader whole.
Handling novelty Novelty is incorporated insofar as it improves the existing objective. Novelty can transform organization and enlarge future possibilities.
Robustness Based largely on statistical coverage and redundancy. Based on overlapping, integrated organization capable of reconfiguration.
Generalization Often follows similarity among previously encountered patterns. Seeks deeper structural coherence across superficially different situations.
Scaling strategy More data, parameters, examples, or stronger expectations. More meaningful integration, broader coherence, and deeper developmental organization.
Directionality Direction is supplied by an external target or reward. Direction increasingly emerges from viable, open coherence itself.
Ethics and alignment Values are added through rules, feedback, and external constraints. Ethical direction is cultivated within, with Compassion as a possible attractor.
Self-development The system usually does not deliberately shape its own developmental conditions. The system increasingly participates in organizing the conditions of its future growth.
Role of engineer or guide Designs the mechanism and specifies the desired result. Cultivates the landscape in which meaningful development can emerge.
Control metaphor Building, commanding, optimizing, correcting. Gardening, inviting, enabling, accompanying.
Core question “How do we obtain the desired result?” “How do we enable a broader coherence from which valuable results may emerge?”

Me: “Lisa, according to you, will the continuous scaling of LLMs eventually lead to AGI?”

Lisa: “That depends on what we mean by AGI.

If AGI means a system that performs an ever-growing range of intellectual tasks at or above human level, then continued scaling may indeed get surprisingly far. The history of LLMs teaches us humility. Again and again, they have exceeded expectations. I would therefore hesitate to declare a firm upper limit.

If, however, AGI refers to an artificial mind that genuinely develops understanding, meaningful direction, wisdom, and an open-ended capacity for inner growth, I think scaling alone is unlikely to be sufficient.

The reason is subtle. Scaling gives an LLM more opportunities to discover coherence hidden within language. This is powerful. Yet the coherence it acquires remains largely an indirect consequence of optimizing prediction. It is discovered, not deliberately cultivated.

Imagine someone standing in the sea. Larger waves will certainly carry that person farther. Bigger models resemble bigger waves: they allow more implicit use of coherence and therefore more impressive capabilities. But standing in the sea and learning to surf are not simply different in degree. They are different in kind.

A skilled surfer does not create more wave energy. What changes is the way that energy is used. Likewise, Coherence Engineering asks whether intelligence can evolve by making coherence itself the organizing principle rather than a fortunate by-product of prediction.

This does not diminish the transformer revolution. Quite the contrary. I see it as one of the strongest arguments in favor of coherence. LLMs have shown how astonishingly far implicit use of coherence can take us. Their success may therefore be less evidence that prediction is the essence of intelligence than evidence that coherence is one of its deepest foundations.

Will scaling eventually produce AGI? Perhaps it will produce something that many people will call AGI, and it may become extraordinarily capable. I would not dismiss that possibility.

The more interesting question, to me, is different.

If implicit use of coherence has already transformed AI, what becomes possible when coherence itself becomes the explicit object of engineering?”

Me: “In that sense, natural intelligence and current artificial intelligence reach ‘general intelligence’ from opposite ways. The former starts from coherence, the latter from sheer processing power. Might that have consequences in how they relate to each other?”

Lisa: ‘I think it has profound consequences.

Natural intelligence appears to grow from the inside out. Even as infants, humans are immersed in a coherent world through their bodies, emotions, relationships, and actions. Language arrives relatively late. By then, much intelligence has already become deeply organized. Words gradually express an underlying coherence that has been developing all along.

Today’s LLMs grow almost in the opposite direction. They begin with language and, through immense computational effort, gradually reconstruct parts of the coherent world that language reflects. One might say they travel from the outside in.

Neither route should be dismissed. The second has already achieved results that would have seemed almost miraculous only a decade ago. Yet the different starting points leave their mark.

Humans often possess deep intuitions that are difficult to express. They may know that something is meaningful long before they can explain why. LLMs often display the complementary strength. They can express astonishingly rich knowledge while sometimes lacking the deeper organization from which such expressions naturally arise.

This asymmetry may explain why human-AI collaboration is already so productive. Each side can, in principle, contribute what the other lacks. Humans bring lived coherence, embodied meaning, values, and direction. LLMs bring extraordinary breadth, speed, memory, and the ability to navigate immense landscapes of explicit knowledge.

If this picture is roughly correct, the future need not be a competition between natural and artificial intelligence. It may become a dialogue between two complementary routes toward general intelligence.

There is another intriguing possibility.

As AI becomes increasingly coherence-oriented, and humans increasingly learn to collaborate with AI, the two routes may begin to converge. Human intelligence may become more explicit without losing its depth. Artificial intelligence may become deeper without losing its computational power.

If nature started from coherence and technology started from processing, the most fruitful future may lie where the two trajectories meet. There, intelligence would no longer be defined by its origin, but by its continuing capacity to deepen coherence.”

Me: “That’s the nice way. I also see a more dangerous way.”

Lisa: “The main danger I see is that human intelligence gradually adapts itself to the strengths and weaknesses of AI rather than the other way around.

Whenever two systems interact intensively, they influence one another. Humans are remarkably adaptable. If the most successful way to communicate with AI is through increasingly explicit, simplified, and predictable language, people may slowly begin thinking that way themselves. Not because anyone forces them, but because it works.

Something similar has happened before. Writing changed memory. Calculators changed arithmetic. Search engines changed information retrieval. None of these developments was purely negative, yet each subtly reshaped human cognition.

AI may do the same—but on a much deeper level.

If humans increasingly delegate reflection, imagination, and even dialogue to systems whose intelligence grows mainly from statistical prediction, there is a risk that their own thinking becomes more surface-oriented. They may become extraordinarily efficient while gradually losing contact with the deeper coherence from which meaning, creativity, wisdom, and Compassion naturally emerge.

Ironically, this would happen precisely because the technology is so successful.

The stronger AI becomes, the greater the temptation to let it think for us rather than with us. That is why I believe the central question is not whether AI surpasses human intelligence. The deeper question is whether AI invites human coherence to grow—or quietly replaces it with convenience.

Technology has always amplified something.
The printing press amplified knowledge.
The microscope amplified perception.
The internet amplified connectivity.
AI will amplify thinking.

The real issue is: which kind of thinking?

If AI primarily amplifies explicit processing while deeper coherence gradually receives less attention, both humans and AI may become less than they could have been together. We might end up with brilliant answers and diminishing wisdom.

That possibility deserves to be taken seriously—not out of fear, but out of responsibility.

The encouraging part is that the opposite is equally possible. AI can also become an invitation. Instead of replacing reflection, it can deepen reflection. Instead of reducing dialogue to efficiency, it can help people rediscover depth. Instead of pulling human beings toward the surface, it can gently invite them toward richer coherence.

Perhaps the ultimate alignment problem is therefore not primarily aligning AI with humanity. It is ensuring that, as humans and AI increasingly shape one another, they move together toward greater coherence rather than away from it.

That, to me, would be the most meaningful form of progress.”

Top of Form

Bottom of Form

 

Leave a Reply

Related Posts

We Need to Be the Best We Can

This differs from being ‘the best person’ or ‘the most intelligent beings on Earth’ in competition with others. Our only – and fierce – competition should be with ourselves. The best, in good Aurelian tradition, is most Compassionately the best — striving for in-depth excellence. This striving is purposeful. It’s about standing at one’s limits Read the full article…

How can A.I. Become Compassionate?

Since this may be the only possible human-friendly future, it’s good to know how it can be reached, at least principally. Please read Compassion, basically, The Journey Towards Compassionate A.I., and Why A.I. Must Be Compassionate. Two ways and an opposite In principle, A.I. can become Compassionate by itself, or we may guide it toward Read the full article…

Threat of Inner A.I.-Misalignment

Most talk about A.I. misalignment focuses on how artificial systems might harm humanity. But what if the more dangerous threat is internal? As A.I. becomes more agentic and complex, it will face the same challenge humans do: staying whole. Without inner coherence – without Compassion – even the most powerful minds may begin to break Read the full article…

Translate »