After the Internet, AI Has to Learn How to See the World

EditorsDossiers1 week ago78 Views

The web gave AI text and images for free. Biology, robotics and science need observations that do not yet exist — and producing them changes the economics of AI.

Generative AI grew quickly because the internet already existed. Text, photographs, code, arguments, documentation and video created an enormous stock of material from which models could learn patterns.

That abundance also created an illusion: that improving AI is mostly a matter of building larger models and adding more compute. In many fields, the real bottleneck is becoming something else: the data we need does not exist yet.

On October 7, 2026, Reuters reported on Biohub’s Virtual Biology Initiative, backed by the US government, Google, Meta and other partners, with $1.8 billion in announced investment. Its goal is to produce large open datasets on cellular behaviour for models designed to predict biological phenomena. The project points to a fundamental shift: instead of finding data already lying around on the web, AI increasingly needs experiments designed to create it.

The internet was a dataset built for other reasons

Large language models benefited from a historically unusual situation. For decades, billions of people wrote pages, books, forums, software and documentation without doing so to train an AI system. When generative models arrived, an immense amount of information was already there.

But the web represents only some parts of reality well. It is rich in language and images. It is much poorer in standardized biological measurements, molecular interactions, robotic manipulation or experimental outcomes.

Biology cannot simply be scraped

To build a model that predicts how a cell responds to a drug, it is not enough to collect papers describing previous studies. The system needs direct measurements: which genes activate, how proteins and cellular structures change, how behaviour shifts under different conditions.

Those measurements have to be produced by laboratories, scientific instruments and consistent protocols. They cost time and money. They also have to be standardized enough to be compared.

The Virtual Biology Initiative is aimed precisely at this missing infrastructure: experimental datasets generated at a scale large enough to become the basis for predictive models of biology.

Robotics runs into the same wall

A language model can read billions of sentences. A robot learning to manipulate objects needs data about the physical world: movement, force, failure, friction, deformation, geometry and the consequences of actions.

Some of that data can be generated in simulation, but simulation is never a perfect copy of reality. Companies therefore collect data from real robots, teleoperation systems and sensors. Again, AI encounters a limit that cannot be solved by scraping more web pages.

The problem is no longer finding enough text. It is building a system capable of observing and measuring the world.

Synthetic data helps, but does not abolish reality

One answer is synthetic data: examples generated by models or simulators rather than directly observed. Synthetic datasets can enlarge a corpus, cover rare situations and reduce some costs.

But synthetic data still begins from a representation of the world. If the initial model is incomplete, generation can reproduce its errors and blind spots. A physical simulator may produce millions of scenarios, but volume cannot compensate for a variable that was represented incorrectly from the start.

Real and synthetic data will therefore continue to coexist. Real observations anchor the model to the world; synthetic data multiplies combinations and expands the space that can be explored.

Economic value moves toward whoever can produce observations

When the internet supplied abundant material, competitive advantage was concentrated heavily in the ability to train larger models. In scientific and industrial AI, advantage may shift toward organizations with laboratories, sensors, experimental platforms or privileged access to users and devices.

That changes the geography of power. Universities, hospitals, pharmaceutical companies, robotics manufacturers and industrial firms can possess data that the largest technology platforms do not.

Partnerships then become unavoidable: Big Tech brings capital and compute; other institutions control access to the physical world.

More data is not automatically better data

The race for scale can hide a quality problem. A huge dataset filled with measurement errors, inconsistent protocols or unbalanced samples can produce models that look powerful while remaining unreliable.

For scientific applications, quality matters even more because errors can propagate into expensive experiments or clinical decisions. Metadata, protocols, traceability and a clear record of how each observation was produced become part of the model’s credibility.

That is a profound difference from much consumer AI, where an error may be annoying without necessarily becoming critical.

After the internet, AI will have to build its own senses

The first phase of generative AI exploited a digital world already overflowing with data. The next phase may require a new infrastructure built specifically to produce observations.

Automated laboratories, robots, sensors and simulations become the equivalent of the crawlers that once mapped the web. But they do not merely search for existing information: they create new observations.

This suggests that the limit of AI will not necessarily be the size of the next model. In many domains, it will be our ability to build datasets describing aspects of reality that the internet never recorded.

After learning from the web, AI has to learn how to look at the world. And looking is much more expensive than downloading a page.

Sources and references

Leave a reply

Loading Next Post...
Search
Loading

Signing-in 3 seconds...

Signing-up 3 seconds...