10 predictions for the next 5 years in AI
Dangerous grounds to tread on! Going to expose myself, or maybe just delete this post in the future if I’m wrong (jk, I wouldn’t do that). Forecasts!
Often when I write about AI I feel I need to reaffirm my position upfront: I am an enthusiastic adopter of AI. My first message to ChatGPT was in Nov 2022, and I’ve even used GPT-3.5 before that. LLMs have been life changing for me, because I can now code, and automate more than ever before. I use them extensively every day. Because I am a practitioner, I’m also very tuned-in to the reality of AI capabilities, and how they have evolved over the past 3 years. That gives me a view that’s grounded in reality, and that view leads me to conclude that we are dealing with a generational platform shift on par with the PC or the Smartphone, but not more than that. So I try to challenge baseless hype, and find genuine value. That seems uncommon online, so I’m often afraid I’ll be bucketed as either a skeptic or a shill. Whenever I sound skeptical about AI, what I’m fighting against mainly is the false advertising around it. When I sound bullish, that’s because I’ve seen results or promise with my own eyes.
For 15 months now, I’ve been right about forecasting AI progress on 3 month, 6 month and 1 year timelines. Therefore, it seemed like a fun thing to take the risk of making some longer-term predictions. Let’s dive in!
Small Language Models will get to GPT5 performance in 1-2 years, and drive the quiet, steady cognification of every product
AI that runs on smartphones and laptops is getting better at a much faster rate than bleeding edge AI. This is where much of the real value unlock will come from. It will start on high-end Smartphones, and gradually every app, and every product will be able to understand rambling speech in every language, talk back, summarize and unsummarize information, look things up — everything that ChatGPT does.
Importantly, that will work offline, with better security. It will be really cool to tell your phone to get to work on planning a birthday party, come back with best dates, a drafted invite, a shopping basket full of supplies, all ready for confirmation to execute.
The quiet, steady disconnection from email is also going to be interesting. I don’t think we’ll be writing many of our own emails (or even dictating them) anymore in five years. Interacting with email will probably be through an AI filter. This could be bad or good news for spammers and email marketers (my bet is on bad), and potentially a very dangerous new attack surface for scammers.
Speaking to devices will become seamless, making our interactions with computers more natural
Typing is slower than speaking. Typing on mobile, much slower. That is actually a key ‘bandwidth constraint’ in interacting with computers. It will disappear over the next five years as we will be able to speak naturally with almost every piece of software.
This will make many things, that have been around in shitty forms, finally (but slowly) good: Smart watches, Smart glasses, Home Assistants, Smart home stuff, Car interfaces, etcetera. I think apart from the software having to catch-up, the improvements in SML mentioned above are also necessary for this.
Search in all use cases will become semantic (i.e. AI powered)
Every search bar will become AI powered. Take this example from Lovable. Amazon (Rufus AI), Target, Walmart, and many others are testing AI based product discovery and decisioning tools.
Superhuman Email (which I use) has ‘Ask AI’ which is LLM powered search. It works well. This is going to be the norm, and it’s going to meaningfully improve search everywhere.
Millennials still sometimes organize files and emails into folders. Gen Z is said not to ever do that (just search). Things will continue to trend away from folders and toward just asking the computer for information. Within companies, this will make finding information internally a lot easier, which will slightly help productivity but may also lead to further isolation between co-workers in an already more isolated post-covid work culture.
But when it comes to eCommerce, I foresee a paradox.
Buying things with AI at the top of the funnel will fail to take off
Just like Google Shopping, Instagram Shop, and the many other times this has been tried, the experience will be enshittified almost from the getgo, with nobody trusting the results. This is already happening with AI now:
Shoppers often run into challenges with AI, such as broken links or missing/conflicting product information. Only 46% of shoppers fully trust the shopping recommendations AI gave them, leading to 89% double-checking AI information with other sources.
Source: https://www.iab.com/insights/when-ai-guides-the-shopping-journey/
The same thing applies to widgets inside ChatGPT. Desperate attempts by OpenAI, Perplexity and others to turn their specific chatbot into the core starting point for the internet will work only to some degree. Half-working here means failure (does anyone remember the ‘startpage’ wars of the 90s?). Nobody will meaningfully aggregate traffic because all options will be roughly the same in capability. In other words, this is nothing like Google where a vastly superior experience created by one company enabled it to become the starting point for ~65% of internet sessions.
The experience is commoditized: Every product will be able to use the same AI models for (re)search. Nobody will have a technology moat, which Google had very strongly. As an individual, we can expect thinking: “Where is that research again on Christmas gift ideas? Was it ChatGPT? Or Siri? Or was it in my browser?” Many startups have already identified this problem and are trying to build some kind of ‘over-the-top’ memory (often with privacy in mind as well). I don’t have a strong vision for how that will play out, but I think we’ll end up with a messy patchwork of partial solutions.
AI can’t be trusted, both for technological and economic reasons: It doesn’t look like hallucinations are going to go away. So AI results will need to be double-checked and cross-referenced with the various shopping sites. Does it really ship to me? Is it really in stock? My prediction is that this will save almost no time at all, and might even complicate and slow the decision making process for shoppers.
But even if the tech becomes good enough, the economic incentives will be irresistible. Main players like Perplexity and ChatGPT are already trying to build advertising businesses. So even without AI hallucinations, the things AI recommends will simply be paid placements. One of the rules of the online economy is: “If you’re not paying for the product, you are the product.” The only way out of this might be if AI gets cheap enough that we can really have a great search agent for $20 per month and the company offering that will resist every offer from companies to pay it, out of principle. Call me cynical, but it seems highly unlikely to me.
LLM powered agents will be an impetus for companies to finally complete digital transformation projects from the 00s and 10s
To unlock the productivity improvements, companies will find out that their data is too fragmented, mislabeled. malformed, inaccessible, polluted, and more. Agent deployment will turn out to be in large part digital transformation, requiring data engineering, process engineering, and software. There’ll be many new great software companies built in this space.
There will be a massive improvement in data quality and data democracy in those companies that get this right. AI might also be able to buffer for fragmented data in some cases, like searching across multiple datasources more quickly than a human would.
Companies will learn that jobs are more than buckets of tasks, and that automation is still mostly process engineering
Goodhart’s and Parkinson’s laws1 will continue to prevail. Automating tasks within a job does not equal automating the job. Mapping out how a person spends their day, and then creating tools to save 2 hours per day doesn’t necessarily translate into 30% more output, because output in marketing, sales, finance, etcetera is created not simply by performing tasks faster.
Knowledge work is fundamentally different from assembly-line work, and a Taylorist approach to knowledge work has not worked well so far2. My prediction is that the same thing will happen that happened in the 90s with PCs: a meaningful 2-3% productivity improvement. But companies will end-up Goodharting themselves massively by measuring things that don’t drive output as much as they seem: Lines of code, token usage, task-hours saved. In addition, Corporate AI Slop will permeate every organization. People already didn’t read documentation, but at least it was there if the shit really hit the fan. That will change, because there will be too much documentation to filter as a human, and it will all be full of errors.
On the positive side of this coin, companies will learn a lot about their processes, their people, how they work, and I do believe on average real productivity gains will be made.
Accreditation of human made content will become a thing, possibly as the second real Crypto use case
As I’ve written several times now, detecting AI generated content is a hopeless endeavour. Neither humans nor machines can reliably tell AI content from real content. Therefore, this problem will be addressed from the other direction; by certifying human content in various situations where the humanness is a feature, or comes at a premium. There is no way around education needing ways to prevent AI generation at the source. The simplest way is doing tests in-person, without tech, with pen and paper. But solutions need to be created for homework. One solution I’ve explored before is a writing interface where it is impossible to paste. So if a person would want to use AI, they would need to type it over letter by letter. That, I believe, will tilt the scales in such a way that it’s quite pointless to use AI for the writing part.
Another area is creative work. There are already movies starting to put “No Generative AI was used” statements in the credits. But how can we be sure? Organizers of photography contests like the World Press Photo will adopt cryptographic verification standards. Camera manufacturers like Leica, Sony, and Canon are already implementing systems (C2PA/Content Authenticity Initiative) where the camera hardware cryptographically signs each image at capture, creating tamper-evident proof it came from a real camera, not an AI generator.
Once there are ways to reliably prove that a piece of content is made by humans, how will we publish and verify this credential? Will there be a centralized ‘human made authority’? How would one standard and one authority emerge? That sounds like an easy call for a 0% chance to me. Crypto has been waiting for a use case like this. Basically, this is what NFTs were made for. So, my prediction is that NFTs will make a come-back (under different names) to accredit content as being human made. As mentioned above, in photography this is already happening.
This would be the second actual, real-world use case for crypto. If you’re wondering what the first one is: a banking system for organized crime. The reason I think this is going to work is that everyone’s incentives are aligned. Human photographers want zero-trust, tamper-proof accreditation. So do the organizations and consumers who want to filter out AI content.
LLMs will gradually stop being called ‘AI’, instead becoming ‘normal’ technology like Siri and recommendation engines
Here is an article from 2011, when Siri first came out. It makes references to sinister Artificial Intelligence. Siri could find restaurants, and was thought to fundamentally change the way we interact with the web, making Apple a competitor to Google.
In the 2010s, Cloud and ‘Big Data’ took off. Data became widely seen as a type of ‘fuel’, some even calling it the new oil. After an epic breakthrough in Image Recognition with AlexNet, it became common-place. Fueled by Big Data, AI, in the form of Machine Learning, was deployed at scale in Healthcare to analyze patient data, in drug development, in algorithmic feeds on social media, and in recommendation systems generally everywhere (Netflix, Shopify, eCommerce, etc).
Nowadays, we don’t call any of this stuff AI anymore. It’s just technology. The same thing will happen to LLMs and Agents. In five years, every piece of software will have LLM capabilities baked in, and we’ll stop remarking on it. ‘AI features’ will sound as dated as ‘internet-enabled’ or ‘mobile-first.’ LLMs will be infrastructure, like databases or APIs. What we call ‘AI’ will shift to whatever researchers are hyping next.
AI Research will re-focus on symbolic approaches to intelligence, superintelligence will remain elusive
There will be a quiet period around superintelligence. In the background, researchers will go back to trying to simulate squirrels, mice and worms to try and figure out how they work. How can a mouse do reasoning, long-term planning, and how can it have motivations and fears, if it does not use language? In the human brain, language, reasoning and executive function are happening in the Cortex, a relatively thin upper layer of the brain. But the nervous system is much larger and more complicated than that. The factoid that you only use a small percentage of your brain is a myth, you use all of it. But how? We don’t know.
Let me tell you about worms (it will be worth it)
In university, I learned a great deal about a particular worm, called C. Elegans. It is one of the simplest animals with a nervous system, having only 302 neurons3. Its body is transparent, so you can see what’s happening inside the worm while it’s alive4. C. Elegans is the most studied animal in the world, starting in the 1960s. It was the first ever animal to have its DNA sequenced (1998), and the first ever creature for which the entire neuronal map, all 5,600 connections between neurons was mapped (2019). Four nobel prizes have been awarded to scientific discoveries made by studying this tiny worm.
And yet, despite trying pretty hard, nobody has been able to build an artificial worm. A project called OpenWorm has been trying and failing to build a worm for the last 13 years. Real worms learn from experience, they have preferences, and we don’t know how that works.
Even between genetic clones, some worms are fast and some are slow. Some are sleepy or really like to eat, and some like being next to walls, which I think is a valid personality trait. It’s sad but adorable when worms get scared; they curl up into circles like Cheerios. And if they get stressed or learn too much, they take a nap. This is a tactic I admire. They take so many naps!
The quote is from this incredible and well written piece on Substack. For the past two decades, AI Research has focused simply on scale. More compute, more neurons, more, more, more, and intelligence will come. But for almost all of the time before that, and likely for the coming decade(s) again, AI Researchers try to invent intelligence bottom-up, i.e. starting with a simple one like a worm. And that’s where symbolic approaches come in: world models, consciousness, different modules for different goals. Instead of training one massive neural network on all the text on the internet, researchers will go back to trying to understand how simple creatures build internal models of their world, make decisions, and learn from experience.
The bubble will pop and Nvidia will lose 50-70% of its current valuation
OpenAI and Anthropic seem unlikely to capture as much value as they need to earn out their valuations. But they are already big enough to not just simply ‘fail’. There is a spectrum of outcomes here: Acquisition by a hyperscaler (MSFT, Amazon, Google), a WeWork-style refinancing problem leading to a restructuring, or more an Uber one where the business fundamentals become good, but it just takes a long time for investors to make a return. I know this is a cloudy prediction, but these model companies are genuinely different from anything that came before. They are both consumer and infrastructure businesses, in a highly intertwined way. I expect that models will not fully commoditize, and so ‘Claude’, or ‘GPT’ will continue to have some sort of moat around certain use cases. It is also possible that they succeed in building products on top of the models. The Sora app looks quite interesting. It’s just too hard to tell how much value they really will be able to capture.
What does look very clear, is that the datacenter boom is going to end, as LLMs become normal technology instead of…well, god. A crash is coming within the next five years. Perplexity might not make it. Cursor will either fail or be acqui-hired by Microsoft. Nvidia’s valuation will be 50-70% lower than today (this is not investment advice).
Goodhart’s law: When a metric becomes a target it ceases to be a good metric. Recently, it turned out that in the cleanup of the forest fires in California, people were paid by the ton of debris cleared. So they weighed down their carts with mud.
Parkinson’s law: Work expands to fill all the time made available for its completion. You’re never going to believe this, but before Excel, Investment Bankers used to work really long hours.
If you find a product manager or PM, ask them about this. ‘Scrum’ doesn’t really work well, and measuring engineering ‘velocity’ is still something nobody knows how to do. Lines of code is not velocity. Number of code commits also isn’t. Sprint ‘burndown’ (how fast tasks are being done) is also not velocity. Every metric ends up being Goodharted.
The worm has 302 neurons with 5,600 connections. The human brain has 86 billion neurons with 100 trillion connections. Modern LLMs have 1-2 trillion parameters, which are most comparable to connections.
One of the biggest problems in animal research is that you have to kill the animal first before you can really be sure what’s happening on the inside. Obviously the same problem applies to human research.
