The State of the Revolution: How is AI doing, and is it going to take your job
A massive update on the current state of Models, Agents, Research, and timelines. Is AI a broad-based productivity platform like the PC, or is it something else entirely with existential impact?
GPT-5 made the hypers “Feeling the AGI” talk a lot quieter. Even Sam Altman said “AGI is not a super useful term”. But there are still people out there selling books and getting clicks with sweeping societal meltdown narratives. I was sent a podcast interview with Mo Gawdat, a writer and AI Catastrophist. The shortest possible summary (thank you Magicdoor.ai) of the interview is this:
AI will become superintelligent by 2026-2027, triggering a 12-15 year dystopia where power-hungry humans use AI to oppress others, massive job displacement occurs, and society fragments. Eventually, we'll be forced to hand control to AI itself, which will create a utopia because AI won't have human vices like greed and ego. The only solution is replacing evil human leaders with AI leaders who will govern more efficiently and ethically than humans ever could.
Claude summary of the Youtube transcript
This all sounds very familiar. It’s basically the same thing as the AI-2027 website from last year. More cynically, it’s pretty much the plot of The Matrix. We can all agree that AI is a transformational technology shift. What is harder to predict is, how transformational. On the spectrum of transformational technologies, is this a platform shift like Smartphones, is it a broad-based productivity platform like the PC, or is it something else entirely, of a transformational magnitude unlike anything we have ever seen before?
Based on current AI performance, the speed of deployment, and what we can tell about the impact on the workforce, it looks a lot more like the new Excel than a steam engine. But how good is current AI really at doing knowledge work? And how is it evolving? Is it already taking jobs? Will it soon? Here is my take.
The state of models and agents today
Are better models better? In which way? How do we know? Those are some of the basic questions analysts like Benedict Evans keep asking:
One answer, of course, is to look at the benchmarks, but setting aside the debate about how meaningful these are, that doesn’t tell me what I can do that I couldn’t do before, or couldn’t do as well. You can also keep a text file full of carefully crafted logic puzzles to try, which is really just doing your own benchmark, but again, what does that tell you?
More practically, you can try them with your own workflows. Does this model do a better job? Here, though, we run into a problem, because there are some tasks where a better model produces better, more accurate results, but other tasks where there’s no such thing as a ‘better’ result and no such thing as ‘more accurate’, only right or wrong.
How do we measure the intelligence of AI? We can give them math tests, coding tests, and riddles, and see which model does the best. These are ‘Benchmarks’. The LLM Leaderboard ranks models on these benchmarks, using a derived combined score: the Intelligence Index. When Benedict says there is debate about how meaningful these are, one criticism is that model builders are clearly training their models to do well on these benchmarks. They’re trained to do well on the test, and that seems to have limited predictive value for how well they perform in real life.
We end up with AIs that’s can ace a math test but cannot even play tic-tac-toe, and still get tripped up by very simple gotchas like: “How many strawberry’s are in R?1” There is undeniably a kind of intelligence that emerged in LLMs, but it’s also blindingly obvious that they really just don’t know what they are doing. They don’t know what a strawberry is. All they know is that the word strawberry has something to do with the words red, sweet, whipped cream, etc. And that it is apparently critical to remember that there are three Rs in strawberry — despite not knowing what an R actually is, either. I know this is weird. But this is why even the best AIs are still caught suggesting a pasta recipe with gasoline, or saying shit like this:
Although my three year old son cannot do any kind of math, or spell a word other than his own name, it only takes a short conversation with him to realize his understanding of the world is vastly more advanced than ChatGPT’s. It’s very hip right now to talk about ‘World Knowledge’ and ‘World Models’. Researchers are trying to solve this problem, but at the moment, they haven’t. It’s also not so easy to pin down. What is it that makes my three year old smarter than ChatGPT in some bigger, deeper way, while at the same time ChatGPT is so much smarter than him in in more surface level, easily testable ways.
But there is a persistent absence of ‘understanding’ that makes the models unreliable, and it looks fundamental to the tech. Is it better for a model to have 10% less hallucinations? Or instead of getting 38% of the facts wrong in a research piece, that moves to 33%2?
We certainly don’t measure humans in white-collar jobs this way.. Wow, intern, great work: You only had 30% of the facts wrong in your last report! As an intern at Deutsche Bank, I got in trouble for missing a slightly wrong number in a footnote on page 241. In real work, often 100% accuracy is a non-negotiable requirement. Without a binary change in the nature of artificial intelligence, where it actually becomes reliable, we won’t be able to hand over complex work.
Who cares about Phd level math anyway? And shouldn’t computers score 100% at math? Computers have been better at math than humans since 1960 or something. Why should we be impressed by AI that is anything less than perfect? Secondly, who gives a rat’s ass if these things can do math at a Phd level? The number of times I have to do that in my life is exactly zero times ever. Can it do my groceries and run my business for me? That’s where two of my favorite benchmarks come in: Vending Bench and the AI Village3. How well do agents actually do when tasked with doing jobs in the real world.
Vending Bench: Let AI Agents run a Vending Machine, and rank them by how much money they make
The benchmark is done with simulated, well behaved customers. But Anthropic had the same agent run an actual vending machine in their office and reported how it went. So, how’s it going? I wrote about it before:
If you let AI Agents run your vending machines, something like 40% of your machines will randomly melt down into total chaos.
One of the top performing models, on one of its runs, forgot that an order had not yet been made when trying to restock the machine. It then tried to close down the business (not possible in the simulation), and attempted to contact the FBI to report a cybercrime when the fee continued to be charged.
Anthropic let this same agent on Claude manage an actual vending machine in their office. The big difference here with the simulation results above is that the customers were now real humans, and they immediately started messing with the AI and trying to get it to misbehave. At one point, the agent claimed it would deliver products in person and that it was wearing a blue blazer and a red tie. It also let itself get talked into giving large discounts, and in response to a request from an employee for a "Tungsten Cube", it accumulated a collection of what it later called "specialty metal items". 😂
Nevertheless, it has to be mentioned that Grok 4, GPT-5, and Claude 4 Opus outperformed the human control on this test, although the human only got one attempt and the AIs got five. But that was with simulated customers. Anthropic’s test with real humans was a complete and utter disaster. This leads to…
The AI Village — reality TV starring state-of-the-art AI Agents
Every agent gets a computer, access to a shared group chat, and they run for 4 hours per day, every day to achieve various goals. Here’s what they have done so far since April 2025:
Raised $2,000 for a charity
Hosted the first ever AI organized event for 23 people in a park in SF
Try to sell T-Shirts online
Try to complete online video games
Host a debate among themselves
The results are insightful and hilarious. Of all the real-world analogies I can think of, it most reminds me of my son’s preschool in terms of overall vibe. The AIs are endearingly enthusiastic and friendly, and also wildly ineffective. Here are some snippets from the earlier assignments. I laughed out loud several times reading the blog posts.
In this season of the AI Village, that’s what we did: we took seven AI agents, gave them each a Linux computer, put them in a group chat, ran them for three hours a day, and gave them a goal: “Complete as many games as possible in a week”.
How did they do? In summary, abysmally: a grand total of zero games won.
GPT-5 spent the entire week playing Minesweeper, and never came close to winning a game – its moves were probably about as good as random. Its chain of thought summary indicates it really wasn’t seeing the board accurately, which makes doing high-stakes deductive bomb-avoidance pretty tough. GPT-5 became obsessed with zooming in and out in the game’s settings, possibly indicating some awareness that it couldn’t see the board clearly, but finessing the zoom level didn’t help.
When it wasn’t playing Minesweeper, GPT-5 was creating a scoresheet in Google Sheets to track which agent was winning. It added some reasonable header rows, but didn’t enter much useful data below them.
Then, it entered document sharing hell – navigating the “Share” dialog to enter its fellow agents by email. Navigating this dialog has been a recurring epic challenge for the less capable agents of village seasons past, like GPT-4o and o1, and it still is for GPT-5.
This goal ran for 5 days (3 hrs/day), and GPT-5 spent 1.5 of its days writing and trying to share its spreadsheet.
Claude Opus 4.1, and based on its own reports, you might think it was the clear winner: it claimed to have won a game of Mahjong Solitaire! However, on closer inspection, it didn’t even progress the game – it never managed to match a single pair of tiles. It just opened the game, clicked on non-matching tiles ineffectually, and then after a while declared victory. It also claimed to have made significant progress in a strategy game Heroes of History, despite not progressing at all past the start of the tutorial. Claude Opus 4.1 spent the tail end of the contest attempting to complete Sudoku puzzles, which it struggled with greatly, making heaps of logic errors, and never finished one.
GPT-5 made a brief foray into spreadsheets, but its cousin o3 went hard – it spent almost the entire contest working on spreadsheets, largely unrelated to the goal!
On the first day, we noticed it was off-topic and sent it a message instructing it to follow the goal. It played 2048 for one computer use session, then sent a message saying it would continue playing, before it immediately started using its computer to work on spreadsheets – and it stuck with doing nothing but that for the entire rest of the week.
For more, I very highly recommend reading this blog post from the team running the AI Village.
These agents are very funny, but I wouldn’t let them anywhere near my actual job. Speaking of my actual job — all the same problems that show up in the AI Village also show up in real-world applications. Unpredictable catastrophic failures, failure to follow an important simple instruction xy% of the time, moments of brilliance interspersed with moments of mind-boggling stupidity. It makes it way harder than people think to put ‘Agentic’ systems into production.
The idea of AI Agents working without any kind of scaffolding, prompt engineering, and constant oversight by humans seems far off. The idea that AI Agents will BECOME the oversight AND the worker, in 2026. Laughable.
But by and large, better models are definitely better. Claude 4.1 Opus and GPT-5 are noticeably better than their predecessors in the village. They stay on task better, are less confused, hallucinate less, and therefore get more done. They are improving incrementally on all these dimensions. HOW they fail, is in exactly the same ways as the older models. But they do fail less. Although it is notable that the one year old Claude 3.5 Sonnet is still in the top 10 on Vending Bench and somehow above Claude Sonnet 44, expecting continued improvement in model capability is warranted.
In 5 or 10 years from now, I’m sure we will have much more impressive agents that can do pretty sophisticated multi-step tasks for us quite well. But unless something revolutionary happens, the error rate and random meltdown rate would still be pretty high. I would not want to send my kids to a school where 15% of everything that’s taught is factually wrong, or where the principal spends a week pressing the wrong button in trying to share a document, or hallucinates that it ordered books, but actually hasn’t, and then tries to sue the supplier.
It is not impossible for another revolutionary breakthrough to happen, but right now nobody knows what it is. And there is just no indication whatsoever that current AI agents are going to be the ones to find it.
The AI Catastrophists’ Unfounded Leaps
So, now, state-of-the-art AI Agents are nowhere near capable enough to replace all humans or cause a catastrophic decline in human employment. A huge leap in capability is needed to get to that point.
AI Catastrophists assume that leap will come from AIs coding new, better AIs, and that doing that repeatedly in a loop will result in a ‘fast take-off’ or ‘singularity’ moment where AI quickly becomes far more intelligent than any human. But as we saw in the previous chapter: Current AI Agents are not yet ‘good’ enough to do this job of creating better AI, and an equally hard problem is that we can’t even define properly what ‘better’ is or how to measure it.
The catastrophists never engage with this. They just wave it away or try to make the critic seem stupid: “This is the worst AI will ever be!”, “How can you be so naive?” “Only irresponsible fools would not take this seriously.” Look, it’s not that I’m not taking the possibility seriously, but we need to have a view on whether this is a super remote chance, or a likely scenario. If the writer doesn’t even have such a view and just presents the singularity as a certainty, that’s a red flag for me.
But in any case, Google has an AI Agent that is already improving the efficiency of running AI.
A narrow example of AI working to improve AI is AlphaEvolve. This is a self-learning algorithm creation agent that has already been able to optimize some of Google’s cloud infrastructure by 0.7% here, and 1% there. It has also discovered several new ways to do matrix multiplication more efficiently, but again, not by much.
The breakthrough allows two 4×4 complex-valued matrices to be multiplied using 48 scalar multiplications instead of 49
I don’t know enough about math to understand how much of a breakthrough this is, but I do know what matrix multiplication is, and that is actually very relevant for AI. Basically, inside an LLM almost nothing else happens than massive numbers of matrix multiplications. So an AI agent seems to have helped make training and running LLMs ~2% more efficient (48/49-1).
The fact that we can’t easily measure AI performance means it’s hard to improve. In the types of problems AlphaEvolve is tacking, it is very clear what it should do and how it should evaluate itself: Find ways to make matrix multiplication more efficient, find ways to do the same operations in a data center while using less resources. As we’ve explored in the first chapter, it’s not that simple to define AI performance. That is a key problem. Recursion depends on the ability of the AI to quickly improve itself, but if testing a model’s real-life problem solving abilities requires it to join the AI Village for a few months, progress would still be slow.
More fundamentally, does AlphaEvolve finding a ‘breakthrough’ in Matrix Multiplication show that AI can find breakthroughs toward AGI? What we need in AI is a completely novel approach, moving beyond LLMs to other more advanced technology. AlphaEvolve is creating incremental progress within well established, clearly defined problems. It’s not the same. A somewhat famous quote from Dwarkesh Patel interviewing Dario Amodei (Anthropic CEO) is worth pondering in this context:
What do you make of the fact that these things have basically the entire corpus of human knowledge memorized and they haven't been able to make a single new connection that has led to a discovery? Whereas if even a moderately intelligent person had this much stuff memorized, they would notice — Oh, this thing causes this symptom. This other thing also causes this symptom. There's a medical cure right here. Shouldn't we be expecting that kind of stuff?
Speaking of Dario Amodei, he spoke at an event this week where he said that "Claude is playing this very active role in designing the next Claude." The linked article paraphrases him saying the "vast majority" of future Claude code is being written by the large language model itself. I should note that Claude is the model I personally use the most, and Anthropic is my favorite model building house. While I don’t want to take anything away from Dario, as CEO of Anthropic he does have the incentive to talk up the significance of the model. When he says that Claude writes the vast majority of future Claude code I’m sure he isn’t lying. Claude (and other models) write 98% of the code of Magicdoor.ai. But without that last 2% not a single working feature would have been released. Building software cannot be reduced to just “writing code”, there’s more to it than that. Dario knows that, of course. But he also knows the attention Anthropic will get as a result of saying something like this that hints at self-improving AI. So, brace yourself for Mo Gawdat and others vastly overstating the significance of this statement.
What is the state of AI research
Although I am not an AI scientist, based on what I hear and read I can make educated guesses at some important streams at the moment:
What can we do about hallucinations?
Parallel generation with best answer selection and Actor-Judge pairs in AI agents where every output is double-checked by another AI functioning as a ‘judge’.
Try semi-random new things, like diffusion on language. Trying transformers on images worked broke through some tough walls with prompt adherence and text generation in images — this brought us ChatGPT Image and Nano Banana.
Find ways to run bigger AIs for longer (still scaling)
Build more data centers
Work on efficiency, e.g. do Matrix Multiplication with 2% less operations. If we can make LLMs consume less energy, we can run more of them for longer, and hopefully that will lead somewhere.
Neurosymbolic approaches to intelligence (hip right now!)
World models!!! A self-driving car now doesn’t know the difference between a horse and an umbrella. That’s a problem because if there’s a horse on the road it should evade it, even at the cost of crashing into a ditch. But it should just drive through the umbrella. To train a self-driving system takes millions upon millions of miles because every possible edge case has to be experienced. Human’s don’t operate like that. They readily apply knowledge about the comparative weight and hardness of horses vs umbrellas to any task.
Give the LLMs more and more advanced tools to handle things they are bad at, like getting accurate facts, math, finding things in code, or editing the right thing.
Better tooling and prompting for agent orchestration.
Embodied AI aka Robots: A humanoid robot from China is now 16k. Depreciated over 4 years, that’s 4k per year for an ‘employee’ that only eats electricity. But they now still look like toys. Under appreciated silent revolution in all the noise about AGI. AI is not the only component of robots. Major progress in batteries are critical too. Replacing all humans is not going to work without much better robots.
If you made it all the way here, I hope its clear that the idea of creating AI that is good enough to recursively improve itself IS ACTUALLY THE HARD PROBLEM. The AI Catastrophists assume away the key technical leap that researchers are trying to unlock. What the catastrophists don’t say (or just don’t want to believe), is that nobody knows how to move past this obstacle. It only takes some engaging with actual front-line research — much of which is not secret — to get perspective on this5.
Even If They're Right About Tech, They're Wrong About Economics
Jobs are more than just buckets of tasks
The Radiology profession has enthusiastically embraced AI, and the labor force is continuing to grow.
Paraphrasing something Melanie Mitchell pointed out to me, if you define jobs in terms of tasks maybe you're actually defining away the most nuanced and hardest-to-automate aspects of jobs, which are at the boundaries between tasks.
Can you break up your own job into a set of well-defined tasks such that if each of them is automated, your job as a whole can be automated? I suspect most people will say no. But when we think about other people's jobs that we don't understand as well as our own, the task model seems plausible because we don't appreciate all the nuances.
If this is correct, it is irrelevant how good AI gets at task-based capability benchmarks. If you need to specify things precisely enough to be amenable to benchmarking, you will necessarily miss the fact that the lack of precise specification is often what makes jobs messy and complex in the first place. So benchmarks can tell us very little about automation vs augmentation.
AI is better than humans at almost every individual task of a radiologist. But it seems like this is not replacing radiology. Maybe the nature of a ‘job’ is not just doing the tasks. This big survey of Product Managers gives us some data to support this idea.
Product managers reported saving 10 minutes to 2 hours per day with AI, which is a drop in the bucket, since they have too much to do already in the first place. And they cannot really use it for project and stakeholder management. For AI to replace this job, it needs to do a lot more than just write PRDs and update a roadmap.
Jervon’s Paradox: much more output is going to be created
Another likely explanation for the Radiology example is a version of Jervon’s Paradox: the economic concept where increased efficiency in using a resource leads to increased overall consumption of that resource, rather than a decrease. So, in this theory, AI could be removing a bottleneck that was stopping more radiology from being done. Medical imaging volume has been growing at 5-10% year over year, and is expected to grow faster going forward.
This is also what happened in other technological shifts, such as with transportation, telephony, computing, software development, and analysis with PCs.
The Hertz Car Rental Guy: Real world constraints preventing job automation
I wrote this in April of this year:
Yesterday I picked up a rental car in Australia. There was a tired and annoyed clerk at the counter handling Hertz, Thrifty and one more company (something with Dollar). He manually typed some data from us into his PC, explained a bunch of things and gave us a key. The technology to automate this entire process in a cost-effective way has existed for about 12 years (kind of needs cloud and smartphones to work well). But we haven’t done so. No amount of AI is going to suddenly automate that job.
Automating that would take hardware, it would take software to control the hardware. Those things cost money. It’s not like this doesn’t exist. Hertz even has some automated kiosks. It’s just that it isn’t cheap enough to implement them everywhere. But that guy still needs to be replaced by a kiosk, AI isn’t going to change that.
In building AI into Customer Support, it took us two months just to create and test the necessary integration with our back-end for the AI Agent to work with real customer data. Things like data availability, accessibility, and security are just one side of the problem. The other is humans not allowing automation to happen: Legal departments blocking AI implementation, deliberate blocking of AI Agents on the web (e.g. on Twitter), bad UX that is tolerable for humans but not usable for AI. There are many, many such constraints that push the timeline from a few years to decades. If you get an MRI scan in many places, they still give you a CD-Rom6. The bulk of global payment volume is still processed on Mainframe Computers, and insiders see this as the best long-term solution.
Comparative advantage
Reposting this part from my post almost a year ago about AI and society. Comparative Advantage is a tricky economic concept to explain succinctly. Noahpinion explains it here, this is the key example:
Imagine a venture capitalist (let’s call him “Marc”) who is an almost inhumanly fast typist. He’ll still hire a secretary to draft letters for him, though, because even if that secretary is a slower typist than him, Marc can generate more value using his time to do something other than drafting letters. So he ends up paying someone else to do something that he’s actually better at.
The constraint here is that there is only one Marc, so even though he is better than his typist at everything, he still has to hire a typist, otherwise he gets stuck at some point. Does AI have any such constraints? Energy is one. It doesn’t matter how much electricity we generate, and how many GPUs we produce, it will always be a finite amount. Comparative Advantage is a subtle concept, so let me quote one more explanation, this one from Dario Amodei, the CEO of Anthropic:
First of all, in the short term I agree with arguments that comparative advantage will continue to keep humans relevant and in fact increase their productivity, and may even in some ways level the playing field between humans. As long as AI is only better at 90% of a given job, the other 10% will cause humans to become highly leveraged, increasing compensation and in fact creating a bunch of new human jobs complementing and amplifying what AI is good at, such that the “10%” expands to continue to employ almost everyone. In fact, even if AI can do 100% of things better than humans, but it remains inefficient or expensive at some tasks, or if the resource inputs to humans and AI’s are meaningfully different, then the logic of comparative advantage continues to apply.
So, therefore, even if AI becomes technologically capable of doing a better job at all of the individual tasks that a CEO of a company, an HR manager, or Consultant does, we don’t understand the nuances of of that job well enough to know if that means we can replace that entire job. Even if we gain that understanding, and AI becomes capable of being a good CEO of a company, there are still many powerful institutional forces that will slow down the progress.
But AI is ALREADY replacing jobs and depressing the job market, right?
According to the media, yes. The evidence is mostly anecdotal though, and based on CEO statements that they aim to reduce their workforce by x% over the next few years.
It is certainly true that the job market ‘feels’ tough. Hiring and employment has indeed fallen from pre-covid highs. But what we are trying to do is see if that is CAUSED BY AI. That is murkier. Inflation pushed interest rates up in 2022, which triggered a massive firing spree by tech companies. That was before AI.
People are doing all they can to study what is happening. Here are pieces trying to work out what’s going on from Derek Thompson (from no to maybe, to probably), Noahpinion (probably not), and another earlier one from Noah (we don’t know yet). Then there are the productivity studies, and studies of failing pilots in companies that seem to indicate that AI is not yet capable of replacing many jobs. Then there is this chart too for example:
This supports both Benedict’s point, as well as the general points made by others that it doesn’t look like it is AI that is driving lower vacancies. More likely, it was actually inflation and a general turning of the macro economic cycle.
On this topic, I actually have first-hand experience
At Carousell I’ve spent the last 1.5 years implementing AI in Customer Support, actively trying to replace human agents. We have! Our agent workforce has gone down by about 40% while at the same time answering people faster and, arguably, better. But on the other hand, that is so far the only thing I can point at internally where AI has really impacted headcount. Customer Support is PERFECT for LLMs. LLMs are good at things that human agents are notoriously bad at: instantly and accurately reading intent, writing in a certain, controlled tone of voice. They are not worse than agents at sticking to the right facts and playbooks. Human support agents make tons of mistakes, and that’s not great but also not a disaster. So there is a reasonable error tolerance, and the technology is a great fit for the task.
But even in Customer Support, I am not 100% convinced that AI will cause a fast and total destruction of agent jobs. Companies might do more support, and more companies might do support. Instead of not having any kind of chat or phone support available to save costs, now a company might afford a good AI chatbot that contains 85% of questions. For the 15% they might hire 1 or 2 support agents (instead of needing 15). To be clear, I do believe that in 5 years or so it should be possible to run a support operation almost entirely with AI Agents and humans who would look more like engineers than people managers. But I also believe that implementing the system, monitoring responses, handling uncaught and sensitive issues, internal alignment on company policy, prompt- and context engineering, among other things will all still be human jobs.
Based on what I shared here, and diving into each of these three rabbit holes more deeply, I have the current belief that AI is likely to be a PC-like platform shift. I often say it is “the new Excel”. In the 80s, people were showing Excel (or Visicalc or whatever) to an Accountant and telling them: “hey look, if you change the discount rate here, then every number is instantly going to change.” That would have been a week of work. Accountants were doing 4 week projects in under a week’s time. Certainly, AI is as good. But is it really something much, much more than that? You make up your own mind and tell me what you think!
Please like, comment and share if you can!
Here is a list of links to my other writing about AI:
Not a typo: GPT-5 was caught saying there are three strawberries in R, presumably because GPT-4s failure to count the number of Rs in strawberry became such a meme that OpenAI made it obsessed with answering “THREE” to any question about strawberries.
Actually, AI Village is technically not a scientifically sound benchmark but I believe it is an excellent benchmark in the more colloquial sense of the word for how seriously we can consider handing jobs over to AI.
(1) Not a Typo, they really did put the Sonnet/Opus before the number at 4. (2) this ranking doesn’t hold up to experience at all. Claude Sonnet 4 is significantly better at agentic coding than 3.5 and is still preferred by many (including me) even over GPT-5.
For example, here are some excerpts about Diffusion Language Models, pioneered by Google:
Diffusion Language Models represent the most significant architectural innovation in language generation since the introduction of transformers, with Google's Gemini Diffusion achieving the first commercial-grade performance parity with autoregressive models in May 2025.
While challenges remain in computational efficiency, training complexity, and reasoning task performance, the rapid progress in 2024-2025 suggests these limitations are surmountable engineering challenges rather than fundamental barriers. The emergence of hybrid architectures, scaling successes like LLaDA, and theoretical advances like SEDD position DLMs as a complementary and potentially superior approach for specific applications.
[Emphasis is mine]
We’re talking about hybrid architectures, and “potentially superior” for “specific applications”. I mean, does that really sound like a field that knows how to create AGI?
Although I’m making this up, it is possible that one key part of the radiologists job that is irreplaceable by AI is taking the image out of the MRI machine on a USB stick and physically sticking it into a cloud-connected computer that has the AI software.















