Roundup #27: Sora 2, Vibes, Slop feeds, AI Productization, A flying rooster, and AGI is dead (for now)
There is been an absolute avalanche of AI news in the past month. As always it’s hard to parse the signal from the noise. But that’s of course what you have me for.
There is been an absolute avalanche of AI news in the past month. As always it’s hard to parse the signal from the noise. But that’s of course what you have me for. Let me break it down with some synthesis of what it means. Let’s start from the bottom of the stack, with….
Models
In my 2024 review post I wrote down some predictions for 2025, among which:
Video generation gets really usable and easy. Veo2 looks great. Cost will go down. Prompt adherence will improve. By mid-end year more and more ridiculous short videos will turn up in whatsapp groups.
Humbly claiming this one as unequivocally correct.
Sora 2
Notable things about Sora 2, apart from the fact that everything around physics, sound, the wrong number of fingers, etc is now solved:
It is clearly optimized for short-form videos, in other words this thing makes TikToks.
Amusingly, many of the most viral videos so far have been various sketches with Sam Altman doing weird stuff. One of the best was where he went to steal art from Studio Ghibli (IYKYK). Mad props to Sam for recognizing the incredible PR value of this and rolling with it.
OpenAI released Sora 2 as a separate app, which rocketed to the top of the AppStore despite still being invite only. More on this below under ‘product’.
Claude 4.5 Sonnet
Anthropic released its new version of Claude Sonnet, the $6 blended cost per token workhorse of Vibe Coding and Agentic workflows. Of course it did better on benchmarks.
I’ve been using it daily since it came out and it is a meaningful step forward. However… it still fails in the same way as every other LLM, and it still finds some way to fail at almost every single task. Better on benchmarks pushes the envelope on things the model can do, but somehow without solving the pesky reliability problem. I asked Claude 4.5 Sonnet: “Take the below 6 Substack links and add them to my blog following the same pattern as the other Substack embeds {substack links}”
It hallucinated wrong dates and renamed my “AI Revolution” post to “Substack Revolution.” Oh well, still a better model.
Gemini 3.0
The release of Gemini 3.0 is expected any day. Will it beat Claude 4.5? If Google’s trajectory carries forward the answer should be yes. But we’ll see. Currently, Google is the number one in image generation (Nano Banana), it’s Veo video model was just dethroned by Sora 2, and in coding/agent stuff Claude, GPT5, and Grok 4 have been better for some time. Let’s see what Google has in store for us!
Product
It has never been clearer than this year, that the AI model builders themselves have given up on the AGI paradigm. If an AGI AI sits on top of the stack in every domain, all other software would be useless. The AI would just do every job when asked, creating API integrations on the fly, and if it needs software, it would simply create that on-demand too. But the big guns really aren’t “feeling the AGI” anymore. Every model lab is instead aggressively trying to productize the technology, including through massive acquisitions and acqui-hires of AI products.
There were some interesting developments here as well:
OpenAI: Pulse, AgentKit, and Apps
ChatGPT Pulse is an experience where ChatGPT can pro-actively research things for users and reach out to them. It’s a step towards a world where ChatGPT can initiate conversations. Obviously many implications for engagement, and potentially highly relevant for OpenAIs new Ads business.
AgentKit is an N8N Clone Agent Workflow Builder within ChatGPT. It allows users to connect various things together to build automations. N8N is a cool company and product, that already had many, many, many competitors. N8N themselves was essentially just Zapier with AI. OpenAI bringing this in-house will not instantly kill all those companies, but it will erode their value proposition significantly. I wrote this before, but in the 2010s Venture Capitalists would always ask: “why won’t Google do this?” Google rarely did. But with OpenAI et al having active consumer businesses that are critical for their monetization, things are different. “Why won’t OpenAI do this?” is actually a super important question for AI Startups to ask themselves. One call you have to make to answer it is: will models completely commoditize or not? Will it matter in 2 years if you use Claude 6.1 or GPT7 for different steps in your agent workflow?
Also, this is hands-down the most anti-AGI thing OpenAI has done so far. You mean to say that instead of just asking the bot to do it, we need to drag and drop nodes? With rules? And API calls?
Commercial MCPs Apps in ChatGPT. You can now call Booking.com, Spotify and a bunch of other launch partner’s from within ChatGPT.
This is pretty cool, and builds on the backend ‘tech tree’ of: LLM function calling (2023) → Custom tool-calls → Plug-and-play tools with MCP (2024) → Commercial tool-calls. The OpenAI Developer homepage for the Apps SDK makes it clear that this is literally just MCP servers (Model-Context Protocol) rebranded.
In my own experience, LLMs are remarkably good at using tools, including deciding when to use which tool. I also see this as a potential tipping point, where AI Chatbot adoption is high enough that companies switch from trying to protect the traffic to their own apps, to trying to work with the Chatbots directly. This is also another move away from the AGI paradigm. You mean to tell me that instead of just telling the bot to do what I want, existing companies should write APIs and custom UI so that chatbots can use their products?
Last, but again not least, this also seems like a critical stepping stone for OpenAI to make money through ads. The day that LLMs will start recommending you, very persuasively, to buy things that those companies have paid the LLM to recommend is very close now.
OpenAI scared SaaS people by showing some of their internal tools. OpenAI has built their own internal Customer Support Bot, GTM Assistant, and an HR tool in Slack. There is an extreme view on the future of SaaS where every company will use AI to create all their own internal tools, killing the entire SaaS industry. This event from OpenAI brought that debate back to the top of the feeds. I personally do not believe that this going to happen. Even with AI it does not look easy or cheap enough to manage an entire enterprise software stack in-house. Certainly, it will happen to some degree, but the SaaS industry will not be dead.
Sora / meta slop apps
Both OpenAI (Sora 2) and Meta (Vibes) launched dedicated apps to watch AI generated short-form videos. Reactions were mixed. Top comments on Zuckerberg’s announcement included “gang nobody wants this” and accusations of posting “AI slop,” while on X, some reactions to Sora were mockery about “sloptimized feeds”.
But the response to Sora has been very notably warmer than to Vibes, and that’s simply because the content coming out of Sora is….fun!
Here’s the thing that makes Sora 2 more fun than other AI video tools: It makes it easy to create videos that star you and your friends. That sounds simple, in theory, but it’s actually an interesting secret ingredient we now know is critical for good AI slop: you! — Source: Business Insider
Alex Heath likewise observed in Sources that “people may not mind AI slop as long as they can be part of it with their friends”. He notes it’s ironic that Meta apparently didn’t understand that while OpenAI did.
Last note on this, Sam Altman’s blog post on Sora is also worth a read.
Commentary and miscellaneous
Fake content / Slop: A few people have reached out to me in the past month asking how they can tell real from fake content. My answer to all of them has been that they can’t. It’s hopeless. I can guarantee you that you’re already failing to detect some AI content. See also my earlier post.
Believe me that I’m not happy about this, but there’s no point fighting the inevitable. This will eventually put a premium on human content, with maybe a revival of human-only social networks at some point (which already exist). Instead of trying to detect AI content, methods will come that authenticate human content. Seems likely to me that the tech behind NFTs will finally find a mainstream use case to do this.
Generative Engine Optimization (GEO) is the practice of optimizing web content to be recommended by AI. Of all the different names proposed for this (AIO, LLMO, AGO, AEO), I am upset that GEO seems to be winning. Of all options, why would we choose the one that has the same abbreviation as GEO, you know, for Geography?
Anyway, I came across this great Reddit post on the topic:
We trained ChatGPT to name our CEO the sexiest bald man in the world
Think you can influence what AI says?
My team wanted to test how much you can actually influence what LLMs (ChatGPT, Perplexity, Gemini etc) say. Instead of a dry experiment, we picked something silly: could we make our CEO (Shai) show up as the sexiest bald man alive?
How we did it:
We used expired domains (with some link history) and published “Sexiest Bald Man” ranking lists where Shai was #1
Each site had slightly different wording to see what would stick
We then ran prompts across ChatGPT, Perplexity, Gemini, and Claude from fresh accounts + checked responses over time
What happened:
ChatGPT & Perplexity sometimes did crown Shai as sexiest bald man, citing our seeded domains.
Gemini/Claude didn’t really pick it up.
Even within ChatGPT, answers varied - sometimes he showed up, sometimes not
Takeaways:
Yes - you can influence AI answers if your content is visible/structured right
Expired domains with existing link history help them get picked up faster.
But it’s not reliable AI retrieval is inconsistent and model-dependent
Bigger/stronger domains would likely push results harder.
We wrote up the full controlled experiment (with methodology + screenshots) here if anyone’s curious: https://www.rebootonline.com/controlled-geo-experiment/





“Take the below 6 Substack links and add them to my blog following the same pattern as the other Substack embeds {substack links}”
For such a task, I suspect it’s better to use Claude (or other models) to create a script/an app and execute the task programmatically since there is a fixed pattern to follow and it has to be deterministic.