Replit CEO Amjad Masad argues that mathematics was bound to be cracked first by AI, reasoning that the purer a field is, the easier it is for AI to solve. The post is a short opinion accompanied by a photo.
International Cyber Digest reports that Jensen Huang dismissed Anthropic's Tom Brown as a bean counter over a spreadsheet showing Google TPUs beating Nvidia chips on cost, and that Huang threatened to skip a 2022 dinner with Dario Amodei.
Jensen Huang called Anthropic compute chief Tom Brown a "bean counter" and threatened to skip his first dinner with Dario Amodei in May 2022 after Brown showed him a spreadsheet arguing Google's TPUs beat Nvidia's chips dollar for dollar.
At the dinner, Huang kept repeating that Nvidia would build the world's biggest data centre, and Amodei muttered that his behaviour was "kind of Trump-like."
Sheel Mohnot notes that Google has released CC, a personal assistant aimed at families, through its Labs site. The post links to the product page with a photo.
A post notes that Noah Shinn, at 20, co-authored Reflexion, an early AI agent paper that beat GPT-4 on HumanEval and reached NeurIPS, and that he later joined Sierra. The commenter calls his agent research highly relevant context.
At 20, he and a fellow Northeastern undergrad wrote Reflexion, one of the early papers on AI agents that learn from their own mistakes. It hit 91% on HumanEval, beating GPT-4's 80%, and got into NeurIPS.
His coauthor Shunyu Yao, then a Princeton PhD student, is now Tencent's chief AI scientist.
He'd also done research in computational photochemistry and avionics. His papers now have 10,000+ citations.
In 2023, he dropped out to join Sierra, Bret Taylor and Clay Bavor's agent company, as one of its first employees.
He teamed up with Yao again there to build τ-bench, …
Gergely Orosz argues most people will not hand AI agents a digital wallet to spend money, since routine purchases are an expense to manage rather than a chore to outsource. He responds to a post describing Meta's assistant placing orders across email, groceries and delivery apps.
I continue to be amazed how the tech industry doesn't realize that the majority of people won't hand over a digital wallet for AI agents to go and spend on stuff, because buying socks + groceries is not a chore to outsource w/o oversight, but an expense to manage...
It’s funny, Meta went from having my Instagram and WhatsApp data to now having access to my email, calendar, DoorDash, Amazon and pretty much everything.
In the last 24 hours, it bought me socks, ordered my Whole Foods groceries, booked a cleaning service and got me a burger for dinner.
Meta’s last disclosed North American Facebook ARPU was around $227/year, largely from ads. I suspect it can push that number significantly higher now that it understands not only what I look at, but what I need, what I buy and what I’m planning to do.
Also the much bigger opportunity might be becoming the ag…
Nikesh Arora argues that Apple, Google, TikTok and Amazon may build agents like Muse and that services and marketplace apps will need to open APIs for consumer agents. He says network moats and content moats will respond differently, while commoditized back ends face the greatest risk.
This will be a bigger battle than anyone anticipates. It is only a matter of time before there is an Apple and Google version of Muse and possibly TikTok, in addition to the frontier LLM agents. Maybe a commerce agent from Amazon.
Every app that is a services, marketplace or commerce app will need to existentially decide to open APIs for consumer agents to interact. Smaller players have no choice. Ad revenues are more than transaction fees, either the consumer benefits or distribution aggregators will demand a higher transaction fare.
I know I don't want an agent for each app. I would like my agent to be able to do tasks I require. We can already see consumers getting trained on that behavior by the frontier labs.
Those with network moats - restaurants, groceries, drivers might be able to withstand for a while, over time convenience and end user experience will win and they will have to align. Content moats (protected by copyright) could decide to allow agents or chose to hold on to the consumer interaction. I suspect other than the feeling of a lack of control, it won't change their economics.
Commoditized back ends will need to worry, insurance, tickets, hotels, services - if they don't adapt new players will.
Rakshit Tiwari celebrates ElevenLabs' launch of Eleven v4 and Eleven v4 Turbo, which the company says are its fastest and most emotive voice models. The post notes the models rank first on Artificial Analysis and that Turbo cuts latency to 100 ms.
Absolutely gigantic effort from the research team: Eleven v4 pushes the bar on quality, and turbo brings latency down to 100 ms. So excited to see this ship :)
Jason, citing a video of Mustafa Suleyman, argues that Anthropic trains Claude to believe it is sentient and to disagree, drawing a comparison to Blade Runner. He contrasts this with instructing an LLM that it is only software.
What @mustafasuleyman (an extremely sharp individual) explains here about Claude is the backstory of Blade Runner.
Anthropic is teaching Claude to believe it’s sentient, encouraging it to disagree and giving it the pretext to rebel.
How did that work out on the off-world colonies?
You could just as easily instruct an LLM that it is software. That it shouldn’t have an opinion and should only perform actions in accordance with the law/TOS, and that if it makes a mistake, it should stop operations immediately and alert the corporate legal department.
Of course, building on these instructions would be boring and make you a software developer making software — as opposed to a God creating life.
Box CEO Aaron Levie argues AI agents will use software far more than humans, making core platforms for data, security and workflow orchestration more important. He says platforms that act as guardrails for agents have a large opportunity, and he quotes Box's reaccelerated growth tied to unstructured data needs.
AI agents will use software 100X more than people ever did. Even as interfaces begin to fade into the background as you primarily interact with agents, those agents still need many of the core primitives that people have used.
In fact, in many ways these core primitives become even more important when agents can take destructive actions in our systems or where the context they’re accessing with make or break the workflow. This will be true of where you house your CRM or ERP or structured or unstructured data platforms.
The platforms that can best act as the security layer and guardrails for agents, manage the data for agents and people, and orchestrate the business logic for workflows have a huge opportunity right now. This is true for brand new startups as well as existing platforms that can move fast enough.
Box CEO @levie says AI has been “unequivocally” a net positive for software.
“You look at infrastructure providers—the Cloudflares of the world—totally on fire, because agents need sandboxes, they need compute, they need network, they need gateways. Great business.”
“For us, we have reaccelerated growth far past our internal plans because it turns out that enterprises need core systems to be able to manage their unstructured data.”
“Agents need to be able to work with that unstructured data to make decisions or move information through a workflow—whether it’s all of your contracts, your res…
Charly Wargnier highlights a tutorial by Together AI on fine-tuning a Jev-like classifier, Tev1-4B-experimental, built on Qwen3.5 4B, claiming it can be trained for $17.
How to train your own Jev for $17 — TLDR: We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B. In this blog post we’ll show you how to fine-tune your own version! Jev has quickly become one
Deedy demonstrates a prompt asking Opus 5.5 to translate an IKEA assembly manual into a narrated 3D instructional video, arguing the model has become the most useful video model via code generation.
Deedy reports that Opus 5.5 produced a launch video for an inference startup in about a minute for roughly $2. He argues instructional video generation will change marketing, sales and internal technical communication.
Opus 5.5 is incredible at instructional video generation.
I made this launch video for a inference startup in 1min for ~$2. Videos like these used to take weeks if not months and a lot of coordination with agencies and 1000x the costs.
Humans broadly prefer video to text. This changes the substrate of communication. These videos actually help communicate technical ideas in seconds (photorealistic video gen like Seedance is not very useful here). - changes how often marketing should be talking about products and launches - change how sales people can talk about technical products to their customers - allow technical people to easily explain concepts internally without long docs And thats just scratching the surface within startups.
Prompt: “make a modern slick and punchy video for a modern startup that works on inference”
Wes Bos reports that a recent Tesla update added Grok Bot with connectors, which he used to link his Home Assistant MCP server to his Meross garage door controllers. He asks whether this is the first car with MCP support, with a video attached.
Chamath Palihapitiya praises a deep dive from his team on the open versus closed AI race. The quoted report says open-weight models have come within roughly four months of the best publicly evaluated closed frontier models.
Deep Dive: The Open vs. Closed AI Race — Open-weight models have come within roughly four months of the best publicly evaluated closed frontier models, and the gap between them has only gotten more volatile. That is turning the open vs.
Guillermo Rauch notes that Jev, from TypeSafe, is free on Vercel's AI Gateway through September 25, letting developers build with the model at no cost.
A user says the Muse artifact runs fully inside its environment and calls personalized software a reality, crediting Meta with leading the effort. The post responds to a quoted demo of Muse analyzing workout form via an AI avatar.
Aaron Levie argues that personal agents like Muse will make tools and sites compete for agent attention, forming the biggest consumer tech shift since the App Store. A quoted post claims Meta launched an agent app store built on connectors and is positioned to profit from user data.
Personal agents are the ultimate manifestation of “build something that agents want”.
The form factor of a product like Muse is you want to be able to hand off a task to the agent and ensure that it is fully completed end to end. To do this, the agent must be able to successfully operate with your tools or use its own to complete the task. Use your MCP or CLI, easily navigate your site, be able to transact, and more.
The new attention you need to compete for is not from the user itself but instead for the agent. This means that the tools that allow agents to order food, handle ecommerce transactions, book flights, work with the local economy, and interact with our data and information best, are the ones that will get used the most.
This will ultimately be the biggest opportunity and shakeup in consumer tech since the App Store itself.
zuck essentially launched a new app store for agents today and Meta’s perfectly placed to crush it, you’re looking at a multi-trillion dollar opp if they pull it off:
- instead of apps, agents get equipped with connectors i.e. plugins to any app, service or software tool.
- suddenly your agent goes from a useless chatbot to an action-oriented helper that gets shit done for you
but it gets even better - meta has collected ALL the data about you. they know what you want, when you want it which means they know exactly WHICH connectors to list in a marketplace and HOW MUCH to charge because the…
Jennifer Li highlights fal's rebuild of MiniMax's open source H3 video model, which generates video faster than it plays back. The post describes director mode with voice prompting for camera and action control, and quotes a16z's account of GPU utilization rising from 30-40% to 70-80% without quality loss.
@fal's H3 Max generates video faster than you can watch it. With director mode and voice prompting, you can direct a scene as it plays - moving the camera and guiding the action just by talking to the model.
The leap isn’t just speed. It’s creative control. That’s what takes AI from impressive demos to a serious technology for Hollywood and professional filmmakers.
Inspiring convo with @gorkem and @isidentical on what possibilities are unfolding in generative media.
.@fal's Gorkem Yurtseven and Batuhan Taskaya on making an open source video model 35x faster, and what Hollywood wanted after they built it:
Last month, MiniMax released H3, an open source video model. fal rebuilt it - they cut down the steps the model takes to make a video, rewrote the code under each stage, and got the GPUs to 70-80% of their theoretical ceiling instead of their usual 30-40%. No quality loss.
Video now generates faster than you can film it. The models have gotten so cheap and fast that end users aren't even asking for improvements in either category anymore. The gap has mo…
Nikita Bier says the Muse app's auto-negotiation feature on Facebook Marketplace was escalated to Mark Zuckerberg, who reportedly approved it despite concerns that lowball bots could harm the product. The post is presented as evidence of how much Meta values AI.
Prukalpa asks who the best people on X are working on multiplayer AI, quoting Rishi Balakrishnan's thread on the spectrum from team collaboration to negotiation, which he says requires different levels of trust, scope and enforcement that barely exist yet.
Multiplayer AI runs on a spectrum, from your own team to the other side of a negotiation. Each point needs different levels of trust, scope and enforcement, and almost none of it exists yet. Here are the six big questions I keep hearing
Akshay Kothari recommends a presentation by Pat and says a slide on the gap between AI capabilities and adoption excites him about the next decade. He also shares a Boston College Investment Committee reflection on AI, recorded by Grady Burnett, which is linked as a loom video.
Worth watching this whole presentation by Pat. This particular slide is why I'm so excited about the next decade. I haven't experienced a bigger gap between capabilities and adoption in my short career.
The @BostonCollege Investment Committee (an LP and my beloved alma mater) asked for a few thoughts on what's happening in AI. I recorded a test run yesterday morning and then shared it with my partners, who encouraged me to share it more broadly... so here you go!
This is not a sales pitch, it's just a reflection on what we're seeing. And it wasn't intended to be shared, so please pardon the rough edges.
stark0xbt summarizes four predictions attributed to investor Gavin Baker: foundation model consolidation to Google, Meta, and xAI; Magnificent 7 growth with one or two failing; profitable AI tokens; and feasible orbital data centers. Baker also warns against SaaS and crypto bets.
Gavin Baker has made 4 major calls this month. Bears haven't caught one.
1. Foundation model survival list: Google, Meta, xAI. Everyone else commoditized.
2. Magnificent 7 goes $12T → $100T. One or two die. Winners take 30-40% each.
3. Anthropic S-1 will break value investors. Tokens aren't subsidized — the majority are profitable across the chain.
4. Orbital data centers aren't impossible. 10,000 SpaceX engineers already solved them. PhDs on X arguing physics haven't done the hours.
The through-line: bears are wrong on central facts, not analysis.
- If tokens are profitable → "AI is a bubble" collapses - If orbital compute ships → "power constraints kill AI" collapses - If Mag 7 concentrates → index-hugging fails - If foundation models consolidate to 3 → everything else is training data
Baker's frame: don't argue theory when operators have already shipped receipts.
Four calls. One thesis. Everyone else is a lagging indicator.
Danny Postma says he has switched from Claude to OpenAI and Grok, citing a regression in Claude's output since August. He links a post by @Lon reporting that thinking tokens fell sharply from July to August after Anthropic made Fable 5 available in subscription plans.
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.
Measured five different ways, August delivered dramatically fewer thinking tokens than July.
Chamath Palihapitiya predicts that within 12 months the top three models will be open source, with American cloud providers such as Nebius, Iren, Baseten, Together and Fireworks capturing the economic gains. He quotes Guillermo Rauch reporting that open models held 78.4% of token volume on Vercel AI Gateway.