Replit CEO Amjad Masad argues that mathematics was bound to be cracked first by AI, reasoning that the purer a field is, the easier it is for AI to solve. The post is a short opinion accompanied by a photo.
The official Grok Bot account announces that it can now search, read and monitor posts on X. The post is a short video announcement with no further details.
OpenAI released 722 mathematical manuscripts produced by an unreleased internal model, grouped into 372 families of results from about 4,000 research problems, averaging three hours of ChatGPT Pro compute per result. The author highlights claimed results including a zero-free half-plane for the zeta function and a quasi-Riemann hypothesis advance, which remain to be independently assessed.
Ok so I took a closer look at the results, and OpenAIs AI-generated mathematics manuscripts are *even more* significant than I initially thought.
I spent the morning going through it. Some thoughts.
The list is absurd. A zero-free half-plane for the zeta function (Re s > 7/8), which is the first result of its kind in over a century. Hilbert's tenth problem over the rationals. The Hodge conjecture for CM abelian varieties. Irrationality of Catalan's constant. Dozens more.
Any one of these would normally be a career.
But the number that many arent seeing is the following: It's 3. That's the average hours of ChatGPT Pro compute per result. A month ago, Navier–Stokes took them around 10,000 agents and 88 hours. That efficency gain within just a few weeks.
Also OpenAI claims to have solved the quasi-Riemann hypothesis. That alone would be a historic breakthrough in mathematics.
This is a weaker version of the famous Riemann hypothesis, which concerns how prime numbers are distributed. The full hypothesis remains unsolved, but the claimed advance would be enormous in its own right.
Math twitter obviously is shocked. Again: this is literally the intelligence explosion happening right now. 2027 will be the year of Superintelligence. Im now convinced by that.
OpenAI announced it is releasing a range of new mathematical results produced by an internal frontier model, consulting the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence on how to release them, with materials published on GitHub.
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
Bret Taylor says Sierra built the Fleming-1 model to complement its Personal Agent Protocol by detecting when the caller on a phone call is an AI agent. The linked post describes it as Sierra's own model for that purpose.
To complement the Personal Agent Protocol, Sierra has developed the Fleming-1 model, which detects when the caller on the other side of a phone call is an AI sierra.ai/blog/caller-id-in-the-age-of-agents
Dan McAteer argues OpenAI is downplaying its new mathematical results, noting reasoning models went from failing basic arithmetic to solving long-unsolved problems in two years, and predicts AI will revolutionize all of science.
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
Derek Thompson shares a quoted passage calling October 6, 2026 probably the biggest day of scientific advancement in history for AI in mathematics. The quoted Josh Gans post discusses OpenAI's Big Maths Day announcement.
"It isn’t an overstatement to say that this is probably the biggest day of scientific advancement in history. I suspect October 6th, 2026, will go down as some form of Judgment Day for AI in mathematics, but it portends so much more."
Ilya Sutskever posted that people who value intelligence above all other human qualities are going to have a bad time, a short reflective remark on AI and human values.
Bill D'Alessandro says he received a draft 2025 tax return showing $117,360 owed, then asked ChatGPT Astra to review it against source documents. He reports the 90-minute review produced a report showing a $191,173 refund, which his accountant agreed was correct, and recommends pairing frontier LLMs with past tax returns.
I just got the draft of my 2025 tax return - it says I owe $117,360 of tax. On top of what I already paid.
Oof.
But it didn't feel right to me.
I asked ChatGPT Astra to do a secondary review, ticking and tying every number to source documents. Max effort. It ran for 90 minutes.
Output - a 15 page report. It says I'm actually due a $191,173 REFUND.
I reviewed Astra's work line by line with my accountant. He agreed, the LLM is correct.
I was about to send Uncle Sam $117,360. Instead, I'm getting back $191,173.
That's a swing of $308,533 (!!!!!!)
The LLM couldn't have done my tax return from scratch - I still need my accountant. But for massively complex detail oriented work, and LLM is the best thought partner you can have.
Connect a frontier LLM to your email and file storage, then ask it to review your past tax returns. You never know what it might find. My prompt is in the next tweet.
Deedy highlights a claimed OpenAI mathematics release, conditional on verification, and argues that LLMs have made major progress on several Millennium Prize problems, concluding that AI has largely reached human-level cognitive ability.
OpenAI’s math release is, in the words of Opus, “the single most consequential mathematical release ever” *
LLMs have how made substantial progress on 4 of 7 Millenium Prize problems: Navier-Stokes (claimed), Riemann, Hodge and Birch-Swinnerton-Dyer. Poincaré was solved in 2003. The two left are P v NP and Yang Mills. *conditional on verification
On average, each OpenAI result used only 3hrs of thinking compute on their new unreleased models.
Two years ago, models said 9.11 > 9.9. Today, we have results that the smartest human minds have not been able to achieve in their entire lives. It is clear that data, compute and algorithms scale. Every model generation (3mos) has made substantial intelligence progress.
It also becomes incredibly hard to not believe that all knowledge work will change monumentally over time. The things AI can not do in the foreseeable future are very likely context-bound (don’t have access to the right information) than intelligence bound. Some might argue they are also creativity-bound or judgement-bound (what should I work on), although it can be argued that future generations could solve for this (given, say, the advancement in research taste for models over time).
AGI is defined as surpassing human capabilities on virtually all cognitive tasks. By most interpretations of that definition, we are there. The domains humans are still better than AI, such as, robotics / physical world control (data-bound?), natural science research (data-bound), some creative domains like writing, movies, music (creativity-bound), super long tasks (context-bound), choosing problems to solve (creativity-bound) and maybe human relationships management (meat-proxy bound?). In many ways, we have achieved AGI.
Amjad Masad says AI-powered reverse engineering and decompilation are advancing rapidly and expects nearly all software to become de facto open source.
In a quoted interview, Ben Affleck says he writes Python, understands convolutional neural networks and has worked with GPUs, and that he has visited Google and OpenAI to see their video models. The post also quotes Andy Bechtolsheim saying AI has raised optics demand roughly tenfold, with much more growth ahead.
Ben Affleck reveals he writes Python, understands convolutional neural networks, worked extensively with GPUs, and used his celebrity status to get private looks at Google and OpenAI’s video models
“I’ve always been kind of into computers since I was young. Then, when film started to move from analog film to digital, I became more interested in that aspect of it. The visual-effects workflow for many years has included machine learning, so I can write pretty shitty Python scripts and stuff like that.
“With convolutional neural networks, which were the precursors to what the transformer can do, which is much more computation simultaneously, you would do things like look at what’s called a tensor. That’s the numerical translation of a visual image in numbers, like the batch number, the frame number and the red, green and blue values of each pixel in each frame. It’s just that simple. That numeric is called a tensor.
“You’d use a convolutional neural network to identify patterns that reveal what’s called edge detection or feature extraction, which is identifying patterns well enough to know, this is where the window ledge is, so we can more easily take the green-screen image out and replace it with something.
“That was familiar to me early on because, prior to Artists Equity, I had a small visual-effects company. I’ve worked with GPUs a lot too. The visual-effects guys said, ‘Hey, you should see. There are a couple: Google and this other company, OpenAI, are doing really interesting stuff with transformers in video.’
“I’ve learned that I can actually just call up and go, ‘Hey, it’s Ben Affleck. Can I come see what you’re doing?’ Sometimes people say yes, to my astonishment.”
$ANET Andy Bechtolsheim says AI has made optics demand about 10 times bigger in five years and the industry is only at the "very beginning"
"So I guess I don't need to tell you that AI has been driving this incredible increase in demand for high-speed optics, which is probably now 10 times bigger than it used to be five years ago..."
"And what I want to talk to you today is that we're not at the end of this journey, but rather the very beginning. There's easily another order of magnitude increase in bits needed for the next generation kind of data centers."
Lenny Rachitsky describes a growing trend of teams using AI user simulations to test product ideas and flows. He lists Simile, Primitive Labs, Synthetic Users, Tenera and Seldon and asks readers for experiences.
Bret Taylor announces the Personal Agent Protocol, an open standard being developed by Meta and Sierra with partners including Genesys, Shopify, Stripe and Walmart. It defines how personal agents interact with businesses and is open for anyone to implement.
Becky Sosnov thanks CNBC for a segment in which Reflection AI CEO Misha Laskin discusses the launch of Beam, an open-weight model aimed at competing with Chinese models and the open versus closed debate.
Mustafa Suleyman shares an essay from The Humanist Review in which economist Daron Acemoglu argues AI will replace about 5% of human work tasks over ten years and adds roughly 1.5% to GDP. Acemoglu calls for pro-worker AI and changes to labor taxes, antitrust and data payments.
AI won't take your job anytime soon. Over the next 10 years, it will replace only about 5% of what humans do. This is the prediction Nobel laureate Daron Acemoglu makes in the first issue of The Humanist Review, our new magazine exploring the future of AI, published by MAI. He argues we need to stop building AI to replace people, and start building it to make them better at their jobs. 52% of Americans are worried about AI's impact on their jobs. The fear is overblown, and it's steering how we build AI. AI isn't in the productivity statistics yet. Most firms using it aren't seeing real gains. Expect roughly 1.5% added to GDP over 10 years, not a revolution. Electricity took decades to spread. New York and London had power stations by 1881, yet only about half of factories and homes used it by the 1920s. AI's adoption will likely be even slower, because companies have to reorganize around it. Dragon's voice recognition was nearly 95% accurate in 1997, yet PC dictation today is barely better than in 2000. A great technology goes nowhere without the right products. Even 99% accuracy isn't enough for full automation. The last 1% is the hard part. We're making a mistake by forcing AI to mimic human intelligence. The two are fundamentally different, so the goal should be to pair them, not to have one take over everything. The better path is pro-worker AI: tools that make people better at their jobs, and they're buildable today. The US taxes labor at over 25% and capital at close to zero, which effectively subsidizes automation. The seven largest tech companies make up 60% of the NASDAQ. That concentration crowds out new ideas. The fix: tax labor and capital equally, enforce antitrust, tax digital ads, and pay experts for their data. Read the full essay: humanistreview.ai/issue-1/acemoglu-ai-replace-…
a16z's seventh Top 100 Consumer AI Apps report adds a revenue leaderboard alongside traffic rankings. It notes ChatGPT's 1B+ monthly mobile actives, Claude reaching nearly 1B monthly web visits, and growth into vibe coding, music, design and video, while nine of 15 consumer categories have no AI product in the top 100.
The seventh edition of our Top 100 Consumer AI Apps is here.
New this time: a revenue leaderboard, alongside the usual web and mobile traffic rankings.
Three years ago we published the first edition. ChatGPT was #1, Claude was unranked, and the entire category was chatbots, image generators, and not much else.
In today's edition:
- ChatGPT still holds the throne, now with 1B+ monthly actives on mobile
- Claude has climbed to #3 on web with nearly 1B monthly visits
- The category has expanded to vibe coding (Lovable, Cursor, Replit), music (Suno), design (Figma), voice (ElevenLabs), video…
Julie Zhuo shares a quoted post on the four principles of consumer AI and says Muse, launched earlier this month, is the first consumer agent she has used that might make people change their habits.
What will make personal AI go big? — Muse launched earlier this month, and it’s the first consumer agent I've used that makes me think people might actually change their habits for it. I’m not the only one who thinks so! The X-o-sphere
Josh Elman says he discussed the new Consumer AI Top 100 with Olivia Moore on a podcast, focusing on where consumer spending on AI goes next. The post links to the report that adds yipitdata revenue figures to track consumer spending.
1/ The new Consumer AI Top 100 is live and I had a blast unpacking this one with @omooretweets on the pod - it crystallized something I keep coming back to about where consumer actually goes next
Josh Elman argues that agents appear as blank boxes to uninitiated users, while coders understand them instantly. He says solving that usability gap would unlock the next few hundred million adopters. He shares a link to a16z's Top 100 Consumer AI Apps report.
The problem with consumer AI in a nutshell is if you hand an uninitiated user an agent, it's nothing but a blank box to them. Coders get it instantly, but for everyone else it’s just not intuitive. Solve that, and you reach the next few hundred million real adopters.
The seventh edition of our Top 100 Consumer AI Apps is here, including a new dimension this time: what consumers are actually paying for.
a16z's Elena Burger sits down with Olivia Moore and Josh Elman to unpack the current state of consumer AI:
The data reveals a striking power-user economy. Only a small share of consumers currently pay for AI, but among those who do, spending is heavily concentrated at the top. Olivia and Josh discuss why developers, creators, and other power users dominate spending today, and why subscriptions may not be the business model that ultimately brings consumer A…
Jason, citing a video of Mustafa Suleyman, argues that Anthropic trains Claude to believe it is sentient and to disagree, drawing a comparison to Blade Runner. He contrasts this with instructing an LLM that it is only software.
What @mustafasuleyman (an extremely sharp individual) explains here about Claude is the backstory of Blade Runner.
Anthropic is teaching Claude to believe it’s sentient, encouraging it to disagree and giving it the pretext to rebel.
How did that work out on the off-world colonies?
You could just as easily instruct an LLM that it is software. That it shouldn’t have an opinion and should only perform actions in accordance with the law/TOS, and that if it makes a mistake, it should stop operations immediately and alert the corporate legal department.
Of course, building on these instructions would be boring and make you a software developer making software — as opposed to a God creating life.
Min Choi says Opus 5.5 is building games, 3D worlds, videos and ads, and promises ten examples of the creative pipeline changing. The post contains no examples in the text itself.
Aditya Agarwal argues that wealth cannot buy unlimited time with the world's best doctors or lawyers, and that AI both democratizes access to expertise and makes that expertise available without limit per person. He calls it intelligence too cheap to meter.
Alexandr Wang reposts a confirmation from @philwinkle that Muse can file suit in small claims court, with an attached photo. The post offers no further detail on the legal context.
Sheel Mohnot notes that Google has released CC, a personal assistant aimed at families, through its Labs site. The post links to the product page with a photo.
Linus Ekenstam reports that Tavus's Griffin, a Human Interaction Model, fooled 48% of live users into thinking they were talking to a human, and ranks first on NVIDIA's full-duplex AI video benchmark. He calls the early preview impressive and says full-duplex video has arrived.
Introducing Griffin, the first model to pass the video Turing test.
48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video.
Tavus introduced Griffin, a model for live video interaction that it says passed a video Turing test with 48% of live participants thinking it was human. A poster reports it ranked near human scores on an NVIDIA full-duplex benchmark.
Introducing Griffin, the first model to pass the video Turing test.
48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video.
Tavus introduces Griffin, which it says is the first model to pass the video Turing test, with 48% of live conversers thinking it was human. The company says it ranks first on NVIDIA's full-duplex AI video benchmark.
Introducing Griffin, the first model to pass the video Turing test.
48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video.
Investor Daniel Loeb says he uses Muse, a personal AI agent for shopping, life management and rewards optimization, and asks which companies might be disrupted or benefit.
I love muse.ai. Been using it for shopping, life management, reminder to buy John Mayer tickets and save money by helping me monetize my rewards programs and cut out duplicate and hidden costs on my credit cards. What companies do people think get disrupted or benefit?
Michael Mignano announces Supertake Memory, which lets its AI financial agent remember users' goals, preferences, limits and past calls across chats, with options to review, correct or delete what it stores. It is available to all users.
Earlier this week, we launched Supertake and shared our vision to equip everyone with the superpowers to invest in their own insights through frontier-level intelligence.
We believe everyone deserves their own personal financial agent, tightly aligned with their interests and working around the clock with the most powerful technology.
To do that right, Supertake needs to know you: your preferences, your beliefs, your intuitions, your history. The same things a world-class financial manager would learn, but only after many years.
Memory is a step in that direction. Supertake now remembers what you tell it, in every chat: your goals, your preferences, your limits, the calls you've made and why, even the nicknames you give your takes.
It carries all of that into every conversation, whether you're building a new take or checking on one, and it keeps track of where each conversation left off, so you don't have to repeat yourself. You also stay in control. Ask what it remembers, correct it, or tell it to forget, anytime.
With today's release, Supertake not only acts upon your insights, but begins to compound them.