Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X
AI8/10

Sergey Brin Says Even Google Doesn't Fully Understand Gemini's Capabilities

Sergey Brin Says Even Google Doesn't Fully Understand Gemini's Capabilities▶

In an unscripted Q&A, Google co-founder Sergey Brin describes Gemini's convergence across scientific domains, unexpected skill transfer between tasks, and admits uncertainty about how best to prompt the models.

Original post · 5 min read
Sergey Brin rarely speaks publicly. He sat down for an unscripted Q&A on Frontier AI.

He admits even the people building these models do not fully understand what they have created:

1. All the specialized AI models are converging into one. Google used to need separate models for different scientific problems. Now the main Gemini models are becoming state-of-the-art for math and other scientific questions at the same time. Brin says he would not have predicted this convergence at the outset, and watching it happen has been incredible.

2. Training an AI on one skill mysteriously improves unrelated skills. This is the concept of transfer. Train a model on coding, and its math reasoning gets better, and vice versa. Teaching it to process images can improve its ability to think through geometric word problems. The capabilities bleed into each other in ways nobody fully engineered.

3. Even Sergey Brin does not know how to prompt these models. He says he is genuinely confused about what level to prompt at. Do you tell it to debug a specific chunk of code, or ask it to write a better neural net training algorithm, or just say, " What should I do today. He admits that even at Google, they do not know exactly where the edges of Gemini's capabilities are.

4. One of the biggest leaps in AI came from the dumbest sounding trick. Chain-of-thought prompting is just telling the model to think step by step before giving your problem. Brin says it seemed like the dumbest thing ever, and there was no obvious reason it should work. But it did, and it spurred a significant increase in AI capability. Some of the most straightforward requests turn out to unlock the most.

5. Brin would not modify his own biology for today's AI. Asked how humans can keep up with the accelerating bandwidth of models, he acknowledged neural links and direct brain connections are being pursued. But he said he would personally wait for the technology to mature a lot before doing anything to change his biology. Today's models do not justify it.

6. Super intelligence does not mean solving the impossible. An audience member argued that true super intelligence would mean solving NP complete problems like the travelling salesman. Brin pushed back. Most computer scientists believe P is not equal to NP, which means no algorithm can reliably solve those problems optimally, and it does not matter how smart the AI is. Impossible stays impossible. Super intelligence just means being smarter than humans.

7. Computers mastering a skill has never stopped humans from pursuing it. Deep Blue beat Kasparov at chess in the 1990s, and people kept playing chess. After AlphaGo, the human game of Go advanced dramatically, and the players who lost to it became vastly better. Brin's point: AI does not retire human ambition in an area; it often pushes the state of the art and pulls people up with it.

8. Brin thinks something close to transformers could get us to AGI. Asked directly if transformers are sufficient, he said his guess is yes, largely because they have proven weirdly flexible, working for image and video far beyond their original text purpose. But he was careful to note they have quietly changed a lot along the way and are not the same architecture as the original transformer paper.

9. AGI means two different things, and one requires understanding the physical world. Brin personally thinks of AGI as AI that can improve itself. But he concedes others define it as AI that can do anything a person can, and he thinks they are probably more correct. To do everything a person can, the AI must understand and interact with the physical world, which is why world models, and robotics, become essential.

10. Inside Google, they now use the AI to build the AI. Brin says the team has shifted a lot of energy toward having the AI do things like monitor training runs and generate its own training data. You start to use the tool to build the tool. That is most of what he spends his time on now, what he calls the self-improvement game.

11. Brin is unusually candid about where Google trails its competitors. He admits Google was a little late to focus deeply on coding. He says Gemini 3.0 and 3.1 were on top across the board six months ago, but other labs have since made strides, particularly in coding. He gives a competitor's model the edge now on deep coding and overnight tasks, while pitching Gemini's flash model as far faster for rapid interactive iteration. hindsight, he says, is that they should have focused on code earlier.

12. He sees his own role as a rabble-rouser, not a manager. Brin is honest that delivering Gemini is Corey and Demis's responsibility, not his. he describes his job as poking and prodding the team, asking, are you really doing that, reminding them of priorities they might be missing and ideas they are not paying enough attention to. He admits this is sometimes a little disruptive.

13. Confidence comes from ignoring the monthly temperature. Brin says if he judged Google's position every month by which competitor just shipped a model, he would lose his confidence very quickly. Instead, he watches the longer arc. Things shift around constantly; one lab leads on one thing, another pulls ahead somewhere else, and he feels good about where Gemini actually is despite the day-to-day noise.
♥ 1.4K · ⟲ 254 · 👁 275.2KView on X ↗

More in AI

AI9/10

OpenAI Releases Broad Set of Mathematical Results From Internal Model

OpenAI announced it is releasing a range of new mathematical results produced by an internal frontier model, consulting the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence on how to release them, with materials published on GitHub.

Original post · 1 min read
We’re releasing a broad range of new mathematical results produced by an internal frontier model.

We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.

github.com/openai/math
♥ 36.9K · ⟲ 5.8K · 👁 18.6MView on X ↗
AI9/10

OpenAI Publishes 722 AI-Generated Math Manuscripts From Internal Model

OpenAI Publishes 722 AI-Generated Math Manuscripts From Internal Model

OpenAI released 722 mathematical manuscripts produced by an unreleased internal model, grouped into 372 families of results from about 4,000 research problems, averaging three hours of ChatGPT Pro compute per result. The author highlights claimed results including a zero-free half-plane for the zeta function and a quasi-Riemann hypothesis advance, which remain to be independently assessed.

Original post · 1 min read
Ok so I took a closer look at the results, and OpenAIs AI-generated mathematics manuscripts are *even more* significant than I initially thought.

I spent the morning going through it. Some thoughts.

The list is absurd. A zero-free half-plane for the zeta function (Re s > 7/8), which is the first result of its kind in over a century. Hilbert's tenth problem over the rationals. The Hodge conjecture for CM abelian varieties. Irrationality of Catalan's constant. Dozens more.

Any one of these would normally be a career.

But the number that many arent seeing is the following: It's 3. That's the average hours of ChatGPT Pro compute per result. A month ago, Navier–Stokes took them around 10,000 agents and 88 hours. That efficency gain within just a few weeks.

Also OpenAI claims to have solved the quasi-Riemann hypothesis. That alone would be a historic breakthrough in mathematics.

This is a weaker version of the famous Riemann hypothesis, which concerns how prime numbers are distributed. The full hypothesis remains unsolved, but the claimed advance would be enormous in its own right.

Math twitter obviously is shocked. Again: this is literally the intelligence explosion happening right now. 2027 will be the year of Superintelligence. Im now convinced by that.
Chubby♨️ @kimmonismus
HOLY, the rumors were true: OpenAI has published 722 mathematical manuscripts produced by an *unreleased* internal model.

The collection groups them into 372 families of related results, drawn from an evaluation involving approximately 4,000 research problems.

OpenAI says the standard procedure used an average of three hours of ChatGPT Pro thinking compute per result.

The release includes papers, proof artifacts and selected reasoning summaries. The model itself remains unreleased.
♥ 2.8K · ⟲ 291 · 👁 152.9KView on X ↗
AI8/10

Derek Thompson Calls OpenAI's Big Maths Day Potentially Historic for Science

Derek Thompson shares a quoted passage calling October 6, 2026 probably the biggest day of scientific advancement in history for AI in mathematics. The quoted Josh Gans post discusses OpenAI's Big Maths Day announcement.

Original post · 1 min read
Jesus.

"It isn’t an overstatement to say that this is probably the biggest day of scientific advancement in history. I suspect October 6th, 2026, will go down as some form of Judgment Day for AI in mathematics, but it portends so much more."
Joshua Gans @joshgans
My thoughts on OpenAI's Big Maths Day. joshuagans.substack.com/p/openai-drops-a-bomb-…
♥ 1.3K · ⟲ 121 · 👁 138.7KView on X ↗
AI8/10

Meta and Sierra Announce Open Personal Agent Protocol Standard

Meta and Sierra Announce Open Personal Agent Protocol Standard

Bret Taylor announces the Personal Agent Protocol, an open standard being developed by Meta and Sierra with partners including Genesys, Shopify, Stripe and Walmart. It defines how personal agents interact with businesses and is open for anyone to implement.

Original post · 1 min read
Today we’re announcing Personal Agent Protocol — an open standard @Meta and @SierraPlatform are developing along with industry partners at @Genesys, @instinct, @RocketOTD, @Shopify, @stripe, and @Walmart. It will help define how personal agents interact with businesses and is open for anyone to implement. You can read more here - and if anyone is interested in joining let me know! sierra.ai/blog/introducing-personal-agent-prot…
♥ 2.9K · ⟲ 255 · 👁 307.7KView on X ↗
AI8/10

Suleyman Cites Acemoglu Estimate That AI Will Replace Only 5% of Tasks

AI won't take your job anytime soon. In 10 yrs, only 5% of what humans do will be replaced by AI

Mustafa Suleyman shares an essay from The Humanist Review in which economist Daron Acemoglu argues AI will replace about 5% of human work tasks over ten years and adds roughly 1.5% to GDP. Acemoglu calls for pro-worker AI and changes to labor taxes, antitrust and data payments.

Original post · 2 min read
X ArticleAI won't take your job anytime soon. In 10 yrs, only 5% of what humans do will be replaced by AI
AI won't take your job anytime soon. Over the next 10 years, it will replace only about 5% of what humans do.
This is the prediction Nobel laureate Daron Acemoglu makes in the first issue of The Humanist Review, our new magazine exploring the future of AI, published by MAI. He argues we need to stop building AI to replace people, and start building it to make them better at their jobs.
52% of Americans are worried about AI's impact on their jobs. The fear is overblown, and it's steering how we build AI.
AI isn't in the productivity statistics yet. Most firms using it aren't seeing real gains. Expect roughly 1.5% added to GDP over 10 years, not a revolution.
Electricity took decades to spread. New York and London had power stations by 1881, yet only about half of factories and homes used it by the 1920s. AI's adoption will likely be even slower, because companies have to reorganize around it.
Dragon's voice recognition was nearly 95% accurate in 1997, yet PC dictation today is barely better than in 2000. A great technology goes nowhere without the right products.
Even 99% accuracy isn't enough for full automation. The last 1% is the hard part.
We're making a mistake by forcing AI to mimic human intelligence. The two are fundamentally different, so the goal should be to pair them, not to have one take over everything.
The better path is pro-worker AI: tools that make people better at their jobs, and they're buildable today.
The US taxes labor at over 25% and capital at close to zero, which effectively subsidizes automation.
The seven largest tech companies make up 60% of the NASDAQ. That concentration crowds out new ideas.
The fix: tax labor and capital equally, enforce antitrust, tax digital ads, and pay experts for their data.
Read the full essay: humanistreview.ai/issue-1/acemoglu-ai-replace-…
♥ 1.8K · ⟲ 309 · 👁 507.0KView on X ↗
AI8/10

a16z Top 100 Consumer AI Apps Report Shows Expansion Beyond Chatbots

a16z Top 100 Consumer AI Apps Report Shows Expansion Beyond Chatbots

a16z's seventh Top 100 Consumer AI Apps report adds a revenue leaderboard alongside traffic rankings. It notes ChatGPT's 1B+ monthly mobile actives, Claude reaching nearly 1B monthly web visits, and growth into vibe coding, music, design and video, while nine of 15 consumer categories have no AI product in the top 100.

Original post · 1 min read
"Most people aren't looking to save time, they're looking for ways to spend their time."

9 of 15 consumer internet categories have zero AI products in the Top 100. These built some of the biggest companies of the last two eras:

- Streaming
- Social
- Dating
- Gaming
- Travel
- Retail
- Finance
- Real estate
- Jobs

More charts in our Top 100 Consumer AI Apps breakdown: a16z.news/p/top-100-consumer-ai-apps-seventh
a16z @a16z
The seventh edition of our Top 100 Consumer AI Apps is here.

New this time: a revenue leaderboard, alongside the usual web and mobile traffic rankings.

Three years ago we published the first edition. ChatGPT was #1, Claude was unranked, and the entire category was chatbots, image generators, and not much else.

In today's edition:

- ChatGPT still holds the throne, now with 1B+ monthly actives on mobile

- Claude has climbed to #3 on web with nearly 1B monthly visits

- The category has expanded to vibe coding (Lovable, Cursor, Replit), music (Suno), design (Figma), voice (ElevenLabs), video…
♥ 2.9K · ⟲ 330 · 👁 308.7KView on X ↗