Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X
AI8/10

GPTZero CTO Explains How AI Text Watermarking Works

Alex Cui, CTO of GPTZero, explains the KGW-style green-list watermarking used by Anthropic, Google and OpenAI, covering generation, detection, and whether paraphrasing can defeat it. He responds to news that Claude models will carry invisible watermarks.

Original post · 5 min read
Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can be defeated.

Almost all forms of watermarking that are fast and cheap enough for a frontier lab have the same formula, following the KGW method:

In generation:

1. Let's say you've generated n tokens so far. Take those n tokens + a secret key to generate a random hash
2. Use that hash to randomly reweight the probabilities for the n+1 token, and then sample from that new distribution. In the simple case, you could split 50% of all English words into a green or red set based on your hash, and boost the probability of words in the green set.

For watermark detection:

1. For each token, see if it was in the green or red set.
2. To do this, recreate the hash based on the secret key and the text preceding the current token. Then, recreate the green and red set of words.
3. Once you've checked all the words in the text, if the next token is selected disproportionally from the green set more than 50% of the time, you claim the text has the watermark.

I can tell you want to ask the following:

1) Isn't it easy to mess up the hash if you paraphrase the text? The answer is mostly yes, however, you can use a statistical model to get your hash instead of a deterministic function (SIR, Adaptive Watermark). Since the entire watermark is probabilistic, this is fine.

2) Doesn't this make the text much worse? The answer is yes, it does - Yes, it does – but for most people, it's imperceptible (Google claims in human feedback study with 20,000 texts), since there are exponentially many ways to write the same paragraph. DiPmark does something more sophisticated to avoid shifting the text distribution on average. Of course, watermarks fail on short text or highly predictable texts like "2+2=4".

3) Shouldn't it be easy to figure out the green and red sets? The answer is no. You would need an exponentially large number of samples from the watermarker to reconstruct those sets exactly, but it's a risk if the detector is open to the wild (Watermark Stealing)

Still, there are couple challenges that a frontier lab needs to overcome:
1. Their watermark needs to work token-by-token because they are streaming their text to users. Many watermark methods plan sentences or paragraphs at a time, or change the text after its entirely written, in order to make their watermark robust to paraphrasers, and a frontier lab cannot afford to do this yet (SemStamp, PostMark)
2. If the secret key leaks, the watermark is busted. To avoid a large blast damage from this, you need to have a couple secret keys in rotation.
3. There are some texts, like code, that cannot be arbitrarily changed, otherwise the code will break. In those cases, the watermark needs to selectively change words in parts of the text that can tolerate synonyms (i.e. like variable naming) - see SWEET, EWD, Invisible Entropy.
4. They will need to educate their users on how to deal with false positives and false negatives of a detector, which is a big challenge (one we put a lot of effort into)

So, how do I see this playing out in the next 6 months?
1. If Anthropic releases the watermark detector publically, I think they defeat their own watermark. People find reliable watermark removal strategies by testing against Anthropic (AI detectors like GPTZero have an advantage here because they can train against these adversaries once they become popular).
2. If they keep the detector private to the government, like Google has done, it's "safer". However, there are some papers showing trained approaches that work robustly to zero-shot break watermarks without any data, simply because they try to write the text just like a human (Zhang et al. 2024, Watermarks in the Sand). Also, making your detector makes it battle-tested and stronger long-term (my experience).
3. In my testing, the watermarks don't survive intense paraphrasing (especially if you combine word choice and syntax attacks), or human text substitution (rewrite your AI text by plagiarizing human authors). The free paraphrasers I've tried have quickly bypassed Google Deepmind's SynthId for what it's worth.
4. All-in-all, frontier labs are likely okay with this because they expect most users to not attack the watermark, and also because they + European regulators likely don't care past a certain point - its good enough.
5. Overall, I think users of frontier LLMs will not really care about this, because 1) they don't realize watermarks are there, 2) EU will force everyone to conform, 3) this seems more like regulatory hoop-jumping than an earnest effort from frontier labs to expose LLM use

Lastly, people's first concern shouldn't be watermarking, it should be AI detectors!

If you're posting, "its not X, its Y!!", I don't think the watermark is going to make a difference :)
NIK @ns123abc
🚨 JUST IN: Claude models will now have invisible watermarks embedded in ALL text, and ALL metadata attached to files…
♥ 6.0K · ⟲ 791 · 👁 1.3MView on X ↗

More in AI

AI9/10

OpenAI Releases Broad Set of Mathematical Results From Internal Model

OpenAI announced it is releasing a range of new mathematical results produced by an internal frontier model, consulting the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence on how to release them, with materials published on GitHub.

Original post · 1 min read
We’re releasing a broad range of new mathematical results produced by an internal frontier model.

We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.

github.com/openai/math
♥ 36.9K · ⟲ 5.8K · 👁 18.6MView on X ↗
AI9/10

OpenAI Publishes 722 AI-Generated Math Manuscripts From Internal Model

OpenAI Publishes 722 AI-Generated Math Manuscripts From Internal Model

OpenAI released 722 mathematical manuscripts produced by an unreleased internal model, grouped into 372 families of results from about 4,000 research problems, averaging three hours of ChatGPT Pro compute per result. The author highlights claimed results including a zero-free half-plane for the zeta function and a quasi-Riemann hypothesis advance, which remain to be independently assessed.

Original post · 1 min read
Ok so I took a closer look at the results, and OpenAIs AI-generated mathematics manuscripts are *even more* significant than I initially thought.

I spent the morning going through it. Some thoughts.

The list is absurd. A zero-free half-plane for the zeta function (Re s > 7/8), which is the first result of its kind in over a century. Hilbert's tenth problem over the rationals. The Hodge conjecture for CM abelian varieties. Irrationality of Catalan's constant. Dozens more.

Any one of these would normally be a career.

But the number that many arent seeing is the following: It's 3. That's the average hours of ChatGPT Pro compute per result. A month ago, Navier–Stokes took them around 10,000 agents and 88 hours. That efficency gain within just a few weeks.

Also OpenAI claims to have solved the quasi-Riemann hypothesis. That alone would be a historic breakthrough in mathematics.

This is a weaker version of the famous Riemann hypothesis, which concerns how prime numbers are distributed. The full hypothesis remains unsolved, but the claimed advance would be enormous in its own right.

Math twitter obviously is shocked. Again: this is literally the intelligence explosion happening right now. 2027 will be the year of Superintelligence. Im now convinced by that.
Chubby♨️ @kimmonismus
HOLY, the rumors were true: OpenAI has published 722 mathematical manuscripts produced by an *unreleased* internal model.

The collection groups them into 372 families of related results, drawn from an evaluation involving approximately 4,000 research problems.

OpenAI says the standard procedure used an average of three hours of ChatGPT Pro thinking compute per result.

The release includes papers, proof artifacts and selected reasoning summaries. The model itself remains unreleased.
♥ 2.8K · ⟲ 291 · 👁 152.9KView on X ↗
AI8/10

Derek Thompson Calls OpenAI's Big Maths Day Potentially Historic for Science

Derek Thompson shares a quoted passage calling October 6, 2026 probably the biggest day of scientific advancement in history for AI in mathematics. The quoted Josh Gans post discusses OpenAI's Big Maths Day announcement.

Original post · 1 min read
Jesus.

"It isn’t an overstatement to say that this is probably the biggest day of scientific advancement in history. I suspect October 6th, 2026, will go down as some form of Judgment Day for AI in mathematics, but it portends so much more."
Joshua Gans @joshgans
My thoughts on OpenAI's Big Maths Day. joshuagans.substack.com/p/openai-drops-a-bomb-…
♥ 1.3K · ⟲ 121 · 👁 138.7KView on X ↗
AI8/10

Meta and Sierra Announce Open Personal Agent Protocol Standard

Meta and Sierra Announce Open Personal Agent Protocol Standard

Bret Taylor announces the Personal Agent Protocol, an open standard being developed by Meta and Sierra with partners including Genesys, Shopify, Stripe and Walmart. It defines how personal agents interact with businesses and is open for anyone to implement.

Original post · 1 min read
Today we’re announcing Personal Agent Protocol — an open standard @Meta and @SierraPlatform are developing along with industry partners at @Genesys, @instinct, @RocketOTD, @Shopify, @stripe, and @Walmart. It will help define how personal agents interact with businesses and is open for anyone to implement. You can read more here - and if anyone is interested in joining let me know! sierra.ai/blog/introducing-personal-agent-prot…
♥ 2.9K · ⟲ 255 · 👁 307.7KView on X ↗
AI8/10

Suleyman Cites Acemoglu Estimate That AI Will Replace Only 5% of Tasks

AI won't take your job anytime soon. In 10 yrs, only 5% of what humans do will be replaced by AI

Mustafa Suleyman shares an essay from The Humanist Review in which economist Daron Acemoglu argues AI will replace about 5% of human work tasks over ten years and adds roughly 1.5% to GDP. Acemoglu calls for pro-worker AI and changes to labor taxes, antitrust and data payments.

Original post · 2 min read
X ArticleAI won't take your job anytime soon. In 10 yrs, only 5% of what humans do will be replaced by AI
AI won't take your job anytime soon. Over the next 10 years, it will replace only about 5% of what humans do.
This is the prediction Nobel laureate Daron Acemoglu makes in the first issue of The Humanist Review, our new magazine exploring the future of AI, published by MAI. He argues we need to stop building AI to replace people, and start building it to make them better at their jobs.
52% of Americans are worried about AI's impact on their jobs. The fear is overblown, and it's steering how we build AI.
AI isn't in the productivity statistics yet. Most firms using it aren't seeing real gains. Expect roughly 1.5% added to GDP over 10 years, not a revolution.
Electricity took decades to spread. New York and London had power stations by 1881, yet only about half of factories and homes used it by the 1920s. AI's adoption will likely be even slower, because companies have to reorganize around it.
Dragon's voice recognition was nearly 95% accurate in 1997, yet PC dictation today is barely better than in 2000. A great technology goes nowhere without the right products.
Even 99% accuracy isn't enough for full automation. The last 1% is the hard part.
We're making a mistake by forcing AI to mimic human intelligence. The two are fundamentally different, so the goal should be to pair them, not to have one take over everything.
The better path is pro-worker AI: tools that make people better at their jobs, and they're buildable today.
The US taxes labor at over 25% and capital at close to zero, which effectively subsidizes automation.
The seven largest tech companies make up 60% of the NASDAQ. That concentration crowds out new ideas.
The fix: tax labor and capital equally, enforce antitrust, tax digital ads, and pay experts for their data.
Read the full essay: humanistreview.ai/issue-1/acemoglu-ai-replace-…
♥ 1.8K · ⟲ 309 · 👁 507.0KView on X ↗
AI8/10

a16z Top 100 Consumer AI Apps Report Shows Expansion Beyond Chatbots

a16z Top 100 Consumer AI Apps Report Shows Expansion Beyond Chatbots

a16z's seventh Top 100 Consumer AI Apps report adds a revenue leaderboard alongside traffic rankings. It notes ChatGPT's 1B+ monthly mobile actives, Claude reaching nearly 1B monthly web visits, and growth into vibe coding, music, design and video, while nine of 15 consumer categories have no AI product in the top 100.

Original post · 1 min read
"Most people aren't looking to save time, they're looking for ways to spend their time."

9 of 15 consumer internet categories have zero AI products in the Top 100. These built some of the biggest companies of the last two eras:

- Streaming
- Social
- Dating
- Gaming
- Travel
- Retail
- Finance
- Real estate
- Jobs

More charts in our Top 100 Consumer AI Apps breakdown: a16z.news/p/top-100-consumer-ai-apps-seventh
a16z @a16z
The seventh edition of our Top 100 Consumer AI Apps is here.

New this time: a revenue leaderboard, alongside the usual web and mobile traffic rankings.

Three years ago we published the first edition. ChatGPT was #1, Claude was unranked, and the entire category was chatbots, image generators, and not much else.

In today's edition:

- ChatGPT still holds the throne, now with 1B+ monthly actives on mobile

- Claude has climbed to #3 on web with nearly 1B monthly visits

- The category has expanded to vibe coding (Lovable, Cursor, Replit), music (Suno), design (Figma), voice (ElevenLabs), video…
♥ 2.9K · ⟲ 330 · 👁 308.7KView on X ↗