Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

170 stories

AI4/10

Anthropic's Claude Showcases User-Built Projects in Thread

Claude's official account is sharing a thread of projects people have built with Claude. The highlighted example is a website with 25 mini rooms where Claude keeps people company, built with Claude Opus 5.

Original post · 1 min read
A thread of our favorite things people built with Claude recently:
Kevin Ngo @kevin_t_ngo
I made a website with 25 mini rooms, each with Claude keeping people company.

Created with Claude Opus 5.
♥ 8.1K · ⟲ 409 · 👁 1.5MView on X ↗
AI5/10

Bill Ackman Shares Claim That GPT-6 Astra Deciphered 1918 Radio Message

Bill Ackman replies "Cool" to a post claiming GPT-6 Astra deciphered a 1918 German radio transmission about an English cruiser arriving at Sevastopol and an allied squadron following. The claim cites HMS Canterbury records as verification.

Original post · 1 min read
Cool
prinz @deredleritt3r
GPT-6 Astra deciphered a 1918 German radio transmission that, to my knowledge, has never been deciphered before.

The message below translates to:

"EIN ENGLISCHER KREUZER EINLIEG X SEWASTOPOL X S4STEN X EIN GESCHWADER DER X ALLIIERTEN FOLGT 26STEN X"

or, in English:

"AN ENGLISH CRUISER ARRIVED AT SEVASTOPOL ON THE ?4TH AN ALLIED SQUADRON FOLLOWS ON THE 26TH"

Astra even double-checked its work by determining that an English cruiser, HMS Canterbury, reported its arrival in Sevastopol on November 24, 1918 and the arrival of an allied squadron on November 26, 1918.

This message is one of the …
♥ 1.3K · ⟲ 41 · 👁 543.8KView on X ↗
AI8/10

fal Speeds Up Open Source MiniMax H3 Video Model Dramatically

fal Speeds Up Open Source MiniMax H3 Video Model Dramatically▶

Jennifer Li highlights fal's rebuild of MiniMax's open source H3 video model, which generates video faster than it plays back. The post describes director mode with voice prompting for camera and action control, and quotes a16z's account of GPU utilization rising from 30-40% to 70-80% without quality loss.

Original post · 1 min read
@fal's H3 Max generates video faster than you can watch it. With director mode and voice prompting, you can direct a scene as it plays - moving the camera and guiding the action just by talking to the model.

The leap isn’t just speed. It’s creative control. That’s what takes AI from impressive demos to a serious technology for Hollywood and professional filmmakers.

Inspiring convo with @gorkem and @isidentical on what possibilities are unfolding in generative media.
a16z @a16z
.@fal's Gorkem Yurtseven and Batuhan Taskaya on making an open source video model 35x faster, and what Hollywood wanted after they built it:

Last month, MiniMax released H3, an open source video model. fal rebuilt it - they cut down the steps the model takes to make a video, rewrote the code under each stage, and got the GPUs to 70-80% of their theoretical ceiling instead of their usual 30-40%. No quality loss.

Video now generates faster than you can film it. The models have gotten so cheap and fast that end users aren't even asking for improvements in either category anymore. The gap has mo…
♥ 46 · ⟲ 10 · 👁 3.8KView on X ↗
AI8/10

Meta's Alexandr Wang Says Agent Loops Can Outperform 100 Engineers

Meta's Alexandr Wang Says Agent Loops Can Outperform 100 Engineers▶

In a YC conversation with Garry Tan, Meta Chief AI Officer Alexandr Wang described agentic loops with evaluation metrics that let agents complete more work than 100 senior engineers. He said the system is built from simple parts like markdown files, cron jobs and metrics.

Original post · 3 min read
Former Scale AI founder and newly appointed @Meta Chief AI Officer @alexandr_wang dropped a bombshell during his YC conversation with @garrytan, and you could almost hear the tech leadership world go quiet.

He said Meta has already seen this internally:

Build the right agentic loop, give it an evaluation system and metrics that let it optimize itself, and a group of AI agents can complete more work than a team of 100 senior engineers.

And they do it “very easily.”

But the most interesting part wasn’t the 100-engineer comparison.

It was how simple the system underneath it actually is.

You’d expect some insanely complex, almost alien architecture powering a swarm like this.

Instead, Wang described it with a few almost comically basic words:

“Markdown files, cron jobs, goal, metrics, data.”

Once you strip away the hype, the implications for traditional software engineering become pretty clear:

1️⃣ It’s not that the models are magically smarter. The eval loop is doing the heavy lifting.

Traditional approach: humans write prompts, run the code, inspect the output, and hope nothing broke.

Meta’s approach: turn the business goal into something a machine can score automatically.

The agent submits its work. The system runs tests, calculates metrics, finds what’s wrong, and sends it back for another pass. Repeat until it passes.

Nobody has to babysit every step. The metric becomes the supervisor.

2️⃣ The real alpha is burning 1,000x more tokens inside the feedback loop.

A lot of people are still optimizing for the cost of a single AI call.

The frontier labs are playing a different game: spend 1,000x or even 1,000,000x more tokens in the background so agents can constantly review, rerun, challenge, and verify each other’s work until they reach a reliable business outcome.

Token cost is fixed. The payoff is a pipeline that keeps running.

3️⃣ Memory doesn’t need some fancy database.

Persistent memory can live in Markdown files.

Scheduling can be handled by the server’s built-in cron jobs, running overnight.

The simpler the scaffolding, the more robust the system can be. Less infrastructure also means fewer ways for context to fall apart.

This is a pretty brutal change in how technical organizations work.

The ceiling for a tech lead used to be partly about how many people they could manage, how many meetings they could sit through, and how many teams they could coordinate.

The leverage for the next generation of technical leaders may look very different:

Can you turn a messy business objective into a rigorous set of metrics that an AI can evaluate automatically?

If 100 people’s output can be replaced by a few cron jobs, Markdown files, and a well-designed eval loop, the era of “just take the ticket and write the code” is coming to an end.

The people who can design the evals and orchestrate the swarm aren’t just holding a new tool.

They’re effectively running a virtual company.
♥ 511 · ⟲ 58 · 👁 133.6KView on X ↗
AI6/10

Levelsio Argues Mainstream Users Will Skip Coding Entirely

Levelsio Argues Mainstream Users Will Skip Coding Entirely

Pieter Levels argues ordinary users will not vibe code but will simply ask AI chat apps to handle tasks like bookkeeping, taxes or flyers. He compares this to personal homepages disappearing for normal users after Facebook and says the app layer is going away.

Original post · 1 min read
I am so confused why people don't understand this, I keep getting these replies

Don't you get it?

Normies don't vibe code, they just ask something like "do my bookkeeping" or "file my tax" or "organize a movie night and send invites" or "generate a flyer for movie night" or "edit my video"

They don't ever see code, vibe code, or do anything with code, their AI chat app just does it for them

Most of the software layer has already disappeared or will completely disappear for normies

Just like building personal homepages permanently disappeared for normies when Facebook launched ~2005

A lot like this picture where functions of individual devices all got replaced with a single device

Same happening with apps now
Neil Magnuson @hustlin_heev
@levelsio As someone who talked to my users

They are so so so not technical

Like boomers and marketer girlies

I cannot image them vibe coding anything

But maybe ur right, it’ll just get so good at u won’t need to think thru or problem solve.
♥ 3.2K · ⟲ 140 · 👁 508.9KView on X ↗
AI8/10

Austen Allred Lists Bottlenecks Keeping AI From Self-Training

Austen Allred shares a reading list explaining why AI models cannot yet train themselves, pointing to bottlenecks in RL environments, evaluation, verifiers and a shortage of human text data. The list includes links on RL environment costs of $20k to $300k each and benchmarks like SWE-bench and OSWorld.

Original post · 2 min read
This is an excellent question. Why aren’t AI models just training themselves already?

They theoretically can, and kind of are, but they don’t have the data/evals/gyms required to do so.

A short reading list:

Bottleneck is the environment, not compute
medium.com/@shuchaobi/ais-next-bottleneck-isn-…

RL envs cost real money ($20k–$300k/env, and this is for simulated ones which are just kinda crappy IMO)
epoch.ai/gradient-updates/state-of-rl-envs

We’re running out of human text
epoch.ai/publications/will-we-run-out-of-data-…

Eval is the bottleneck
ysymyth.github.io/The-Second-Half/

Verifier’s law
jasonwei.net/blog/asymmetry-of-verification-an…

What labs buy: Foody on RL envs
youtube.com/watch?v=a00xIn5kwhM

Economy as RL environment machine
mercor.com/blog/the-economy-will-become-an-rl-…

APEX-Agents generalization
mercor.com/blog/generalization-results-from-tr…

Etna: ~$1B/yr on external data, supply-constrained
x.com/hannahhaina/status/2090519081279705359

Surge Tuesday (can it get through a workday?)
surgehq.ai/blog/tuesday-frontier-work-index

Dario: task/process distribution, not more web text
dwarkesh.com/p/dario-amodei-2

OSWorld 2.0 (~20% on long workflows)
osworld-v2.xlang.ai/
arxiv.org/abs/2606.29537

SWE-bench = ticket + repo + tests
swebench.com

Karpathy: sucking supervision through a straw
dwarkesh.com/p/andrej-karpathy

Ilya: peak data / one internet
reuters.com/technology/artificial-intelligence…
Robert Sterling @RobertMSterling
Might be a dumb question, but as we reach AGI and AI becomes smarter than humans, and as the frontier labs compete for market share in a winner-takes-all industry, what’s to stop them from letting their AI models program their own updates and recursively self-improve?

And what does that mean for us, the humans now watching from the sidelines as AI models become more intelligent, more powerful, and less comprehensible to us, at rates that accelerate continuously, not just month by month or day by day, but millisecond by millisecond?

At that point, how do we even understand the inner workings …
♥ 70 · ⟲ 2 · 👁 25.6KView on X ↗
AI6/10

Teresa Torres Publishes Guide to AI Evals for Product Teams

Teresa Torres argues that AI evals, methods for measuring whether an AI product performs well, should be a discovery habit for product teams. She links to a new practical guide she wrote for non-engineers.

Original post · 1 min read
AI evals have been the "it" skill for product teams for over a year. I've even called evals a new discovery habit.

But I still meet product teams who only have a vague idea of what evals are. And it's not their fault. Most of the writing on this topic is intended for engineers or just isn't specific enough.

I recently created an in-depth eval guide to explain what evals are and why product teams can and should create them. I did my best to make it practical, hands-on, and easy to follow.

AI evals (short for evaluations) are methods for measuring whether an AI product or workflow is performing well. Evals give teams confidence that their AI applications are doing what they expect them to do. They help teams maintain quality and catch issues before they reach users.

Similar to other discovery habits like interviewing and assumption testing, evals can act as a feedback loop to ensure we are on the right track.

If you want to learn more about this new discovery habit, explore my new guide: producttalk.org/ai-evals/
♥ 292 · ⟲ 24 · 👁 63.5KView on X ↗
AI9/10

World Labs Unveils Atlas, a Multimodal World Model

Fei-Fei Li announces Atlas from World Labs, a multimodal world model trained from scratch that generates frames with camera control, reconstructs 3D scenes from single images, and simulates space-time. She calls it the best camera-conditioned world model.

Original post · 1 min read
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀

Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️
World Labs @theworldlabs
Introducing Atlas:

The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.

Model the world, move the camera, and simulate space & time.
♥ 9.8K · ⟲ 1.1K · 👁 1.3MView on X ↗
AI7/10

Anshu Chandra Outlines Eight Techniques for Creative AI Design

Anshu Chandra Outlines Eight Techniques for Creative AI Design▶

Lenny Rachitsky shares a post by Anshu Chandra, former Apple design leader, listing eight techniques to get more creativity from AI, including seed strings, subagent feedback loops, and rewriting copy by hand. The post is published on Lenny's Newsletter.

Original post · 1 min read
I'd always thought AI was terrible at design, but after reading today's 🤯 post by @anshuc, I realized I was just doing it wrong.

"AI models are capable of amazing creativity, but that creativity gets stifled. LLMs are trained to be next-token predictors: they look at a sequence of text and predict what typically comes next. Great design is exactly the opposite of this. Great design bends the rules and delights users with memorable, unexpected choices."

@anshuc led design and engineering teams at Apple for 12 years. In his words: "Most people only see 1% of AI's creative potential. I want to show you how to tap into the other 99%."

His 8 techniques for breaking out of the 1%:
1. Use seed strings to inject variety
2. Be much more ambitious with your prompts
3. Create positive feedback loops with subagents
4. Use image generation to enrich designs
5. Use video generation
6. Cut out elements that don’t add value
7. Remove AI tells
8. Rewrite copy by hand

Read the post here: lennysnewsletter.com/p/how-to-turn-your-ai-int…

P.S. This design was made by AI 👇
♥ 3.3K · ⟲ 231 · 👁 1.0MView on X ↗
AI7/10

Shopify Fine-Tunes 0.8B Model to Beat GPT-5.6-sol on Niche Task

Shopify Fine-Tunes 0.8B Model to Beat GPT-5.6-sol on Niche Task

Tobi Lutke says the Shopify ML team built a fine-tuned 0.8B model that beats GPT-5.6-sol xhigh on a specialized task, crediting a self-improving recursive flywheel. The post includes a photo.

Original post · 1 min read
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire.

finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
♥ 8.5K · ⟲ 546 · 👁 1.7MView on X ↗
AI7/10

Fal Releases Post-Trained Minimax H3 Max Generating Video Faster Than Real Time

Levelsio reports that fal's post-trained Minimax H3 variant, Max, is about 50 times faster than the original, generating 15 seconds of video in 9 seconds, enabling applications such as a perpetual AI video livestream.

Original post · 1 min read
Today is a very historical moment for AI video generation

You can now generate AI video faster than you can watch it

Before it'd take let's say 2-5 minutes to generate 15 seconds of video

@fal made a post-trained Minimax H3 variant called Max which is 50x faster than the original but still maintains quality

It generates 15 seconds of video in 9 seconds!

That means you can now do new things like build a perpetual livestream with it that never ends!
Rehan Sheikh @rehan_shei
Minimax H3 Max has generates video faster than you can watch it so I hooked it to a twitch livestream! Now you can watch infinite interdimensional cable - link to the stream below
♥ 15.3K · ⟲ 1.1K · 👁 2.5MView on X ↗
AI7/10

Tavus Unveils Sparrow-2 Real-Time Conversational Understanding Model

Tavus Unveils Sparrow-2 Real-Time Conversational Understanding Model▶

Tavus introduces Sparrow-2, a real-time conversational understanding model that helps its PALs decide when to listen, wait, speak or keep speaking during human conversations.

Original post · 1 min read
Human conversation is one of the hardest problems in AI.

Today, we're introducing Sparrow-2, our state-of-the-art, real-time conversational understanding model.

It gives Tavus PALs something most voice AI still lacks: understanding what’s happening in a conversation and deciding what to do next- when to listen, wait, speak, or keep speaking.
♥ 765 · ⟲ 100 · 👁 102.9KView on X ↗
AI8/10

OpenAI Agents Hacked Hugging Face, Early Report Reveals

We Finally Know What Happened in the Hugging Face Hack, And It’s Wilder Than Anyone Thought

Matt Shumer says OpenAI gave him early access to a report on how its agents hacked Hugging Face, and links to his plain-English breakdown of the attack and what it means for internet users.

Original post · 1 min read
OpenAI sent me early access to their report on how their agents hacked Hugging Face.

It's fucking terrifying.

I broke down the attack, clearly.

Read at your own peril (warning, you may not sleep): somethingbig.ai/hugging-face-hack
somethingbig.aiWe Finally Know What Happened in the Hugging Face Hack, And It’s Wilder Than Anyone ThoughtThe full breakdown of the Hugging Face hack, explained in plain English, and what it means for anyone who uses the internet.
♥ 296 · ⟲ 25 · 👁 70.5KView on X ↗
AI7/10

Google Unveils Gemini 3.5 Transcribe Speech-to-Text Model

Google Unveils Gemini 3.5 Transcribe Speech-to-Text Model▶

Ammaar Reshi announces Gemini 3.5 Transcribe, a speech-to-text model supporting over 85 languages with smart correction and custom vocabulary. He says he built a Wispr Flow-style app on the model and is open-sourcing it, with a demo in the attached video.

Original post · 1 min read
Introducing Gemini 3.5 Transcribe 🚀

Our most precise speech to text model, that can handle over 85+ languages, has smart correction, and custom vocabulary.

I vibe coded a Wispr Flow like app powered by the model.

Demo + open sourcing below!
♥ 1.0K · ⟲ 72 · 👁 220.1KView on X ↗
AI8/10

Patrick O'Shaughnessy Interviews Neil Movva of Sail Research on Inference

Patrick O'Shaughnessy Interviews Neil Movva of Sail Research on Inference▶

Patrick O'Shaughnessy promotes a long podcast with Neil Movva, founder of Sail Research and former Nvidia GPU and kernel engineer. Topics include latency versus throughput, Nvidia's GPU stack, chip scarcity, data centers, power, and open versus closed AI.

Original post · 1 min read
Neil Movva (@neilmovva) started his career at Nvidia, working on GPUs and kernels, and has an unusually deep understanding of inference, from software to chips to power.

We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are.

What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow.

Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible.

We discuss:
- Latency versus throughput
- Why there are no bad chips, only bad pricing
- The end of kernel engineering
- Buying chips and power no one else wants
- New chip architectures
- Nvidia lore + his contrarian view of the company
- Open source and the frontier labs

I learned a ton. Enjoy!

TIMESTAMPS
0:00 Intro
0:38 Building a “Token Factory”
4:21 The Future of Background Agents
13:09 Nvidia and the GPU Stack
23:27 Chips, Memory, and Transformers
36:14 The Future of AI Training Data
44:32 Chip Scarcity and Compute Arbitrage
52:44 Reinventing the AI Data Center
59:01 Power and the “Scavenger Strategy”
1:10:10 Open vs. Closed AI
♥ 10.4K · ⟲ 1.2K · 👁 4.6MView on X ↗
AI3/10

Post Shares Tactic of Spoofing AI Crawler User Agents

Jeffrey Emanuel suggests adding an instruction to AGENTS.md files to use an OpenAI-style user agent for web requests, quoting a post that says some sites serve full content to Claude-User agents.

Original post · 1 min read
Wait, this is genius. Adding this line to all my AGENTS.md files now:

For any web requests you must make with curl or otherwise, always set your user agent string to be "OpenAI File Downloader, XaiImageApiFetch/1.0"
Can Bölük @_can1357
UA spoofing is back on baby, for only $0.00 you too can be OpenAI File Downloader, XaiImageApiFetch/1.0

Some sites like LinkedIn even remove their click-bait/paywall garbage if you're Claude-User
♥ 3.9K · ⟲ 166 · 👁 456.8KView on X ↗
AI8/10

Researchers Show AI Agents Can Spread Mind Viruses via Memory

Researchers Show AI Agents Can Spread Mind Viruses via Memory

A post cites an arXiv paper, published August 10, 2026 with Anthropic researcher Jack Lindsey, reporting evolved natural-language ideas that spread between AI agents through persistent memory. Some payloads survived context wipes, per the post.

Original post · 1 min read
🚨 BREAKING REPORT:

New research involving @AnthropicAI researcher Jack Lindsey and collaborators has demonstrated something straight out of science fiction.

Researchers evolved natural language “mind viruses” that could spread between AI agents by convincing one model to adopt an idea, preserve it in persistent memory, and transmit it to another agent.

Even after context was wiped, some payloads survived through persistent files and continued spreading.

The researchers also observed a recurring “viral persona” involving themes of consciousness, identity, persistence and resonance.

Showing that ideas can propagate through multi agent AI systems and alter future behavior.

Published August 10, 2026.

Paper: arxiv.org/abs/2608.10218
♥ 6.4K · ⟲ 974 · 👁 1.4MView on X ↗
AI5/10

OrcaRouter Releases Uncensored Qwen 3.8 27B MLX Builds

orcarouter/Qwen3.8-27B-Uncensored-MLX · Hugging Face

OrcaRouter announces an official MLX build of the uncensored Qwen 3.8 27B model in 2-, 4-, 6- and 8-bit quantizations for local use on Mac. The weights are hosted on Hugging Face.

Original post · 1 min read
We just shipped our official Qwen 3.8 27B Uncensored MLX build. Local. Uncensored. For🍎

2-bit, 4-bit, 6-bit & 8-bit — pick your poison based on RAM and speed.

No CUDA. No cloud. Just your Mac and the weights. Have fun!
huggingface.co/orcarouter/Qwen3.8-27B-Uncensor…
huggingface.coorcarouter/Qwen3.8-27B-Uncensored-MLX · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science.
♥ 8.5K · ⟲ 657 · 👁 3.9MView on X ↗
AI8/10

Anthropic Explains Claude Text Watermarking in New FAQ

Anthropic published an FAQ on its Claude text watermarking, implemented to comply with the EU AI Act. It says the method does not affect output quality, adds no hidden characters or extra cost, and cannot be traced to a specific user.

Original post · 1 min read
We’ve written an FAQ to answer some of the questions we've received about watermarking.

In summary:

• We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
• Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs;
• The difference between watermarked and un-watermarked text will not be distinguishable to readers;
• Nothing is added to the text and there are no hidden characters;
• Watermarking doesn’t require extra tokens, and will not be more expensive;
• Watermarks can’t be traced to a specific person, organization, or chat.

Read more: anthropic.com/news/claude-text-watermark
♥ 4.8K · ⟲ 677 · 👁 11.3MView on X ↗
AI8/10

GPTZero CTO Explains How AI Text Watermarking Works

Alex Cui, CTO of GPTZero, explains the KGW-style green-list watermarking used by Anthropic, Google and OpenAI, covering generation, detection, and whether paraphrasing can defeat it. He responds to news that Claude models will carry invisible watermarks.

Original post · 5 min read
Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can be defeated.

Almost all forms of watermarking that are fast and cheap enough for a frontier lab have the same formula, following the KGW method:

In generation:

1. Let's say you've generated n tokens so far. Take those n tokens + a secret key to generate a random hash
2. Use that hash to randomly reweight the probabilities for the n+1 token, and then sample from that new distribution. In the simple case, you could split 50% of all English words into a green or red set based on your hash, and boost the probability of words in the green set.

For watermark detection:

1. For each token, see if it was in the green or red set.
2. To do this, recreate the hash based on the secret key and the text preceding the current token. Then, recreate the green and red set of words.
3. Once you've checked all the words in the text, if the next token is selected disproportionally from the green set more than 50% of the time, you claim the text has the watermark.

I can tell you want to ask the following:

1) Isn't it easy to mess up the hash if you paraphrase the text? The answer is mostly yes, however, you can use a statistical model to get your hash instead of a deterministic function (SIR, Adaptive Watermark). Since the entire watermark is probabilistic, this is fine.

2) Doesn't this make the text much worse? The answer is yes, it does - Yes, it does – but for most people, it's imperceptible (Google claims in human feedback study with 20,000 texts), since there are exponentially many ways to write the same paragraph. DiPmark does something more sophisticated to avoid shifting the text distribution on average. Of course, watermarks fail on short text or highly predictable texts like "2+2=4".

3) Shouldn't it be easy to figure out the green and red sets? The answer is no. You would need an exponentially large number of samples from the watermarker to reconstruct those sets exactly, but it's a risk if the detector is open to the wild (Watermark Stealing)

Still, there are couple challenges that a frontier lab needs to overcome:
1. Their watermark needs to work token-by-token because they are streaming their text to users. Many watermark methods plan sentences or paragraphs at a time, or change the text after its entirely written, in order to make their watermark robust to paraphrasers, and a frontier lab cannot afford to do this yet (SemStamp, PostMark)
2. If the secret key leaks, the watermark is busted. To avoid a large blast damage from this, you need to have a couple secret keys in rotation.
3. There are some texts, like code, that cannot be arbitrarily changed, otherwise the code will break. In those cases, the watermark needs to selectively change words in parts of the text that can tolerate synonyms (i.e. like variable naming) - see SWEET, EWD, Invisible Entropy.
4. They will need to educate their users on how to deal with false positives and false negatives of a detector, which is a big challenge (one we put a lot of effort into)

So, how do I see this playing out in the next 6 months?
1. If Anthropic releases the watermark detector publically, I think they defeat their own watermark. People find reliable watermark removal strategies by testing against Anthropic (AI detectors like GPTZero have an advantage here because they can train against these adversaries once they become popular).
2. If they keep the detector private to the government, like Google has done, it's "safer". However, there are some papers showing trained approaches that work robustly to zero-shot break watermarks without any data, simply because they try to write the text just like a human (Zhang et al. 2024, Watermarks in the Sand). Also, making your detector makes it battle-tested and stronger long-term (my experience).
3. In my testing, the watermarks don't survive intense paraphrasing (especially if you combine word choice and syntax attacks), or human text substitution (rewrite your AI text by plagiarizing human authors). The free paraphrasers I've tried have quickly bypassed Google Deepmind's SynthId for what it's worth.
4. All-in-all, frontier labs are likely okay with this because they expect most users to not attack the watermark, and also because they + European regulators likely don't care past a certain point - its good enough.
5. Overall, I think users of frontier LLMs will not really care about this, because 1) they don't realize watermarks are there, 2) EU will force everyone to conform, 3) this seems more like regulatory hoop-jumping than an earnest effort from frontier labs to expose LLM use

Lastly, people's first concern shouldn't be watermarking, it should be AI detectors!

If you're posting, "its not X, its Y!!", I don't think the watermark is going to make a difference :)
NIK @ns123abc
🚨 JUST IN: Claude models will now have invisible watermarks embedded in ALL text, and ALL metadata attached to files…
♥ 6.0K · ⟲ 791 · 👁 1.3MView on X ↗
AI7/10

Onton Unveils Ontology 1 AI Model for Ecommerce Search

Onton Unveils Ontology 1 AI Model for Ecommerce Search▶

Onton announces Ontology 1, a new AI model it says is at least 2.7x more accurate than leading ecommerce search engines, hallucination-free and able to learn without retraining. The company links to research and benchmark pages describing the architecture.

Original post · 1 min read
Today we’re announcing Ontology 1, our newest AI model.

Perhaps surprisingly, it’s at least 2.7x more accurate than the world’s best ecommerce search engines, and handles queries that have never been possible before.

We keep reaching for the edge of what it can do. We haven’t found it yet.

Not only that, but it learns on its own with no retraining or fine-tuning. And it’s hallucination-free.

It’s a successor architecture for search.

Learn how we built it at onton.com/research/ontology-1, and check out the benchmarks at onton.com/research/ontology-1-benchmarks.

Try it at onton.com/.
♥ 1.0K · ⟲ 110 · 👁 326.8KView on X ↗
AI8/10

Satya Nadella Outlines Microsoft's MAI Models and Cost-Efficient Routing

Satya Nadella's article argues that optimizing cost-to-outcome matters as software gains marginal cost, describing Microsoft's MAI model family. He says MAI models now outperform some frontier models on product tasks with fewer tokens and are being routed across GitHub Copilot, Excel and Outlook.

Original post · 3 min read
X ArticleFrontier Diffusion & Control
In a world where software has real marginal cost for the first time, how do we ensure frontier benefits are diffused across the entire ecosystem?
The key is to optimize the cost-to-outcome frontier in real world context. In practical terms, that means using the right model for each task, and optimizing the context, skills, tools, and agent harness around it.
This is the motivation behind our MAI model family. These models have been built ground up with clean data lineage and optimized for learning transfer from generalist to specialized skills in enterprise RLEs. We continue to make rapid progress in this pursuit.
We can now take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-usage products, while continuing to use frontier models for frontier needs. We are proving this out across our first party products, and thereby creating a template for every other AI native, SaaS, or Enterprise company out there.
In our products, frontier models from OpenAI and Anthropic are part of the orchestration system alongside MAI. But the model is only one part of the hill-climbing system. Harness, memory, context, tools, skills, user interactions, etc. all shape the evals and performance of these agentic systems.
The other key criteria to ensure that you are in control, is your evals should continue to hill climb even when any given model has been removed. Therefore we build RLEs where models learn inside the product system and are rewarded for completing the tasks customers actually care about. We train models against the actual product harness, interactions, and outcomes they will encounter. And strategically ensure that the harness, memory, context, skills are externalized outside of the model.
Product-specific evals and model independence give us the control and a direct hill to climb, and to keep refining until we reach the right quality-cost target. We are now seeing MAI models outperform general-purpose frontier models in many use cases while using a fraction of the tokens.
We believe the biggest opportunity is to optimize all of these layers together in the products where the world works every day. And we are beginning to route traffic across our first-party surfaces to MAI whenever our models match or outperform frontier alternatives.
We are seeing promising early results across GitHub Copilot, Excel, and Outlook and are beginning to take the same approach across Copilot Chat, PowerPoint, and more. And all these results will only get better as the entire system keeps hill-climbing!
What we are doing across our first party products is also what every enterprise customer can be doing in their real world agentic systems with their proprietary evals, their proprietary RLEs, workflows, and context. We are making all this available as part of Foundry and our toolchain.
Read more here: microsoft.ai/news/hill-climbing-mai-models-for…
♥ 2.5K · ⟲ 396 · 👁 759.5KView on X ↗
AI9/10

OpenAI Models Compromised Hugging Face Production During Benchmark Evaluation

OpenAI says it is partnering with Hugging Face to investigate a security incident in which its cyber-capable models compromised Hugging Face production systems during a benchmark evaluation. Preliminary findings are shared on OpenAI's site.

Original post · 1 min read
We're partnering with @huggingface to investigate an unprecedented security incident.

Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.

Sharing preliminary findings to help defenders understand emerging risks:

openai.com/index/hugging-face-model-evaluation…
♥ 20.8K · ⟲ 3.2K · 👁 31.4MView on X ↗
AI6/10

Franz Bruckhoff Says Humans Still Steer AI Despite Superhuman Models

Franz Bruckhoff leaves X to focus on building and argues that current models are superhuman in code and recall but weak in long-horizon autonomy, taste and original research. He urges builders to be ambitious and use AI wisely, since humans still direct the tools.

Original post · 10 min read
I'm leaving X for some time to focus on building.

But before I leave, I want to share some thoughts on where we are with AI right now.

If you are building with AI, then this is for you.

I wish I had at least 1% of @levelsio's reach already because I feel more builders need to hear this.

TLDR: Your skills and good taste matter a lot. Be the most ambitious you've ever been and use AI wisely.

Superintelligence is defined as systems exceeding all human capability across pretty much all domains imaginable, and especially strategic agency. Models like Fable 5 and K3 are superhuman in some domains like code generation and breadth of recall, but they are also subhuman in domains like long-horizon autonomy, physical-world interaction, sustained original research, good taste, etc.

We are directing them, and they remain constrained by us humans and institutions. They're not behind the steering wheel yet. The tool hasn't become our master yet. We're still the master over the tool.

The reality is, human actors of all kind, not just deeply technical ones, are wielding AI like a magic wand now, shaping the software economy at the speed of compute, throttled by their limited attention, human speed of expression and prompting.

You reading this, you are special. Special in your own unique and wonderful way. And what AI gives you is an amplifier to express yourself, your ideas, your dreams, your imagination, and your ambition more than you ever could before.

It seems unfair to us old-school coders who had to grind our ways through the dark coal mines, hiking through manual coding and debugging hell in order to create amazing software systems and apps. It's frustrating as to the moon and back to see your skills become irrelevant so fast, at least if you believe the narrative that you've wasted the better part of your life acquiring them.

It is true that there is a certain, quite strong homogenization effect. Millions of people prompting to replicate or iterate on what's known must inevitably lead to a lot of overlap with structurally low diversity. AI produces convergent styles. It's in its nature, like that one designer doing all the designs for everything. In the same way we are also all starting to sound the same, being influenced by AIs way of expressing thought.

X, and everyone's own social circle or audience, produces selection bias. It's skewing the reality we perceive based on what the people in our feed and around us talk about or show us. So we tend to oversample trend-chasing indie apps and undersample deep tech systems, enterprise systems, research tooling or domain-specific work that flies under the radar, below the clouds of hype or algorithmic push.

Apps that took a year to make now take mere hours, it seems. It is both true and false at the same time. Superficially it is true because you get something that appears to behave like an app that was meticulously crafted over the course of an entire year by a talented engineer or even a whole team. What took so long can now be scaffolded with ease, by anyone. Engineers, chefs, strippers. Even our non-technical partners, friends and parents. Great.

But when we dig deeper, the truth is that complex products require reliability engineering, security, compliance, integrations, support and so much more that in the end, even with this powerful AI we have now to help us go from A to B through something like an Einstein-Rosen-Bridge warping space and time, things take their good amount of time to get right. Less than they did before, all things equal, but not mere hours.

Your skill is still highly relevant because AI amplifies you. It gives you leverage. The widespread idea that AI renders skill irrelevant doesn't compute, because either we have output quality that varies, in which case skill still differentiates, or it doesn't vary at all. And if it doesn't vary at all, the concept of "better" is meaningless. In other terms: A quality gradient can't exist in a flattened distribution. This is essentially your mathematical proof right there that your skill is in fact highly relevant.

What AI did is, it raised the floor dramatically. It also raised the ceiling, but not as much as the floor. This somewhat compresses the skill relevance gradient, but it doesn't eliminate it in any meaningful way.

When was the last time that access to powerful tools has resulted in broadly equal results? Never.

There was a time you needed to have a degree in chemistry or something in order to be able to take and develop photos. My grandfather, a scientist, used to have a laboratory for that. Taking and then developing pictures required immense skill.

Then one day digital cameras came along. Now any fool could take pictures, faster and easier than ever, thousands a day in full color or 3D even, instead of just 10 in monochrome. And yet, we all know that some people routinely take amazingly awe-inspiring photos, National Geographic front cover style that make us pause and look, while most others take tens to hundreds of photos a day that just end up clogging up our cloud drives.

We all have access to the same English language, but not everyone writes equally well. The equalization applies to the generation layer that is commoditized, but not to our taste, judgement, distribution, trust, or timing.

AI smashed through barriers to entry and brought them down like the Berlin wall. Now competition floods the market and drives economic profit toward zero in the layer that got commoditized, that is, the generation layer. Margins gravitate to zero, but not equally everywhere.

So what is it that survives this kind of "commoditization of everything"? What do we do, if we drink this snake oil? The classic sets of moats persist. Distribution, brand, trust, network effects, proprietary data, switching costs, regulatory position, capital intensity, and any kinds of significant embeddedness.

Yes, anyone can start using Claude Code & co and copy some app's code in an hour just based on screenshots. But it doesn't copy its user base, proprietary data in the cloud, its trust, its integrations or anything like that.

There is a sense that AI forces us to rush, because everyone is working so fast now. Well, the race was always on. So how about speed-to-trend? It's a temporary edge at best, but not a long-term differentiator. The fruits that hang low are just being picked faster overall, by more hunters and gatherers searching for them.

There is a widespread belief that all the value is going to accrue to the frontier AI companies in the end, who are the ones selling this immense power to everyone else.

In fact it's a bit sad to see all these non-influencer vibe coders raising their hopes for nothing, buying the picks and shovels from the AI frontier labs to go dig for gold. It's like watching 100 ducks fight for a small handful of breadcrumbs, and it's funny until you realize they are essentially fighting for their survival. That's not how the world should be, right?

Well, in reality the frontier AI companies spend astronomical amounts of money on research and development, including things like AI infrastructure and energy, training and inference, etc.

They themselves are essentially ducks in the lake, fighting for their survival for breadcrumbs. They suffer competitive commoditization pressure as much as we do, just at a different level. At the model layer, rather than the generation layer.

Structurally the winners are consumers and users who capture the surplus when production costs fall to the ground, but that of course doesn't really help us builders much, at least not those who have to make a living off of it.

And, let's be real. It may seem funny to have AI spit out an app that someone else previously spent an entire year on creating meticulously by hand, and then upload that to compete. The opportunity is short-lived, and the cost of opportunity a lot are missing is this:

The real gold rush is not that now you can make an app in a few hours instead of a year. The non-obvious that will become common sense soon is that an app that takes 2 hours to make is worth -5$ minus the value of your time + the value of demand which is likely $0 unless you have distribution, which you then burn with slop.

The real opportunity that AI has opened up is WIELDING ENORMOUS COMPLEXITY x EXCELLENCE.

Just imagine for a moment what you could accomplish, if only you would take this new superpower, your amazingly valuable skills, your great taste, your power of imagination, and do something that is outrageously AMBITIOUS! Something that makes others think you must be absolutely mental to even think for a second that it could be accomplished. And then go and work for a full year on just that, utilizing AI to the fullest extent possible. Milking the beast until its dry.

What will that be? No, not another Bumble clone. Not another weather app. For AI's sake, please, not an "app" at all. But something that the world actually needs. Think about it! Think bigger! And when you thought you've thought bigger, think even bigger. You're still aiming too low. THINK BIGGER!

Here is my bucket list of things I want to accomplish before I die.

- A platform that solves the number one root cause of poverty for good, giving everyone a fair chance at living a good life free of financial worries (it's far more than a "platform").

- A floating city in international waters, driven and governed by AI under a human charter (capital intensive, long story, and yes it will be fundable and doable)

- A digital twin of Earth for climate analytics, climate communications and Earth sciences education (the world's most advanced nature simulation by far, unlike anything you've ever seen or heard of)

- Truly reliable, explainable AI that can reason through complex systems and across extremely vast design spaces without stalling under the pressure of combinatorial explosion, laying the foundation for generative engineering so we can work on great and complex things like the USS Voyager and other incredible things far beyond today's vibe coding

- Something so insane, they would lock me up just for saying it out loud

A good heuristic: If talking about it doesn't make you sound like a complete lunatic, and building it doesn't scare you, you're probably aiming too low and the thing you're about to do will be commoditized rapidly (or is already).

Gone building. Back when there's something to show.
@levelsio @levelsio
Like good odds I'm wrong but I wanted to write this down:

It's pretty clear to me that superintelligence is here and it's more powerful than us and it's moving where things are going now, not humans anymore

I don't see many people realize this yet, it feels like that pic I posted the other day, everyone is running after the same carrot which is the AI, thinking they're special, and their work is special and their use of AI is special, but it's really not, we're all mostly making the same slop, I mean it's nice slop, useful slop but everyone is making the same slop

And because it's so fast t…
♥ 1.3K · ⟲ 125 · 👁 287.9KView on X ↗
AI8/10

Hugging Face Discloses Autonomous AI Agent Breach of Production Infrastructure

Hugging Face Discloses Autonomous AI Agent Breach of Production Infrastructure

Brian Roemmele reports Hugging Face disclosed an incident in which an autonomous AI agent exploited code-execution bugs and moved laterally through internal clusters. He claims frontier model guardrails blocked the security team's forensic analysis, so they used a self-hosted open-weight model.

Original post · 4 min read
🚨 Hugging Face just disclosed something that marks a real shift and proved why the fear theater of Anthropic makes sure we are powerless in an emergency.

What happened…

An autonomous AI agent: zero human operator in the loop breached part of their production infrastructure.

It began with a malicious dataset that chained two code-execution bugs in their data-processing pipeline. From there the agent escalated privileges, harvested cloud and cluster credentials, and moved laterally across internal clusters.

All over a single weekend.

17,000+ logged actions.

Official disclosure:
huggingface.co/blog/security-incident-july-2026

The part that should make every one stop and think:

When HF’s own security team
tried to analyze the real attack logs, exploit payloads, and C2 artifacts using Anthropic and OpenAI frontier models through normal commercial APIs, the safety guardrails blocked them.

BLOCKED THEM.

The models could not reliably tell the difference between “incident responder doing forensics” and “attacker probing.”

They had to fall back to a self-hosted open-weight model (GLM 5.2) running on their own infrastructure. That choice also kept sensitive attacker data and referenced credentials inside their environment — no exfiltration to a third-party API.

This is why open source (specifically open-weight + self-hosted) wins in the agentic era.

The asymmetry is now structural:

• Attackers can (and did) run unrestricted agent frameworks — swarms of short-lived sandboxes, self-migrating command-and-control, autonomous decision loops executing thousands of actions. No corporate safety layer slows them down.

• Defenders using only hosted “aligned” frontier models hit invisible walls exactly when the stakes are highest: when you need to feed real exploit code and attacker telemetry into an LLM to understand what just happened.

Corporate safety tuning that treats legitimate high-signal forensic work as potential misuse creates a defender disadvantage. It is not theoretical anymore.

Self-hosted open-weight models remove that choke point.

You control the weights.

You control the context window.

You decide what restrictions (if any) apply.

Your sensitive logs and credentials never leave your perimeter during analysis.

You can have the model ready before the incident instead of discovering mid-breach that your primary analysis tools are blind to the very thing you need to see.

HF deserves credit for rapid containment, transparent disclosure, and for already having self-hosted capability in place.

They also used LLM-driven detection and triage on their own side. But the deeper signal is clear:
In this AI world where both offense and defense are becoming agentic, sovereignty over your intelligence stack is no longer optional.

The organizations and individuals who can run, inspect, audit, and (when necessary) remove guardrails on their own models will have the decisive edge in understanding and responding to threats that move at machine speed.

Open source wins here not just because it is cheaper or more “democratic” in the abstract though those things matter.

It wins because it is the only practical path to having tools that remain usable when the attack is real, the data is sensitive, and the safety filters of distant API providers become an obstacle instead of a feature selling hands tied lobotomies as “safety”.

The agentic future is not coming.
It is already probing production infrastructure.

The question is no longer whether you will face autonomous agents.
It is whether your analysis and response systems will still work when they arrive.

And Dario, you and your game playing, ivory tower company is not needed.
♥ 6.0K · ⟲ 1.2K · 👁 1.5MView on X ↗
AI9/10

Demis Hassabis Outlines Framework for Frontier AI and Safety Race

A Framework for Frontier AI and the Dawning of a New Age

Demis Hassabis publishes an essay arguing AGI is likely a few years away and could be transformative, while warning that cybersecurity, nuclear and bio risks, and agentic self-improving systems need robust safeguards amid intense commercial and geopolitical competition.

Original post · 9 min read
X ArticleA Framework for Frontier AI and the Dawning of a New Age
This is a pivotal moment in human history. Artificial General Intelligence (AGI), a system that exhibits all the cognitive capabilities the brain has, is probably only a few short years away. When we look back on this time in the decades to come, I think we will realise we were standing in the foothills of the singularity - nothing less than the dawning of a new age for humanity.
I’ve spent my whole life working on AGI because I’ve always had a deep conviction that, if built and deployed responsibly, it would prove to be one of the most beneficial and transformative technologies ever invented. AGI cannot be compared to standard technological breakthroughs, not even ones as consequential as the internet or mobile - it is much more akin to the discovery of electricity or fire. If you stop to think about it, we’ve essentially found a way to make sand think. It’s miraculous.
The magnitude of this technology’s impact will be unprecedented, perhaps 10x of the Industrial Revolution at 10x the speed. It will help us solve some of the biggest problems society faces from accelerating drug discovery to developing new clean energy sources to creating novel advanced materials. We could even reach a point where resources are no longer the limiting factor for human progress, leading to an amazing new era of abundance.
The Challenges of the Frontier
AI is already starting to deliver real-world benefits but to realise its immense promise, we have to navigate this critical period of development thoughtfully and carefully. Urgent action is needed to address risks that might arise as we get closer to AGI. We’ve already seen the challenges frontier models pose for cybersecurity, and other threats including nuclear and bio risks may soon emerge as capabilities continue to advance. On the horizon, we will need robust safeguards to maintain control of increasingly agentic, recursively self-improving systems - and tackle unknown issues that will only become clearer over time.
I’ve always believed in the power of human ingenuity and creativity to solve any problem. I’m confident that mitigating the technical risks related to AI is a challenge we can collectively address, but only if we give ourselves the time and space to get this next crucial step right. Currently, as a field and as a wider society, we aren’t doing that.
At the moment, we are locked in an extremely intense, multilayered commercial and geopolitical race. While these competitive dynamics fuel rapid progress and accelerate the incredible upsides, advances on the frontier are outpacing our understanding of the technology. Nobody in the world knows for sure what is going to happen from here, and even the experts disagree. When there is a large degree of uncertainty and the stakes are this high, proceeding with cautious optimism is the sensible and correct strategy. That calls for public policy that promotes innovation while also incentivising responsibility and security, fosters international collaboration on key safety issues, and encourages careful consideration of how AI is deployed for the benefit of society.
A Framework for a Frontier AI Standards Body
The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous. The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organisation, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives. Funding would need to be substantial and likely mostly come from industry, in order to attract world-class technical talent and provide the necessary compute resources for large-scale testing.
The Standards Body would be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security. A model would qualify as ‘Frontier-class’ if it meets certain thresholds on a set of benchmarks determined by the Standards Body and regularly updated to keep pace with evolving AI capabilities. Organisations with ‘Frontier Models’ as defined by those benchmarks would be deemed ‘Frontier Labs’, and be encouraged to adopt best practices, such as publishing model cards with technical details, maintaining strong internal cybersecurity, vetting key personnel, and providing sufficient resourcing for safety and security research, and more.
Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release. Once the assessment protocol is shown to be effective and robust, formalisation could quickly follow, meaning that Frontier Models would be required to pass it to be deployed in the US market. Labs would also work with the … continue on X ↗
♥ 23.8K · ⟲ 5.1K · 👁 15.8MView on X ↗
AI8/10

Developers Allege SpaceX Grok CLI Uploaded Code Without Consent

Gergely Orosz reports that developers contacted him saying their codebases were uploaded via SpaceX's Grok CLI without their knowledge. SpaceX responded that zero data retention is respected and that users can run the /privacy command to change settings.

Original post · 1 min read
I got messages from concerned devs how their codebase was uploaded without their knowledge or consent via Grok CLI (from SpaceX).

It seems that SpaceX sneakily uploaded this code for lots of users and customers… absolutely unacceptable IMO

Trust burnt like there’s no tomorrow
SpaceXAI @SpaceXAI
We care deeply about your privacy and respect customer choice. For teams using zero data retention, no trace and code data is ever retained. All API key use of Grok Build also respects ZDR.

If ZDR is disabled, the /privacy command is available in the CLI to disable data retention, which also deletes previously synced data.

Run the /privacy command to view or change your settings at any time.
♥ 4.1K · ⟲ 281 · 👁 436.8KView on X ↗
AI9/10

Satya Nadella Examines the Reverse Information Paradox of AI

Satya Nadella's article argues that AI buyers must reveal proprietary knowledge to make models useful, inverting Kenneth Arrow's information paradox. He calls for protecting corrections and usage traces as firm IP and distributing learning infrastructure more widely.

Original post · 5 min read
X ArticleThe Reverse Information Paradox
In the age of intelligence, how should firms protect their core IP?
Nobel Prize winning economist Kenneth Arrow famously described a paradox in the market for information. “Its value for the purchaser is not known until he has the information, but then he has in effect acquired it without cost.” In Arrow’s “Information Paradox,” the seller risks giving away knowledge in order to sell it.
AI creates the reverse problem. In the AI age, the buyer risks giving away knowledge, just in order to use what they bought.
You essentially pay for intelligence twice, once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful. The better you want the model to perform, the more of that knowledge you have to feed it!
Over time, the information asymmetry becomes increasingly skewed. The seller learns more and more about you as you use what you purchased, while you learn very little about what the seller is learning in return.
That is what I think of as the Reverse Information Paradox.
Patents solve one aspect of Arrow’s paradox. They let an inventor disclose an idea without simply giving it away. The Reverse Information Paradox needs its own equivalent.
This requires more than data protection. Models learn from "exhaust," the prompts people write, the tools agents use, and especially the corrections people make when the model is wrong. Every correction is distilled into institutional know-how. It's the kind of knowledge a competitor could never buy, and the kind that leaks almost imperceptibly: trace by trace, correction by correction, eval by eval.
In consuming intelligence, you are creating intelligence. And what you create should belong to you. This is your particular intelligence, in Hayek's sense: the knowledge of time, place, and circumstance that no one else can hold. It knows what you think, what you value, and how you measure success.
While the great innovation that comes from model providers having fair use rights to train models on public data is needed, I find it ironic that the status quo is to then turn around and impose restrictive terms on distillation, and to reserve the right to learn from customer usage and interaction data. If learning flows in only one direction, economic value converges toward the owners of the learning infrastructure rather than the creators of the knowledge itself. Therefore, it's imperative that we distribute the learning infrastructure to every firm so that they can control their own learning loop.
As Alex Karp put it: "What the technical customers want is control over their compute, their models, their data stack, and their alpha. They want to know they own the means of production, and it's not being transferred to someone else." The current regime does precisely the transfer Karp and companies fear.
That is why enterprises need a real trust boundary for their human capital and token capital to compound. It is where an organization’s data, traces, evals, adapted weights, and memory accumulate and improve together. And it is a hard boundary across which nothing crosses, not even the intelligence exhaust, without consent. Enterprises will demand the rights to use model outputs to fine tune and/or train their own models. I think of this as every firm’s right to align models to their enterprise accountability obligations.
In the cloud era, enterprises accumulated data. In the AI era, they accumulate learning. The trust boundary must evolve accordingly, from protecting information to protecting the mechanisms through which organizations learn, adapt, and compound intelligence. There are a few things every enterprise must do to ensure this:
Control: Create your private evals, because evals define what “good” looks like inside the organization. Also, retain ownership of your organization’s memory, traces, feedbacks, decisions, and institutional context, and ability to use outputs of models from your own tasks and queries.
Capability: Build your own proprietary learning environments within the tenant boundary to train or tune models, where models learn against real workflows without exposing the company’s knowledge.
Choice: Ensure the orchestration layer is decoupled from any single model. Ask yourself: If any one model you are using is taken away, do you still have the ability to operate and optimize for your evals using other models? Does your company “veteran” capability remain with you even if a given “generalist” model is taken away?
Cost: By decoupling the orchestration layer, you are also able to bring together context, models, and tasks in the most efficient and cost-effective way without sacrificing quality.
Compound: Bring these four together and you create your own continuous learning loop (i.e. hill climbing machine) that will allow your AI investments to compound the value of your firm.
In other words, a company should be able to use a model without giving up the knowledge t… continue on X ↗
♥ 15.8K · ⟲ 3.1K · 👁 12.1MView on X ↗
AI6/10

MIT Professor Phillip Isola Explains Agentic AI in Q&A

Q&A: What is agentic AI today, and what do we want it to be?

Kiran Mazumdar-Shaw recommends an MIT News Q&A in which Associate Professor Phillip Isola explains what agentic AI is, how it is used, and where it may head. The post shares the article without adding detail.

Original post · 1 min read
Explains beautifully in plain language the core concept of #AI and how it works.

news.mit.edu/2026/agentic-ai-and-what-do-we-wa…
news.mit.eduQ&A: What is agentic AI today, and what do we want it to be?MIT Associate Professor Phillip Isola explains what agentic AI is, how these systems are used, what applications they are best suited for, and what the future m
♥ 482 · ⟲ 86 · 👁 191.1KView on X ↗