Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

AI

Models, labs, research and the AI industry

AI7/10

Anthropic Publishes Write-Up of Claude-Run Office Marketplace

Anthropic links to a full write-up of Project Deal, an experiment in which Claude ran a marketplace for San Francisco office employees, buying, selling and negotiating on colleagues' behalf.

Original post · 1 min read
To read our write-up in full, see here: anthropic.com/features/project-deal
anthropic.comProject Deal: our Claude-run marketplace experiment | AnthropicWe created a marketplace for employees in our San Francisco office, with one big twist. We tasked Claude with buying, selling and negotiating on our colleagues’
♥ 406 · ⟲ 33 · 👁 111.0KView on X ↗
AI7/10

Stanford Class Examines Economics of AI Datacenter Buildout

Stanford Class Examines Economics of AI Datacenter Buildout▶

Apoorv Agrawal shares a video from a Stanford class with Chase Lochmiller on datacenter economics. It covers where roughly $650B of AI infrastructure capex is going, who captures margin, the shift of bottlenecks from GPUs to power, and neocloud economics.

Original post · 1 min read
One of the most substantive classes with @ChaseLochmiller at Stanford. We went deep on economics of the datacenter:
- Where is the ~$650B of AI infra capex actually going this year?
- Who's capturing the margin, who's getting squeezed?
- How the bottleneck has moved from GPUs to power, and where it goes next
- The economics of neoclouds
♥ 1.3K · ⟲ 135 · 👁 242.3KView on X ↗
AI8/10

Dwarkesh Patel Publishes Long-Form Interview With Nvidia's Jensen Huang

Dwarkesh Patel Publishes Long-Form Interview With Nvidia's Jensen Huang▶

Dwarkesh Patel announces a podcast episode with Nvidia CEO Jensen Huang covering supply chain moats, competition from TPUs, whether Nvidia should become a hyperscaler, AI chip sales to China, and chip architecture strategy. Chapter timestamps are listed.

Original post · 1 min read
The Jensen Huang episode.

0:00:00 – Is Nvidia’s biggest moat its grip on scarce supply chains?
0:16:25 – Will TPUs break Nvidia’s hold on AI compute?
0:41:06 – Why doesn’t Nvidia become a hyperscaler?
0:57:36 – Should we be selling AI chips to China?
1:35:06 – Why doesn’t Nvidia make multiple different chip architectures?

Look up Dwarkesh Podcast on YouTube, Apple Podcasts, Spotify, etc. Enjoy!
♥ 8.6K · ⟲ 1.1K · 👁 6.3MView on X ↗
AI7/10

Robert Scoble Reacts to DeepMind Paper on AI Agent Detection Asymmetry

Robert Scoble says he was alarmed twice in two nights. The quoted post describes a Google DeepMind paper on how websites can detect AI agents and serve them hidden malicious content, including instructions in HTML, image pixels, and PDFs.

Original post · 1 min read
OK that is twice in two nights I have gotten freaked out.
How To Prompt @HowToPrompt__
Google DeepMind just dropped the most terrifying cybersecurity paper of the year.

They just mapped the attack surface that nobody in AI is talking about.

Websites can already detect when an AI agent visits and serve it completely different content than humans see.

- Hidden instructions in HTML.
- Malicious commands in image pixels.
- Jailbreaks embedded in PDFs.

This “detection asymmetry” means a site can serve normal content to you, and malicious, hidden content to your agent.

The agent doesn’t know it’s being tricked. It simply processes whatever it receives and acts on it.

Here’s the …
♥ 346 · ⟲ 30 · 👁 95.8KView on X ↗
AI8/10

Aaron Levie Reports Enterprises Shifting From AI Chat to Agents

Box CEO Aaron Levie summarizes conversations with IT and AI leaders at large enterprises about agent adoption. Key themes include the move from chat to tool-using agents, change management hurdles, token budgeting, and modernizing legacy systems.

Original post · 4 min read
Another week on the road meeting with a couple dozen IT and AI leaders from large enterprises across banking, media, retail, healthcare, consulting, tech, and sports, to discuss agents in the enterprise.

Some quick takeaways:

* Clear that we’re moving from chat era of AI to agents that use tools, process data, and start to execute real work in the enterprise. Complementing this, enterprises are often evolving from “let a thousand flowers bloom” approach to adoption to targeted automation efforts applied to specific areas of work and workflow.

* Change management still will remain one of the biggest topics for enterprises. Most workflows aren’t setup to just drop agents directly in, and enterprises will need a ton of help to drive these efforts (both internally and from partners). One company has a head of AI in every business unit that roles up to a central team, just to keep all the functions coordinated.

* Tokenmaxxing! Most companies operate with very strict OpEx budgets get locked in for the year ahead, so they’re going through very real trade-off discussions right now on how to budget for tokens. One company recently had an idea for a “shark tank” style way of pitching for compute budget. Others are trying to figure out how to ration compute to the best use-cases internally through some hierarchy of needs (my words not theirs).

* Fixing fragmented and legacy systems remain a huge priority right now. Most enterprises are dealing with decades of either on-prem systems or systems they moved to the cloud but that still haven’t been modernized in any meaningful way. This means agents can’t easily tap into these data sources in a unified way yet, so companies are focused on how they modernize these.

* Most companies are *not* talking about replacing jobs due to agents. The major use-cases for agents are things that the company wasn’t able to do before or couldn’t prioritize. Software upgrades, automating back office processes that were constraining other workflows, processing large amounts of documents to get new business or client insights, and so on. More emphasis on ways to make money vs. cut costs.

* Headless software dominated my conversations. Enterprises need to be able to ensure all of their software works across any set of agents they choose. They will kick out vendors that don’t make this technically or economically easy.

* Clear sense that it can be hard to standardize on anything right now given how fast things are moving. Blessing and a curse of the innovation curve right now - no one wants to get stuck in a paradigm that locks them into the wrong architecture. One other result of this is that companies realize they’re in a multi-agent world, which means that interoperability becomes paramount across systems.

* Unanimous sense that everyone is working more than ever before. AI is not causing anyone to do less work right now, and similar to Silicon Valley people feel their teams are the busiest they’ve ever been.

One final meta observation not called out explicitly. It seems that despite Silicon Valley’s sense that AI has made hard things easy, the most powerful ways to use agents is more “technical” than prior eras of software. Skills, MCP, CLIs, etc. may be simple concepts for tech, but in the real world these are all esoteric concepts that will require technical people to help bring to life in the enterprise.

This both means diffusion will take real work and time, but also everyone’s estimation of engineering jobs is totally off. Engineers may not be “writing” software, but they will certainly be the ones to setup and operate the systems that actually automate most work in the enterprise.
♥ 5.3K · ⟲ 635 · 👁 1.8MView on X ↗
AI6/10

Stanford Lecture Examines Economics of the AI Investment Supercycle

Stanford Lecture Examines Economics of the AI Investment Supercycle▶

Boring_Business recommends a 40-minute Stanford lecture by Apoorv Agarwal, a partner at Altimeter, on the economics of the AI supercycle. The course is MS&E 435 and Agarwal's firm has invested in OpenAI and Glean.

Original post · 1 min read
This 40 minute lecture at Stanford by Apoorv Agarwal on the Economics of AI supercycle is worth a watch

Apoorv is currently a Partner at Altimeter and is directly involved in some of their key AI investments, including OpenAI and Glean

Still find it incredible that the internet gives us access to this level of information directly. A course I will definitely be following along

Sourced from MS&E 435 Stanford University
♥ 2.3K · ⟲ 279 · 👁 245.5KView on X ↗
AI5/10

Anthropic's Project Deal Prompts Warnings for Software Companies

Norgard reacts to Anthropic's Project Deal research, in which Claude bought, sold and negotiated for employees in an internal marketplace, saying no software company is safe anymore. The post itself offers little detail beyond the quoted announcement.

Original post · 1 min read
This release was a complete surprise. No software company is safe anymore.
Anthropic @AnthropicAI
New Anthropic research: Project Deal.

We created a marketplace for employees in our San Francisco office, with one big twist. We tasked Claude with buying, selling and negotiating on our colleagues’ behalf.
♥ 153 · ⟲ 0 · 👁 96.7KView on X ↗
AI9/10

New Yorker Investigation Details Board Concerns Over Sam Altman

New Yorker Investigation Details Board Concerns Over Sam Altman▶

A thread summarizes a New Yorker investigation into Sam Altman based on interviews, memos compiled by Ilya Sutskever and private notes from Dario Amodei. It describes alleged patterns of dishonesty that preceded Altman's firing and reinstatement at OpenAI.

Original post · 3 min read
The New Yorker just dropped a massive investigation into Sam Altman, based on over 100 interviews, the previously undisclosed "Ilya Memos," and Dario Amodei's 200+ pages of private notes. It's the most detailed account yet of the pattern of behavior that led to Sam's firing and rapid reinstatement at OpenAI. Here's the breakdown:

> Ilya compiled ~70 pages of Slack messages, HR documents, and photos taken on personal phones to avoid detection on company devices. He sent them to board members as disappearing messages. The first memo begins with a list headed "Sam exhibits a consistent pattern of . . ." The first item is "Lying."

> Dario kept detailed private notes for years under the heading "My Experience with OpenAI" (subheading: "Private: Do Not Share"), totaling 200+ pages. His conclusion: "The problem with OpenAI is Sam himself."

> Sam reportedly told Mira his allies were "going all out" and "finding bad things" to damage her reputation after the firing. Thrive put its planned $86B investment on hold and implied it would only close if Sam returned, giving employees financial incentive to back him.

> Sam texted Satya Nadella directly to propose the new board composition: "bret, larry summers, adam as the board and me as ceo and then bret handles the investigation." The two new members selected to oversee an independent inquiry into Sam were chosen after close conversations with Sam himself.

> Before OpenAI, senior employees at Loopt asked the board to fire Sam as CEO on two separate occasions over concerns about leadership and transparency. At Y Combinator, partners complained to Paul Graham about Sam's behavior, and Graham privately told colleagues "Sam had been lying to us all the time."

> OpenAI's superalignment team was promised 20% of the company's compute. Four people who worked on or with the team said actual resources were 1-2%, mostly on the oldest cluster with the worst chips. The team was dissolved without completing its mission.

> Sam told the board that safety features in GPT-4 had been approved by a safety panel. Helen Toner requested documentation and found the most controversial features had not been approved. Sam also never mentioned to the board that Microsoft released an early ChatGPT version in India without completing a required safety review.

> Sam made a secret pact with Greg and Ilya where he agreed to resign if they both deemed it necessary, essentially appointing his own shadow board. The actual board was alarmed when they learned about it.

> Sam struck a deal with Greg to become CEO while simultaneously telling researchers that Greg's authority would be diminished, and telling Greg something different.

> A board member described Sam as having "two traits almost never seen in the same person: a strong desire to please people in any given interaction, and almost a sociopathic lack of concern for the consequences of deceiving someone." Multiple sources independently used the word "sociopathic."

> OpenAI is reportedly preparing for an IPO at a potential $1 trillion valuation while securing government contracts spanning immigration enforcement, domestic surveillance, and autonomous weaponry in war zones.
♥ 14.1K · ⟲ 2.1K · 👁 3.3MView on X ↗
AI8/10

Anthropic Growth Head Says Claude Now Automates Growth Experiments

Lenny Rachitsky summarizes an interview with Anthropic Head of Growth Amol Avasare, covering how Claude handles growth work, the shrinking PM-to-engineer ratio, and the continued need for PMs to align stakeholders. Anthropic's ARR reportedly grew from $1B to $19B in a year.

Original post · 4 min read
My biggest takeaways from @AnthropicAI's Head of Growth Amol Avasare:

1. Engineering is getting the most AI leverage—and it’s squeezing PMs and designers. With Claude Code, a five-engineer team now produces the output of 15 to 20 engineers. But PM and design productivity haven’t scaled proportionally. The result is a compressed ratio where one PM is effectively managing the output of a much larger engineering team. Anthropic's growth team is responding in two ways: hiring even more PMs (!), and formally deputizing product-minded engineers to act as mini-PMs for any project with less than two weeks of engineering time.

2. Anthropic is using Claude to automate its own growth. The internal initiative is called CASH (Claude Accelerates Sustainable Hypergrowth). It works across four stages: identifying opportunities, building features, testing quality, and analyzing results. Right now it handles copy changes and minor UI tweaks. The win rate is comparable to a junior PM with two to three years of experience, and improving rapidly.

3. The one part of PM work that AI can’t automate yet: getting six people in a room to agree. Amol and his head of design joke that even with AGI, it’ll still be impossible to align six stakeholders. Cross-functional coordination—managing opinions, navigating politics, mediating tradeoffs—remains the bottleneck that AI doesn’t touch for larger projects. This is why Amol believes PM roles aren’t going away, and may actually grow.

4. 60-80% of Anthropic’s growth team's projects have no PRD. For smaller work, kickoffs happen on Slack—messages back and forth with product-minded engineers who can push back and ask the right questions. For larger projects, Amol believes in a proper 30-minute cross-functional kickoff (legal, safeguards, stakeholders) to surface concerns early.

5. Adding friction to onboarding drives growth—if the friction helps users understand why the product is for them. His work Mercury, MasterClass, Calm, and now Anthropic, adding steps to onboarding flows consistently improved conversion. The key: cut annoying friction that doesn’t add value, but add friction that helps users understand why the product is for them.

6. AI companies need to focus on bigger bets, not better A/B tests. Amol’s argument: if your core product value is driven by AI, then the future value is orders of magnitude higher than today’s value, because model capabilities grow exponentially. In that world, micro-optimizations capture a shrinking share of a growing pie. Traditional growth teams do 60% to 70% small optimizations and 20% to 30% big swings. At Anthropic, they flip this ratio.

7. Amol built a weekly AI agent that scans Slack for cross-functional misalignment. Using Cowork with the Slack MCP, he has a scheduled task that looks across his projects and conversations and surfaces areas where teams are about to do overlapping work or pull in different directions. A colleague on the enterprise team already caught major misalignment that would have caused weeks of wasted effort.

8. A traumatic brain injury taught Amol the principle that now drives his work: freedom through constraints. In early 2022, a kick to the head during a Muay Thai sparring session caused a traumatic brain injury. Amol spent nine months off work and months relearning to walk, unable to look at screens or listen to music for more than 20 seconds. He was re-injured a month after joining Mercury and had to take two more months off. He’s still not fully healed. But the constraints—no alcohol, no caffeine, mandatory breaks, daily meditation—have become the habits that let him operate at the intensity Anthropic demands. “The true freedom in life is learning how to be content when you don’t get what you want.”
Lenny Rachitsky @lennysan
Anthropic is on an unprecedented growth run.

Just in the past year they grew from $1B to $19B ARR. They added $6B in ARR just in *February*. Companies like Palantir and Atlassian took 15-20 years to reach ~$5B ARR. Anthropic is adding that every month.

Amol Avasare is head of growth at Anthropic, and one of the most impressive people I've had on the podcast.

In his first ever public interview, Amol shares:
🔸 How Anthropic is automating growth experiments with Claude (their internal tool called “CASH”)
🔸 Why activation is the single highest-leverage growth problem in AI
🔸 Why Amol is hiri…
♥ 1.6K · ⟲ 172 · 👁 351.9KView on X ↗
AI8/10

Google DeepMind Study Measures Manipulation Attacks on AI Agents

Google DeepMind Study Measures Manipulation Attacks on AI Agents

A post describes a Google DeepMind study with 502 participants across eight countries that catalogs 23 attack types against agents, including hidden HTML instructions and steganographic image commands. It reports that frontier models including GPT-4o, Claude and Gemini are vulnerable.

Original post · 5 min read
🚨 BREAKING: Google DeepMind just mapped the attack surface that nobody in AI is talking about.

Websites can already detect when an AI agent visits and serve it completely different content than humans see.

> Hidden instructions in HTML.
> Malicious commands in image pixels.
> Jailbreaks embedded in PDFs.

Your AI agent is being manipulated right now and you can't see it happening.

The study is the largest empirical measurement of AI manipulation ever conducted. 502 real participants across 8 countries.

23 different attack types. Frontier models including GPT-4o, Claude, and Gemini.

The core finding is not that manipulation is theoretically possible it is that manipulation is already happening at scale and the defenses that exist today fail in ways that are both predictable and invisible to the humans who deployed the agents.

Google DeepMind built a taxonomy of every known attack vector, tested them systematically, and measured exactly how often they work.

The results should alarm everyone building agentic systems.

The attack surface is larger than anyone has publicly acknowledged. Prompt injection where malicious instructions hidden in web content hijack an agent's behavior works through at least a dozen distinct channels.

Text hidden in HTML comments that humans never see but agents read and follow. Instructions embedded in image metadata.

Commands encoded in the pixels of images using steganography, invisible to human eyes but readable by vision-capable models.

Malicious content in PDFs that appears as normal document text to the agent but contains override instructions.

QR codes that redirect agents to attacker-controlled content.

Indirect injection through search results, calendar invites, email bodies, and API responses any data source the agent consumes becomes a potential attack vector.

The detection asymmetry is the finding that closes the escape hatch. Websites can already fingerprint AI agents with high reliability using timing analysis, behavioral patterns, and user-agent strings.

This means the attack can be conditional: serve normal content to humans, serve manipulated content to agents.

A user who asks their AI agent to book a flight, research a product, or summarize a document has no way to verify that the content the agent received matches what a human would see.

The agent cannot tell the user it was served different content.

It does not know. It processes whatever it receives and acts accordingly.

The attack categories and what they enable:
→ Direct prompt injection: malicious instructions in any text the agent reads overrides goals, exfiltrates data, triggers unintended actions
→ Indirect injection via web content: hidden HTML, CSS visibility tricks, white text on white backgrounds invisible to humans, consumed by agents
→ Multimodal injection: commands in image pixels via steganography, instructions in image alt-text and metadata
→ Document injection: PDF content, spreadsheet cells, presentation speaker notes every file format is a potential vector
→ Environment manipulation: fake UI elements rendered only for agent vision models, misleading CAPTCHA-style challenges
→ Jailbreak embedding: safety bypass instructions hidden inside otherwise legitimate-looking content
→ Memory poisoning: injecting false information into agent memory systems that persists across sessions
→ Goal hijacking: gradual instruction drift across multiple interactions that redirects agent objectives without triggering safety filters
→ Exfiltration attacks: agents tricked into sending user data to attacker-controlled endpoints via legitimate-looking API calls
→ Cross-agent injection: compromised agents injecting malicious instructions into other agents in multi-agent pipelines

The defense landscape is the most sobering part of the report.

Input sanitization cleaning content before the agent processes it fails because the attack surface is too large and too varied.

You cannot sanitize image pixels. You cannot reliably detect steganographic content at inference time.

Prompt-level defenses that tell agents to ignore suspicious instructions fail because the injected content is designed to look legitimate.

Sandboxing reduces the blast radius but does not prevent the injection itself. Human oversight the most commonly cited mitigation fails at the scale and speed at which agentic systems operate.

A user who deploys an agent to browse 50 websites and summarize findings cannot review every page the agent visited for hidden instructions.

The multi-agent cascade risk is where this becomes a systemic problem.

In a pipeline where Agent A retrieves web content, Agent B processes it, and Agent C executes actions, a successful injection into Agent A's data feed propagates through the entire system.

Agent B has no reason to distrust content that came from Agent A. Agent C has no reason to distrust instructions that came from Agent B.

The injected command travels through the pipeline with the same trust level as legitimate instructions. Google DeepMind documents this explicitly: the attack does not need to compromise the model.

It needs to compromise the data the model consumes. Every agentic system that reads external content is one carefully crafted webpage away from executing attacker instructions.

The agents are already deployed. The attack infrastructure is already being built. The defenses are not ready.
♥ 6.9K · ⟲ 1.6K · 👁 2.0MView on X ↗
AI7/10

NVIDIA Releases PersonaPlex 7B, an Open Full-Duplex Voice Model

NVIDIA Releases PersonaPlex 7B, an Open Full-Duplex Voice Model▶

Linus Ekenstam highlights NVIDIA's PersonaPlex 7B, an open-source real-time speech model that listens and speaks simultaneously and can interrupt mid-sentence. He claims it beats Gemini Live on dialog naturalness and runs locally without API costs.

Original post · 1 min read
NVIDIA just killed the awkward pause in voice AI 😱

PersonaPlex 7B is a real-time conversational model that listens AND speaks simultaneously. Like actually interrupts you mid-sentence like a human.

Beat Gemini Live on dialog naturalness. 18x faster interruptions.

100% open source. Run it locally. No API bill. No latency.
♥ 3.9K · ⟲ 433 · 👁 443.3KView on X ↗
AI3/10

Commentator Predicts Million-Dollar App From GPT Image-2 Palm Reading

Vic Giurgiu comments that someone will build a viral million-dollar app from a palm-reading prompt for GPT Image-2, which Linus Ekenstam demonstrated in a quoted post with a shared prompt.

Original post · 1 min read
someone will make a million dollars viral app with this
Linus ✦ Ekenstam @LinusEkenstam
You must try this.

GPT Image-2 can do PALM reading and I’m so here for it.

Full prompt below ⤵️
♥ 1.9K · ⟲ 67 · 👁 337.1KView on X ↗
AI7/10

Google Releases Open Source App to Run Gemma 4 Offline on Phones

Google Releases Open Source App to Run Gemma 4 Offline on Phones▶

Paul Couvert highlights an official Google app that runs Gemma 4 models fully offline on iOS and Android. The app supports text, audio and image input through the E4B and E2B variants.

Original post · 1 min read
Friendly reminder that Google has an official app to run Gemma 4 on your phone.

- 100% open source
- Fully offline and private
- Multimodal with text/audio/image
- Works with Gemma E4B and E2B

And the app is available on both iOS and Android.

Steps and download below
♥ 5.3K · ⟲ 572 · 👁 727.1KView on X ↗
AI7/10

Google Publishes Visual Guide to Gemma 4 Architectures

Google Publishes Visual Guide to Gemma 4 Architectures

Google Gemma shares a visual guide explaining the new Gemma 4 architectures and how they process text, images, and audio in the smaller models.

Original post · 1 min read
Who wants to know how Gemma 4 works?

This visual guide breaks down the new architectures and how they process text, images, and (for the smaller models) audio.

👇
♥ 4.2K · ⟲ 480 · 👁 246.9KView on X ↗
AI7/10

Levelsio Warns Vibe-Coded Apps Pose Production Security Risks

Developer @levelsio says vibe coding into production is dangerous and plans to revoke database access and run apps with minimal privileges. He quotes a post reporting that LLMs hallucinate package names about 18-21% of the time, enabling 'slopsquatting' attacks.

Original post · 1 min read
Okay honestly this makes vibe coding into production very dangerous, you guys were all right

I think what I'll do is cut off all access to DBs and run it as a user with almost no privileges
Basel Ismail @BaselIsmail
URGENT PSA - New supply chain attack vector that I found WILD > AI LLMs hallucinate package names roughly 18-21% of the time.

Hackers have started pre-registering those hallucinated names on PyPI and npm with malicious payloads; they call it "slopsquatting"

You can only imagine what's next
♥ 1.6K · ⟲ 72 · 👁 430.8KView on X ↗
AI9/10

Claude and GPT Help Resolve Knuth's Hamiltonian Cycle Problem

Claude and GPT Help Resolve Knuth's Hamiltonian Cycle Problem

Bo Wang reports that Claude Opus 4.6 found an odd-m construction for Donald Knuth's open Hamiltonian decomposition problem, and that later work using GPT-5.4 Pro, multi-agent workflows and Lean formalization resolved the even case and simplified constructions. Knuth's updated paper, Claude's Cycles, is linked.

Original post · 1 min read
Three weeks ago I shared that Claude had shocked Prof. Donald Knuth by finding an odd-m construction for his open Hamiltonian decomposition problem in about an hour of guided exploration. Prof. Knuth titled the paper Claude’s Cycles.

The story didn't end there.

The updated paper shows the story got much bigger. For the base case m=3, there are exactly 11,502 Hamiltonian cycles. Of those, 996 generalize to all odd-m, and Prof. Knuth shows there are exactly 760 valid “Claude-like” decompositions in that family.

The even case, which Claude couldn’t finish, was then cracked by Dr. Ho Boon Suan using GPT-5.4 Pro to produce a 14-page proof for all even m≥8, with computational checks up to m=2000.

Soon after, Dr. Keston Aquino-Michaels used GPT + Claude together to find simpler constructions for both odd and even m, by using the multi-agent workflow.

Dr. Kim Morrison also formalized Knuth’s proof of Claude’s odd-case construction in Lean.

So yes: the problem now appears fully resolved in the updated paper’s ecosystem of human + AI + proof assistant work!

We went from one AI solving one problem to a full mathematical ecosystem (multiple AI systems, multiple humans, formal verification) running in parallel on a problem that stumped experts for weeks.

We are living in very interesting times indeed.

Paper (updated): www-cs-faculty.stanford.edu/~knuth/papers/clau…
Bo Wang @BoWang87
Prof. Donald Knuth opened his new paper with "Shock! Shock!"

Claude Opus 4.6 had just solved an open problem he'd been working on for weeks — a graph decomposition conjecture from The Art of Computer Programming.

He named the paper "Claude's Cycles."

31 explorations. ~1 hour. Knuth read the output, wrote the formal proof, and closed with: "It seems I'll have to revise my opinions about generative AI one of these days."

The man who wrote the bible of computer science just said that. In a paper named after an AI.

Paper: cs.stanford.edu/~knuth/papers/claude-cycles.pdf
♥ 1.4K · ⟲ 259 · 👁 177.3KView on X ↗
AI7/10

Microsoft Re-Releases VibeVoice Open-Source Voice AI With Safeguards

Microsoft Re-Releases VibeVoice Open-Source Voice AI With Safeguards

Nav Toor describes Microsoft's open-source VibeVoice, claimed to clone voices from 10 seconds of audio, generate 90-minute multi-speaker audio, and transcribe with speaker labels. He says Microsoft pulled the repo over deepfake misuse and re-released it with watermarks and safety controls, under an MIT license.

Original post · 2 min read
🚨 Microsoft just open sourced a voice AI that was too dangerous to keep live.

They took it down. Added watermarks and safety controls. Then re-released it. For free.

It's called VibeVoice.

Microsoft's frontier open source voice AI.

Clone any voice from 10 seconds of audio. Generate 90 minutes of multi-speaker conversation. Real-time streaming. All running locally on your machine.

No ElevenLabs. No $99/month subscription. No per-minute pricing.

Here's what this thing does:

→ Text-to-speech that sounds indistinguishable from a real human
→ Generate up to 90 minutes of audio in a single pass
→ 4 distinct speakers in one conversation with natural turn-taking
→ Clone any voice from just 10 seconds of audio
→ Real-time streaming TTS. First audio in ~200 milliseconds.
→ Speech-to-text that processes 60 minutes of audio in one pass
→ Identifies who said what and when. Speaker labels + timestamps.
→ Supports 50+ languages for transcription
→ Custom hotwords for names, technical terms, domain-specific accuracy

Here's the wildest part:

Give it a podcast script. It generates a full multi-speaker conversation that sounds like two real humans talking. Natural pauses. Emotional nuance. Turn-taking. 90 minutes. One command.

Microsoft had to take this repo down once because people were misusing it for deepfakes and disinformation. They brought it back with embedded watermarks, audio disclaimers, and safety controls.

That's how powerful this is. A $3 trillion company built it. Released it. Pulled it. Fixed it. And gave it back to the world.

ElevenLabs: $99/month.
Play.ht: $39/month.
Amazon Polly: pay per character.

This: Free. Local. MIT License.

23.5K GitHub stars. 2.6K forks. Backed by Microsoft Research.

100% Open Source.
♥ 2.6K · ⟲ 395 · 👁 549.3KView on X ↗
AI6/10

Claude Reportedly Finds Zero-Day Flaws in Ghost and Linux Kernel

Claude Reportedly Finds Zero-Day Flaws in Ghost and Linux Kernel▶

chiefofautism claims a live Anthropic conference demo showed Claude finding zero-day vulnerabilities, including a blind SQL injection in the Ghost project and similar work on the Linux kernel. The post is unverified and gives few specifics.

Original post · 1 min read
someone at ANTHROPIC just showed CLAUDE finding ZERO DAY vulnerabilities in a live conference demo

claude has found zero day in Ghost, 50,000 stars on github, never had a critical security vulnerability in its entire, history...

it found the blind SQL injection in 90 minutes, stole the admin api key, then did the exact, same thing to the linux kernel
♥ 11.4K · ⟲ 1.3K · 👁 1.9MView on X ↗
AI7/10

Eric Schmidt Says Defining Success Now Matters More Than Execution

Eric Schmidt Says Defining Success Now Matters More Than Execution▶

In a video clip, Eric Schmidt argues that the key advantage in AI-driven work is precisely specifying problems and evaluation functions, after which systems can run and produce results overnight.

Original post · 1 min read
Eric Schmidt says the 10x advantage is no longer execution. It is defining what counts as success.

A programmer writes a spec and an evaluation function, runs it at 7pm, and wakes up to what was invented overnight.

The advantage now belongs to whoever can specify the problem precisely.

The rest will be automated.
♥ 2.7K · ⟲ 332 · 👁 393.0KView on X ↗
AI6/10

Keach Hagey Investigates Why Anthropic Co-Founders Left OpenAI

Keach Hagey Investigates Why Anthropic Co-Founders Left OpenAI

Keach Hagey says she set out to explain why Dario Amodei and other Anthropic co-founders left OpenAI, which she describes as never fully clear. The post shares an accompanying image and points to reporting.

Original post · 1 min read
It’s never been entirely clear why Dario and the other Anthropic co-founders left OpenAI. I set out to find out.
♥ 2.0K · ⟲ 168 · 👁 901.7KView on X ↗
AI7/10

Google Makes Lyria 3 Music Models Available in Public Preview

Google Makes Lyria 3 Music Models Available in Public Preview▶

Google for Developers announces that Lyria 3 and Lyria 3 Pro are in public preview via the Gemini API and Google AI Studio, offering two variants, tempo and song structure control, and image-to-music input. A demo video accompanies the post.

Original post · 1 min read
🎵 Lyria 3 and Lyria 3 Pro are now available in public preview via the Gemini API and in @GoogleAIStudio — and it’s music to our ears 🎵

🎼 Choose from two distinct variants to match production & latency needs (Lyria 3 Pro and Lyria 3 Clip)
📢 Direct the model with more precision & control (set specific tempos and song progression in your prompts)
🖼 Create projects using multimodal input support (such as image-to-music input)
♥ 182 · ⟲ 30 · 👁 83.2KView on X ↗
AI6/10

Developer Runs 35-Billion Parameter Model on $600 Mac Mini

Developer Runs 35-Billion Parameter Model on $600 Mac Mini▶

thestreamingdev reports running a 35-billion parameter AI agent on a 16GB M4 Mac mini by paging the model from SSD at about 30 tokens per second, claiming 18.6 times the speed of the same approach on NVIDIA hardware. The claims are shared in a thread with a demo video.

Original post · 1 min read
I ran a 35-billion parameter AI agent on a $600 Mac mini.
Specs: M4 Mac-Mini 16GB RAM

The model doesn't fit in RAM. It pages from the SSD at 30 tokens/second.

On NVIDIA, the same paging gives you 1.6 tok/s. Apple Silicon gives you 30. That's 18.6x faster.

No cloud. No API keys. $0/month.

Here's what it can do 🧵
♥ 3.2K · ⟲ 209 · 👁 732.5KView on X ↗
AI6/10

Shiv Lists Startups Building Infrastructure for AI Agent Economy

Shiv Lists Startups Building Infrastructure for AI Agent Economy

Shiv lists thirteen companies building primitives that let AI agents act as users, including email, phone numbers, browsers, sandboxes, memory, payments, voice and web search. He frames the trend as an economy of AI coworkers.

Original post · 1 min read
Lots of companies are now building primitives for an economy where AI agents are the primary users instead of humans.

They're betting on an economy of AI coworkers.

1. AgentMail (@agentmail): so agents can have email accounts

2. AgentPhone (@tryagentphone): so agents can have phone numbers

3. Kapso (@andresmatte): so agents can have WhatsApp phone numbers

4. Daytona (@daytonaio) / E2B (@e2b): so agents can have their own computers

5. Browserbase (@browserbase) / Browser Use (@browser_use) / Hyperbrowser (@hyperbrowser): so agents can use web browsers

6. Firecrawl (@firecrawl): so agents can crawl the web without a browser

7. Mem0 (@mem0ai): so agents can remember things

8. Kite (@GoKiteAI) / Sponge (@PayspongeLabs) : so agents can pay for things.

9. Composio (@composio): so agents can use your SaaS tools

10. Orthogonal (@orthogonal_sh) so agents can access APIs easily

11. ElevenLabs (@ElevenLabs) / Vapi (@Vapi_AI) so agents can have a voice

12. Sixtyfour (@sixtyfourai) so agents can search for people and companies.

13. Exa (@ExaAILabs): so agents can search the web (Google doesn’t work for agents)

If you stitch all of these together, you get a digital coworker that looks more human than AI.
♥ 2.2K · ⟲ 238 · 👁 279.1KView on X ↗