Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Edition of Thursday, September 3, 2026

2 stories

AI8/10

Austen Allred Lists Bottlenecks Keeping AI From Self-Training

Austen Allred shares a reading list explaining why AI models cannot yet train themselves, pointing to bottlenecks in RL environments, evaluation, verifiers and a shortage of human text data. The list includes links on RL environment costs of $20k to $300k each and benchmarks like SWE-bench and OSWorld.

Original post · 2 min read
This is an excellent question. Why aren’t AI models just training themselves already?

They theoretically can, and kind of are, but they don’t have the data/evals/gyms required to do so.

A short reading list:

Bottleneck is the environment, not compute
medium.com/@shuchaobi/ais-next-bottleneck-isn-…

RL envs cost real money ($20k–$300k/env, and this is for simulated ones which are just kinda crappy IMO)
epoch.ai/gradient-updates/state-of-rl-envs

We’re running out of human text
epoch.ai/publications/will-we-run-out-of-data-…

Eval is the bottleneck
ysymyth.github.io/The-Second-Half/

Verifier’s law
jasonwei.net/blog/asymmetry-of-verification-an…

What labs buy: Foody on RL envs
youtube.com/watch?v=a00xIn5kwhM

Economy as RL environment machine
mercor.com/blog/the-economy-will-become-an-rl-…

APEX-Agents generalization
mercor.com/blog/generalization-results-from-tr…

Etna: ~$1B/yr on external data, supply-constrained
x.com/hannahhaina/status/2090519081279705359

Surge Tuesday (can it get through a workday?)
surgehq.ai/blog/tuesday-frontier-work-index

Dario: task/process distribution, not more web text
dwarkesh.com/p/dario-amodei-2

OSWorld 2.0 (~20% on long workflows)
osworld-v2.xlang.ai/
arxiv.org/abs/2606.29537

SWE-bench = ticket + repo + tests
swebench.com

Karpathy: sucking supervision through a straw
dwarkesh.com/p/andrej-karpathy

Ilya: peak data / one internet
reuters.com/technology/artificial-intelligence…
Robert Sterling @RobertMSterling
Might be a dumb question, but as we reach AGI and AI becomes smarter than humans, and as the frontier labs compete for market share in a winner-takes-all industry, what’s to stop them from letting their AI models program their own updates and recursively self-improve?

And what does that mean for us, the humans now watching from the sidelines as AI models become more intelligent, more powerful, and less comprehensible to us, at rates that accelerate continuously, not just month by month or day by day, but millisecond by millisecond?

At that point, how do we even understand the inner workings …
♥ 70 · ⟲ 2 · 👁 25.6KView on X ↗
AI6/10

Teresa Torres Publishes Guide to AI Evals for Product Teams

Teresa Torres argues that AI evals, methods for measuring whether an AI product performs well, should be a discovery habit for product teams. She links to a new practical guide she wrote for non-engineers.

Original post · 1 min read
AI evals have been the "it" skill for product teams for over a year. I've even called evals a new discovery habit.

But I still meet product teams who only have a vague idea of what evals are. And it's not their fault. Most of the writing on this topic is intended for engineers or just isn't specific enough.

I recently created an in-depth eval guide to explain what evals are and why product teams can and should create them. I did my best to make it practical, hands-on, and easy to follow.

AI evals (short for evaluations) are methods for measuring whether an AI product or workflow is performing well. Evals give teams confidence that their AI applications are doing what they expect them to do. They help teams maintain quality and catch issues before they reach users.

Similar to other discovery habits like interviewing and assumption testing, evals can act as a feedback loop to ensure we are on the right track.

If you want to learn more about this new discovery habit, explore my new guide: producttalk.org/ai-evals/
♥ 292 · ⟲ 24 · 👁 63.5KView on X ↗