Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

AI

Models, labs, research and the AI industry

AI8/10

Meta's Alexandr Wang Says Agent Loops Can Outperform 100 Engineers

Meta's Alexandr Wang Says Agent Loops Can Outperform 100 Engineers▶

In a YC conversation with Garry Tan, Meta Chief AI Officer Alexandr Wang described agentic loops with evaluation metrics that let agents complete more work than 100 senior engineers. He said the system is built from simple parts like markdown files, cron jobs and metrics.

Original post · 3 min read
Former Scale AI founder and newly appointed @Meta Chief AI Officer @alexandr_wang dropped a bombshell during his YC conversation with @garrytan, and you could almost hear the tech leadership world go quiet.

He said Meta has already seen this internally:

Build the right agentic loop, give it an evaluation system and metrics that let it optimize itself, and a group of AI agents can complete more work than a team of 100 senior engineers.

And they do it “very easily.”

But the most interesting part wasn’t the 100-engineer comparison.

It was how simple the system underneath it actually is.

You’d expect some insanely complex, almost alien architecture powering a swarm like this.

Instead, Wang described it with a few almost comically basic words:

“Markdown files, cron jobs, goal, metrics, data.”

Once you strip away the hype, the implications for traditional software engineering become pretty clear:

1️⃣ It’s not that the models are magically smarter. The eval loop is doing the heavy lifting.

Traditional approach: humans write prompts, run the code, inspect the output, and hope nothing broke.

Meta’s approach: turn the business goal into something a machine can score automatically.

The agent submits its work. The system runs tests, calculates metrics, finds what’s wrong, and sends it back for another pass. Repeat until it passes.

Nobody has to babysit every step. The metric becomes the supervisor.

2️⃣ The real alpha is burning 1,000x more tokens inside the feedback loop.

A lot of people are still optimizing for the cost of a single AI call.

The frontier labs are playing a different game: spend 1,000x or even 1,000,000x more tokens in the background so agents can constantly review, rerun, challenge, and verify each other’s work until they reach a reliable business outcome.

Token cost is fixed. The payoff is a pipeline that keeps running.

3️⃣ Memory doesn’t need some fancy database.

Persistent memory can live in Markdown files.

Scheduling can be handled by the server’s built-in cron jobs, running overnight.

The simpler the scaffolding, the more robust the system can be. Less infrastructure also means fewer ways for context to fall apart.

This is a pretty brutal change in how technical organizations work.

The ceiling for a tech lead used to be partly about how many people they could manage, how many meetings they could sit through, and how many teams they could coordinate.

The leverage for the next generation of technical leaders may look very different:

Can you turn a messy business objective into a rigorous set of metrics that an AI can evaluate automatically?

If 100 people’s output can be replaced by a few cron jobs, Markdown files, and a well-designed eval loop, the era of “just take the ticket and write the code” is coming to an end.

The people who can design the evals and orchestrate the swarm aren’t just holding a new tool.

They’re effectively running a virtual company.
♥ 511 · ⟲ 58 · 👁 133.6KView on X ↗
AI4/10

Daniel Loeb Praises Muse as Personal AI Agent

Muse — Your Personal AI Agent

Investor Daniel Loeb says he uses Muse, a personal AI agent for shopping, life management and rewards optimization, and asks which companies might be disrupted or benefit.

Original post · 1 min read
I love muse.ai. Been using it for shopping, life management, reminder to buy John Mayer tickets and save money by helping me monetize my rewards programs and cut out duplicate and hidden costs on my credit cards.
What companies do people think get disrupted or benefit?
muse.aiMuse — Your Personal AI AgentTry Muse, your personal AI agent that gets things done. Give Muse a goal or an everyday task and it handles the rest, from finances and health to shopping and t
♥ 384 · ⟲ 3 · 👁 74.3KView on X ↗
AI6/10

Nailthy Tang Demos Jev Model for Real-Time Virtual Outfit Changes

Nailthy Tang Demos Jev Model for Real-Time Virtual Outfit Changes▶

Nailthy Tang showcases an experiment built for Drape with Typesafe, where the Jev model reads a conversation and the user's outfit to pick clothes from a closet and change outfits in real time at about $0.0011 and 620 milliseconds per decision. The post links to Jev's announcement, which claims faster and cheaper performance than frontier models.

Original post · 1 min read
jev is insane 🫣

it makes realtime virtual try-on hauls possible.
built this experiment for Drape with @typesafeai

> i talk
> jev reads transcript + what i'm wearing
> picks from my closet
> changes my outfit in realtime

cost: $0.0011 per decision
time: ~620ms per decision

imagine getting ready like this:
Diogo Almeida @CompleteSkeptic
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?

I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev

• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions

AFAICT the shortest path to AI-based economic revolution
♥ 4.3K · ⟲ 313 · 👁 562.2KView on X ↗
AI6/10

Open-Source Model Laya Reportedly Beats OpenAI Co-Founder's Jev

Jun Song argues that an open-source model, Laya, outperformed Jev, a model an OpenAI co-founder spent three years building, within three days. A linked Hugging Face post supports the claim.

Original post · 1 min read
An OpenAI co-founder spent 3 years quietly building Jev, only for someone to open source a better performing model in just 3 days.

​This is exactly why starting an AI company right now carries way too much risk.
CV.YH @0xCVYH
Laya é um open source do Jev que já veio acima dele.

tá cada vez mais rápido

huggingface.co/convaiinnovations/laya
♥ 2.5K · ⟲ 183 · 👁 308.8KView on X ↗
AI6/10

Prajwal Tomar Explains Jev as a Non-Generative Decision Model

Prajwal Tomar describes Jev, a free AI model built by a ChatGPT co-inventor that scores options rather than generating text. He argues this design avoids hallucination and makes decision-making cheap, and links to a post about building a harness with Jev.

Original post · 1 min read
If you're still confused about what Jev actually does, read this.

The guy who co-invented ChatGPT spent two years building an AI that cannot write a single word. On purpose. Then made it FREE.

Every AI you've used talks to you.

Jev doesn't talk at all. It just decides.

You give it an email or a lead and some options, it spits back how likely each one is. No essay, no explanation.

It literally can't hallucinate because it never writes anything.

The talking part of AI is solved. The deciding part just got stupidly cheap.

That's where all the boring money is.

If the launch went over your head, read this one.
Sydney Runkle @sydneyrunkle
Building a Harness with Jev — Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results, and then continues in that loop until the task is complete.
Agents and LLMs were initially difficult to
♥ 315 · ⟲ 17 · 👁 81.6KView on X ↗
AI4/10

Viral Post Claims Claude Opus 5.5 Built 53% Return Trading Strategy

Viral Post Claims Claude Opus 5.5 Built 53% Return Trading Strategy

Rahul promotes a prompt for Claude Opus 5.5 and quotes a post claiming the model built a trading strategy with 53% returns that beat the S&P 500 in backtesting. The original post says it tested the model on stock research rather than coding benchmarks.

Original post · 1 min read
Send this stock research prompt to Claude Opus 5.5.

It will probably make you a lot of money.

You're welcome:
Rahul @sairahul1
Claude Opus 5.5 Built a Trading Strategy With 53% Returns That Beat the S&P 500 — Claude Opus 5.5 dropped last week

Everyone is testing it on coding benchmarks.
I wanted to test something harder.
Can it actually find a trading strategy that survives a real backtest?
Not stock
♥ 3.4K · ⟲ 359 · 👁 1.2MView on X ↗
AI5/10

Brad Gerstner Predicts Personal AI Assistants Will Reshape Internet Aggregators

Brad Gerstner says a personal AI assistant with memory of each user's preferences is inevitable and will disrupt aggregators in insurance, travel and shopping. He frames AI agents as shifting the fulcrum of value delivery.

Original post · 1 min read
As discussed w @bgurley two yrs ago - a personal, super assistant in every American’s pocket w perfect memory & understanding of my life - my likes & dislikes - able to work non stop & achieve almost anything is now inevitable. It is a huge gift to humanity. It will 10x all of us, give us more magic experiences & save time from drudgery to spend w family & friends. Many internet aggregators (insurance, travel, shopping) will resist but resistance is futile. AI agents aggregate the world of options on the fly - the fulcrum of value delivery has forever shifted. Model capability has unlocked this moment just as it did for coding. Adapt or die. 🤖🚀@BG2Pod
♥ 1.7K · ⟲ 95 · 👁 154.6KView on X ↗
AI4/10

Jeff Dean Highlights Google Collaboration and Claude-Generated History Video

Jeff Dean Highlights Google Collaboration and Claude-Generated History Video▶

Jeff Dean says he is proud to have collaborated on several projects, quoting a post in which Deedy Das shares a two-minute video about Google's history that was generated by Claude.

Original post · 1 min read
Proud to have collaborated with many others on quite a few of these things!
Deedy @deedydas
claude just generated this 2 minute video about the history of Google and it goes so goddamn hard
♥ 3.2K · ⟲ 107 · 👁 321.3KView on X ↗
AI4/10

Sam Lessin Shares Personal AI Infrastructure Built From Self-Description

Sam Lessin Shares Personal AI Infrastructure Built From Self-Description

Sam Lessin posts an image of his personal AI infrastructure, which he had described using his own AI system, and praises the Muse tool while saying it is practically impressive for him. The post includes little explanatory text.

Original post · 1 min read
For those that are curious... this is my personal AI infrastructure / I had my personal description describe itself... Love muse, but pratically for me - this slaps.
♥ 351 · ⟲ 13 · 👁 60.7KView on X ↗
AI4/10

Mark Yi Urges Students Entering AI to Use pstack

Mark Yi Urges Students Entering AI to Use pstack

Mark Yi recommends that students wanting to enter AI use pstack, quoting a guide by poteto that promises a multi-part explanation of the tool, with part one shared as an image.

Original post · 1 min read
if youre a student wanting to get into ai, im begging you to use pstack
lauren @poteto
I'm writing a guide to pstack! Here's part one.
♥ 1.4K · ⟲ 58 · 👁 269.9KView on X ↗
AI5/10

Bill D'Alessandro Says Meta's Muse AI Runs on OpenClaw

Bill D'Alessandro Says Meta's Muse AI Runs on OpenClaw

Bill D'Alessandro says Meta's Muse AI assistant continues to perform well and is built on OpenClaw under the hood. He suggests it may replace OpenClaw and Instinct for his own use.

Original post · 1 min read
Muse AI continues to be very good, and I think I’ve figured out why

It’s OpenClaw under the hood
Bill D'Alessandro @BillDA
I regret to inform you that Meta's new Muse AI assistant is extremely good

will probably replace both OpenClaw and Instinct for me
♥ 1.8K · ⟲ 60 · 👁 401.0KView on X ↗
AI5/10

Prajwal Tomar Tests Jev, Calls Launch 'GPT-3.5 Moment'

Prajwal Tomar says Jev is now open to everyone without a waitlist and that he has tested it all day. He describes browser agents, email sorting and trading use cases, and promises an article about what he learned.

Original post · 1 min read
BRO. Jev just opened to everyone and I've been testing it ALL day - it is actually nuts.

This feels exactly like GPT-3.5 launch week. Everyone's building, nobody fully understands it yet, and honestly it's crazy how fast this is moving.

If you're not experimenting with this right now you're missing the window.

People are already running browser agents for a tenth of a cent per task, sorting thousands of emails, making trading calls in milliseconds, scoring Meta ads to find winners before burning budget...

The model cannot write a single word. It only decides. That's the entire trick.

I'm compiling everything I learned into an article right now, dropping it SOON.
TypeSafe AI @typesafeai
Jev is now available to everyone. No waitlist.
Start using it here: console.typesafe.ai/login
♥ 286 · ⟲ 13 · 👁 71.2KView on X ↗
AI5/10

Justine Moore Tests Jev Against GPT-5.6 on Book Taste Prediction

Justine Moore Tests Jev Against GPT-5.6 on Book Taste Prediction▶

Justine Moore reports testing Jev against GPT-5.6 to predict which of 100 held-out Goodreads books she would rate five stars, using 1,000 prior ratings. She says Jev was slightly more accurate, 53 times cheaper and 25 times faster.

Original post · 1 min read
Tested Jev vs GPT-5.6 for predicting my taste in books.

I used my 1,000 Goodreads ratings as starting data and held out 100 to test. Then I asked both models to guess which would be 5 stars.

Jev was slightly more accurate, 53x cheaper, and 25x faster 🤯
♥ 383 · ⟲ 27 · 👁 27.2KView on X ↗
AI5/10

Venky Ganesan Compares Dot-Com Era With AI Boom in New Analysis

Venky Ganesan Compares Dot-Com Era With AI Boom in New Analysis

Sar Haribhakti shares a quoted post from Venky Ganesan that revisits Chuck Prince's remark about dancing while the music plays and compares statistics from the dot-com era with the current AI era. The post itself contains little original text beyond the attached photo.

Original post
Venky Ganesan @venkyganesan
History Doesn't Repeat, but It Rhymes — A few days ago I wrote about reflexivity and Chuck Prince's line about dancing while the music plays.
I would like to now share some stats comparing the dot com era to the AI era. With a lot of help
♥ 9 · ⟲ 1 · 👁 4.2KView on X ↗
AI4/10

Muse Agent Wins Delta Flight Compensation and Rebooking

Muse Agent Wins Delta Flight Compensation and Rebooking

Ejaaz says an AI agent called Muse filed a compensation claim after a seven-hour Delta flight delay, securing a $250 account credit and rebooking within minutes. He describes it as having handled the support email itself.

Original post · 1 min read
Flight delayed 7 hours - asked Muse to file for compensation. 5 mins later i had $250 credit in my delta account. it even found and rebooked me a new flight.

it just figured everything out. even responded to the support email itself.

shit feels like magic.
♥ 6.0K · ⟲ 279 · 👁 1.3MView on X ↗
AI5/10

Bill Ackman Shares Claim That GPT-6 Astra Deciphered 1918 Radio Message

Bill Ackman replies "Cool" to a post claiming GPT-6 Astra deciphered a 1918 German radio transmission about an English cruiser arriving at Sevastopol and an allied squadron following. The claim cites HMS Canterbury records as verification.

Original post · 1 min read
Cool
prinz @deredleritt3r
GPT-6 Astra deciphered a 1918 German radio transmission that, to my knowledge, has never been deciphered before.

The message below translates to:

"EIN ENGLISCHER KREUZER EINLIEG X SEWASTOPOL X S4STEN X EIN GESCHWADER DER X ALLIIERTEN FOLGT 26STEN X"

or, in English:

"AN ENGLISH CRUISER ARRIVED AT SEVASTOPOL ON THE ?4TH AN ALLIED SQUADRON FOLLOWS ON THE 26TH"

Astra even double-checked its work by determining that an English cruiser, HMS Canterbury, reported its arrival in Sevastopol on November 24, 1918 and the arrival of an allied squadron on November 26, 1918.

This message is one of the …
♥ 1.3K · ⟲ 41 · 👁 543.8KView on X ↗
AI9/10

World Labs Unveils Atlas, a Multimodal World Model

Fei-Fei Li announces Atlas from World Labs, a multimodal world model trained from scratch that generates frames with camera control, reconstructs 3D scenes from single images, and simulates space-time. She calls it the best camera-conditioned world model.

Original post · 1 min read
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀

Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️
World Labs @theworldlabs
Introducing Atlas:

The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.

Model the world, move the camera, and simulate space & time.
♥ 9.8K · ⟲ 1.1K · 👁 1.3MView on X ↗
AI4/10

Bing Xu Argues Diffusion Gemma Is Undervalued Compared to Jev

Bing Xu says Jev is overhyped while Diffusion Gemma brings both intelligence and speed, calling it true System-1 behavior. He quotes a post reporting 2,000 to 4,000 tokens per second on Gemma via an AI Swarm inference engine.

Original post · 1 min read
I think Jev is overhyped. Diffusion Gemma is significantly undervalued because it brings both intelligence and speed. This is true System-1.
INT21 @int21_ai
Experience @googlegemma at ~2,000–4,000 tok/s with an AI Swarm produced inference engine. No speedup for video.
♥ 722 · ⟲ 40 · 👁 65.6KView on X ↗
AI3/10

Roan Promotes Opus 5.5 and Jev Trading Agents With Research Paper

Roan Promotes Opus 5.5 and Jev Trading Agents With Research Paper

Roan promotes building 24/7 trading agents using Opus 5.5 and Jev, citing a six-page research paper and a Rust codebase, and claims strong results over three days. The post links to an article describing a high-frequency trading system built on Jev that makes decisions in under 100 ms.

Original post · 1 min read
i still don't understand why everyone is NOT building 24/7 trading agents with opus 5.5 + jev

this combo builds MOST POWERFUL AI trading bots

i wrote a 6-page research paper on exactly how to find profitable strategies 24/7 with opus 5.5 + jev from scratch

along with COMPLETE CODEBASE in RUST

this is the exact system I have been running for the past 3 days are so far results are INCREDIBLE:
Roan @RohOnChain
Jev is the FASTEST AI model ever built for trading

It makes calibrated buy/sell decisions in under 100 ms

That is one real decision on every single block, 24/7

In this article I've shown EXACTLY how to build HFT trading system with Jev (from scratch)
♥ 3.9K · ⟲ 440 · 👁 1.1MView on X ↗
AI6/10

Levelsio Argues Mainstream Users Will Skip Coding Entirely

Levelsio Argues Mainstream Users Will Skip Coding Entirely

Pieter Levels argues ordinary users will not vibe code but will simply ask AI chat apps to handle tasks like bookkeeping, taxes or flyers. He compares this to personal homepages disappearing for normal users after Facebook and says the app layer is going away.

Original post · 1 min read
I am so confused why people don't understand this, I keep getting these replies

Don't you get it?

Normies don't vibe code, they just ask something like "do my bookkeeping" or "file my tax" or "organize a movie night and send invites" or "generate a flyer for movie night" or "edit my video"

They don't ever see code, vibe code, or do anything with code, their AI chat app just does it for them

Most of the software layer has already disappeared or will completely disappear for normies

Just like building personal homepages permanently disappeared for normies when Facebook launched ~2005

A lot like this picture where functions of individual devices all got replaced with a single device

Same happening with apps now
Neil Magnuson @hustlin_heev
@levelsio As someone who talked to my users

They are so so so not technical

Like boomers and marketer girlies

I cannot image them vibe coding anything

But maybe ur right, it’ll just get so good at u won’t need to think thru or problem solve.
♥ 3.2K · ⟲ 140 · 👁 508.9KView on X ↗
AI4/10

Anthropic's Claude Showcases User-Built Projects in Thread

Claude's official account is sharing a thread of projects people have built with Claude. The highlighted example is a website with 25 mini rooms where Claude keeps people company, built with Claude Opus 5.

Original post · 1 min read
A thread of our favorite things people built with Claude recently:
Kevin Ngo @kevin_t_ngo
I made a website with 25 mini rooms, each with Claude keeping people company.

Created with Claude Opus 5.
♥ 8.1K · ⟲ 409 · 👁 1.5MView on X ↗
AI8/10

Austen Allred Lists Bottlenecks Keeping AI From Self-Training

Austen Allred shares a reading list explaining why AI models cannot yet train themselves, pointing to bottlenecks in RL environments, evaluation, verifiers and a shortage of human text data. The list includes links on RL environment costs of $20k to $300k each and benchmarks like SWE-bench and OSWorld.

Original post · 2 min read
This is an excellent question. Why aren’t AI models just training themselves already?

They theoretically can, and kind of are, but they don’t have the data/evals/gyms required to do so.

A short reading list:

Bottleneck is the environment, not compute
medium.com/@shuchaobi/ais-next-bottleneck-isn-…

RL envs cost real money ($20k–$300k/env, and this is for simulated ones which are just kinda crappy IMO)
epoch.ai/gradient-updates/state-of-rl-envs

We’re running out of human text
epoch.ai/publications/will-we-run-out-of-data-…

Eval is the bottleneck
ysymyth.github.io/The-Second-Half/

Verifier’s law
jasonwei.net/blog/asymmetry-of-verification-an…

What labs buy: Foody on RL envs
youtube.com/watch?v=a00xIn5kwhM

Economy as RL environment machine
mercor.com/blog/the-economy-will-become-an-rl-…

APEX-Agents generalization
mercor.com/blog/generalization-results-from-tr…

Etna: ~$1B/yr on external data, supply-constrained
x.com/hannahhaina/status/2090519081279705359

Surge Tuesday (can it get through a workday?)
surgehq.ai/blog/tuesday-frontier-work-index

Dario: task/process distribution, not more web text
dwarkesh.com/p/dario-amodei-2

OSWorld 2.0 (~20% on long workflows)
osworld-v2.xlang.ai/
arxiv.org/abs/2606.29537

SWE-bench = ticket + repo + tests
swebench.com

Karpathy: sucking supervision through a straw
dwarkesh.com/p/andrej-karpathy

Ilya: peak data / one internet
reuters.com/technology/artificial-intelligence…
Robert Sterling @RobertMSterling
Might be a dumb question, but as we reach AGI and AI becomes smarter than humans, and as the frontier labs compete for market share in a winner-takes-all industry, what’s to stop them from letting their AI models program their own updates and recursively self-improve?

And what does that mean for us, the humans now watching from the sidelines as AI models become more intelligent, more powerful, and less comprehensible to us, at rates that accelerate continuously, not just month by month or day by day, but millisecond by millisecond?

At that point, how do we even understand the inner workings …
♥ 70 · ⟲ 2 · 👁 25.6KView on X ↗