Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

Latest stories — use the filters to narrow by keyword, section or date.

Brian Halligan Recounts Databricks CEO's Take on Staff Meetings

Brian Halligan Recounts Databricks CEO's Take on Staff Meetings▶

Brian Halligan shares a video in which Databricks CEO Ali Ghodsi calls a routine Monday staff meeting unnecessary, arguing that teaching material can go in a shared doc. Halligan, who built a digital twin named Hal, concedes that teams still need time together to function.

Original post · 1 min read
This answer is so Ali.

I asked @Databricks CEO @alighodsi if his Monday staff meeting was necessary, or if it's just routine...theater.

His answer: "why do you have to meet your wife? Or your kids? Just put what you want to teach them in tagoddamn Google Doc."

I built Hal, an digital twin of myself. I'm as pro-async as anyone out there.

But Ali's still right: if you want people to be a team, they have to hang out. This is humans, after all.

Link to the full episodes in the comments.

cc: @matanSF
♥ 90 · ⟲ 5 · 👁 21.7KView on X ↗

Jasper Li Open-Sources Workflow for Agent-Driven AI UGC Video Variants

Jasper Li announces an open-sourced workflow that lets an agent create AI user-generated-content videos, clone styles and generate variants at scale, positioned against Higgsfield, Seedance and MiniMax. A quoted post by Shengkun Ye describes the workflow as costing $0.03 per second.

Original post · 1 min read
You don’t need Higgsfield.
You don’t need Seedance.
You don’t need MiniMax.

We open-sourced the workflow so your agent can create AI UGC, clone any style, and generate variants at scale.

1 clone. 100 variants. 100M views.
Let your agent cook. 🔥
Shengkun Ye @shengkunye
We killed the $60 human UGC.

Introducing GPT-6 Astra for AI UGC.

Open-sourced the workflow to replicate viral videos for just $0.03/sec.

@hypitai × @MonidHQ
♥ 923 · ⟲ 68 · 👁 109.1KView on X ↗
AI8/10

fal Speeds Up Open Source MiniMax H3 Video Model Dramatically

fal Speeds Up Open Source MiniMax H3 Video Model Dramatically▶

Jennifer Li highlights fal's rebuild of MiniMax's open source H3 video model, which generates video faster than it plays back. The post describes director mode with voice prompting for camera and action control, and quotes a16z's account of GPU utilization rising from 30-40% to 70-80% without quality loss.

Original post · 1 min read
@fal's H3 Max generates video faster than you can watch it. With director mode and voice prompting, you can direct a scene as it plays - moving the camera and guiding the action just by talking to the model.

The leap isn’t just speed. It’s creative control. That’s what takes AI from impressive demos to a serious technology for Hollywood and professional filmmakers.

Inspiring convo with @gorkem and @isidentical on what possibilities are unfolding in generative media.
a16z @a16z
.@fal's Gorkem Yurtseven and Batuhan Taskaya on making an open source video model 35x faster, and what Hollywood wanted after they built it:

Last month, MiniMax released H3, an open source video model. fal rebuilt it - they cut down the steps the model takes to make a video, rewrote the code under each stage, and got the GPUs to 70-80% of their theoretical ceiling instead of their usual 30-40%. No quality loss.

Video now generates faster than you can film it. The models have gotten so cheap and fast that end users aren't even asking for improvements in either category anymore. The gap has mo…
♥ 46 · ⟲ 10 · 👁 3.8KView on X ↗

Developer Reports iOS App Ranking Near Top 11 With Daily Revenue

Developer Reports iOS App Ranking Near Top 11 With Daily Revenue

David Ch says his iOS app ranks No. 11 on the App Store, reporting about 1,300 daily downloads and $1,800 in daily revenue. He links to an article and a quoted post describing a manual study of 30,600 iOS apps earning $20k to $100k per month.

Original post · 1 min read
Our app is ranked #11 on the ios AppStore

It only took:
→ 1,300 downloads per day
→ $1,800 revenue per day

Everything you need to know about scaling iOS apps like we have

is in this article.

Quit brainrot for 3 minutes
And read this
David Ch @chhddavid
I manually studied 30,600 iOS apps making $20k-100k/mo+, the recipe is stupidly easy — Only around 1.7% of apps in the iOS App Store make $20k/mo+.
That's still more than 30,000 apps.
I've spent a stupid amount of time studying them manually. No AI summaries, no giant scraped dataset.
♥ 429 · ⟲ 23 · 👁 93.9KView on X ↗

ABC News Reports Trump Administration Downplays US War With Iran

ABC News Reports Trump Administration Downplays US War With Iran▶

ABC News reports the Trump administration continues to downplay the US war with Iran as new images appear to show damage and destruction to US military assets, shared in a video.

Original post · 1 min read
The Trump administration continues to downplay the U.S. war with Iran, as new images obtained appear to show the damage and destruction to U.S. military assets.
♥ 16.2K · ⟲ 6.1K · 👁 3.6MView on X ↗

Full Context Highlights Open Source Jev Tool for Fast Browser Agents

Full Context introduces Jev, an open source tool from Browser Use that reviews a page and picks the next action, such as clicking, typing, scrolling or retrying. The author suggests splitting work between a large model for planning and a small fast model for choosing clicks, and notes the claims are still being verified.

Original post · 1 min read
Just found Jev.

Jev looks at the page, sees the real buttons, and just picks: click this one, type here, scroll, wait, or done.

- Click this
- Scroll down
- Type this
- Buy / sell
- Retry or stop
- Send the job to another agent

That’s it.

This is useful for:
• Browser agents that need to be fast
• Auto QA
• Trading loops
• Agent routing
• Games / robots
• Anything where the AI has to choose an action over and over.

The real trick is splitting thinking from “what do I do next”. Big model does the hard thinking. And the small fast model just picks the next click.

Still looking into it, so forgive me for any false claims.

Open source here:
github.com/browser-use/jev-ultrafast
♥ 1.9K · ⟲ 160 · 👁 134.0KView on X ↗

Dex Horthy Argues Jev Fits Pipeline Design From 12-Factor Agents

Dex Horthy argues that Jev is a strong reason to revisit his 12-Factor Agents guidance, since tool calling can be split into classification and action. He recommends designing AI systems as pipelines that mix classification, structured data, deterministic code and small agent loops, citing a linked guide.

Original post · 1 min read
jev is the best excuse you could possibly have to go re-read 12 factor agents. Tool calling itself can be decomposed into classify+action,

if you learn to design ai programs as pipelines that switch breathlessly between classification, structuring data, deterministic code, AND small agent-shaped append-chat loops, then jev is a WONDERFUL building block

hlyr.dev/12fa
Dillon Mulroy @dillon_mulroy
i think jev is resonating with devs so well b/c it unlocks so many opportunities for composing ai into systems and products rather than ai _becoming_ the product/system

really does feel like it was a missing primitive
♥ 2.7K · ⟲ 137 · 👁 233.6KView on X ↗

Fahad Desmukh Shares Video of Karachi Marathi Ganesh Chaturthi Event

Fahad Desmukh Shares Video of Karachi Marathi Ganesh Chaturthi Event▶

Fahad Desmukh shares a video from a Ganesh Chaturthi event in Karachi featuring the local Marathi community, posted in response to questions he receives from Indians about Marathis in Pakistan.

Original post · 1 min read
I regularly get DMs from Indians mistaking me for a Marathi due to my name and asking me questions about the Marathi community in Pakistan.

So for them, here's a video by @theamarparkash of real Karachi Marathis at an event to mark Ganesh Chaturthi I believe 🚩♥️♥️
♥ 499 · ⟲ 83 · 👁 138.6KView on X ↗
Other1/10

Zatchfcb Praises Standout YouTube Short for Its Editing

Zatchfcb Praises Standout YouTube Short for Its Editing▶

Zatchfcb calls a YouTube Short the best they have seen and praises the editing in an attached video. The post offers no details about the creator or content.

Original post · 1 min read
Hands down the best Youtube short ever made. What an incredible edit! 👏
♥ 25.1K · ⟲ 3.2K · 👁 5.6MView on X ↗

Grok Bot Livestreams Three SpaceXAI Staff Building a Company in Three Days

Grok Bot announces a livestream in which SpaceXAI employees Matt Palmer, Lauren Tan and Roshan Sadanani attempt to build a company in three days using Grok Bot, starting with research, a plan and a product, with sessions on engineering, product and founding.

Original post · 1 min read
Three SpaceXAI employees are building a company in 3 days with Grok Bot. This is Day 1.

Matt Palmer (@mattyp), Lauren Tan (@poteto), and Roshan Sadanani (@roshan_s) start with research, a plan, and a product.

Live now, plus sessions for engineering, product, and founders.

x.com/i/broadcasts/1AxRnZbVpjaxl
♥ 11.0K · ⟲ 1.6K · 👁 4.9MView on X ↗

Anthropic Says Claude Writes 80% of Its Code, Strains CI

Anthropic Says Claude Writes 80% of Its Code, Strains CI

Addy Osmani reports Claude now writes 80% of Anthropic's code and engineers ship 8x more per quarter. The side effects include 10x more tests and a 25x rise in CI jobs over six months, which Anthropic addressed with a scaled test impact analysis.

Original post · 1 min read
At Anthropic, Claude now writes 80% of our code. Engineers ship 8x more code per quarter.

Side effect: Tests grew 10x. CI jobs up 25x in 6 months. Here's what helped us scale:

claude.com/blog/agentic-coding-is-straining-ci…
claude.comAgentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic | Claude by AnthropicOur CI job volume increased 25x over 6 months. We patched our test selection service three times before finding a sustainable solution.
♥ 5.0K · ⟲ 367 · 👁 789.3KView on X ↗

Trump Phones Jensen Huang Live at All In Summit

Trump Phones Jensen Huang Live at All In Summit▶

Ben Pouladian shares a video of the president calling Nvidia CEO Jensen Huang on stage at the All In Summit, with the president saying the U.S. will not lose the AI race. The post is largely commentary and promotes Nvidia stock.

Original post · 1 min read
Holy crap POTUS just phoned in @JensenHuang live on stage at All In Summit

We will not lose the ai race! And whatever Dario said this weekend won’t stop our progress

This made my morning!

$NVDA
♥ 16.1K · ⟲ 1.9K · 👁 6.2MView on X ↗
AI8/10

Meta's Alexandr Wang Says Agent Loops Can Outperform 100 Engineers

Meta's Alexandr Wang Says Agent Loops Can Outperform 100 Engineers▶

In a YC conversation with Garry Tan, Meta Chief AI Officer Alexandr Wang described agentic loops with evaluation metrics that let agents complete more work than 100 senior engineers. He said the system is built from simple parts like markdown files, cron jobs and metrics.

Original post · 3 min read
Former Scale AI founder and newly appointed @Meta Chief AI Officer @alexandr_wang dropped a bombshell during his YC conversation with @garrytan, and you could almost hear the tech leadership world go quiet.

He said Meta has already seen this internally:

Build the right agentic loop, give it an evaluation system and metrics that let it optimize itself, and a group of AI agents can complete more work than a team of 100 senior engineers.

And they do it “very easily.”

But the most interesting part wasn’t the 100-engineer comparison.

It was how simple the system underneath it actually is.

You’d expect some insanely complex, almost alien architecture powering a swarm like this.

Instead, Wang described it with a few almost comically basic words:

“Markdown files, cron jobs, goal, metrics, data.”

Once you strip away the hype, the implications for traditional software engineering become pretty clear:

1️⃣ It’s not that the models are magically smarter. The eval loop is doing the heavy lifting.

Traditional approach: humans write prompts, run the code, inspect the output, and hope nothing broke.

Meta’s approach: turn the business goal into something a machine can score automatically.

The agent submits its work. The system runs tests, calculates metrics, finds what’s wrong, and sends it back for another pass. Repeat until it passes.

Nobody has to babysit every step. The metric becomes the supervisor.

2️⃣ The real alpha is burning 1,000x more tokens inside the feedback loop.

A lot of people are still optimizing for the cost of a single AI call.

The frontier labs are playing a different game: spend 1,000x or even 1,000,000x more tokens in the background so agents can constantly review, rerun, challenge, and verify each other’s work until they reach a reliable business outcome.

Token cost is fixed. The payoff is a pipeline that keeps running.

3️⃣ Memory doesn’t need some fancy database.

Persistent memory can live in Markdown files.

Scheduling can be handled by the server’s built-in cron jobs, running overnight.

The simpler the scaffolding, the more robust the system can be. Less infrastructure also means fewer ways for context to fall apart.

This is a pretty brutal change in how technical organizations work.

The ceiling for a tech lead used to be partly about how many people they could manage, how many meetings they could sit through, and how many teams they could coordinate.

The leverage for the next generation of technical leaders may look very different:

Can you turn a messy business objective into a rigorous set of metrics that an AI can evaluate automatically?

If 100 people’s output can be replaced by a few cron jobs, Markdown files, and a well-designed eval loop, the era of “just take the ticket and write the code” is coming to an end.

The people who can design the evals and orchestrate the swarm aren’t just holding a new tool.

They’re effectively running a virtual company.
♥ 511 · ⟲ 58 · 👁 133.6KView on X ↗

Michigan SAT Data Shows Wide Gaps in Score Distribution

Michigan SAT Data Shows Wide Gaps in Score Distribution

Marc Porter Magee shares a breakdown of Michigan public high school students scoring 1200 or higher on the SAT, which he says is the floor for a competitive college. The post lists the share of students in each racial group, with Asian students at 46% and White students at 16%.

Original post · 1 min read
Michigan’s public high school students are required to take the SAT.

Share scoring 1200+ (the floor for a competitive college):

Asian: 46%
White: 16%
Hispanic: 7%
African American: 2%
♥ 1.2K · ⟲ 144 · 👁 217.1KView on X ↗
AI6/10

Levelsio Argues Mainstream Users Will Skip Coding Entirely

Levelsio Argues Mainstream Users Will Skip Coding Entirely

Pieter Levels argues ordinary users will not vibe code but will simply ask AI chat apps to handle tasks like bookkeeping, taxes or flyers. He compares this to personal homepages disappearing for normal users after Facebook and says the app layer is going away.

Original post · 1 min read
I am so confused why people don't understand this, I keep getting these replies

Don't you get it?

Normies don't vibe code, they just ask something like "do my bookkeeping" or "file my tax" or "organize a movie night and send invites" or "generate a flyer for movie night" or "edit my video"

They don't ever see code, vibe code, or do anything with code, their AI chat app just does it for them

Most of the software layer has already disappeared or will completely disappear for normies

Just like building personal homepages permanently disappeared for normies when Facebook launched ~2005

A lot like this picture where functions of individual devices all got replaced with a single device

Same happening with apps now
Neil Magnuson @hustlin_heev
@levelsio As someone who talked to my users

They are so so so not technical

Like boomers and marketer girlies

I cannot image them vibe coding anything

But maybe ur right, it’ll just get so good at u won’t need to think thru or problem solve.
♥ 3.2K · ⟲ 140 · 👁 508.9KView on X ↗

Andrew Chen Says AI Should Automate Drudgery, Not Human Connection

Andrew Chen Says AI Should Automate Drudgery, Not Human Connection

Andrew Chen says AI wins at automating drudgery, workflows and costs, while people want to spend time on connection, entertainment and humanity. He frames the latter as less verifiable and more driven by novelty-seeking.

Original post · 1 min read
you want to spend zero time on the left. This is where AI wins - automating drudgery/workflows/costs

we want to spend our time on the right. More of it is about connection/entertainment/humanity - less verifiable, more novelty-seeking
♥ 465 · ⟲ 40 · 👁 41.8KView on X ↗

Sergey Karayev Outlines Multiplayer Cloud Agent Workflow

Sergey Karayev describes how his team runs Claude, Codex and other agents in the cloud where any teammate can join a session. Meetings launch subagents for research and implementation, and a Chief of Staff agent tracks review queues; he published a manifesto on the approach.

Original post · 1 min read
For over a year, my team has been working with AI agents in a fundamentally different way than most.

All of our Claude/Codex/Pi/etc agents run in the cloud, and any session is joinable by anyone on the team.

Each of our meetings has an agent that launches subagents to do research, draft posts, and implement features as we discuss things.

A Chief of Staff agent lets me know what's waiting for my review, and can talk to any person or agent on the team to resolve bottlenecks.

Working in this fundamentally multiplayer way has given us a preview of the way everyone will work soon, so I wrote up a short manifesto explaining

• the current problems
• principles for a great solution
• some things that are tricky to get right

Check it out, and let me know what you think!

multiplayer-ai.com
♥ 902 · ⟲ 74 · 👁 420.7KView on X ↗

Post Shows Branded iMessages Sent to Users Who Abandon Paywalls

Post Shows Branded iMessages Sent to Users Who Abandon Paywalls▶

Gabe Roeloffs shares a video claiming app developers can send users branded iMessages when they abandon a paywall. The post is brief and offers little detail on implementation.

Original post · 1 min read
They don’t want you to know this but you can send users branded iMessages when they abandon your paywall
♥ 219 · ⟲ 5 · 👁 42.4KView on X ↗

Shopify Seeks Feedback on Agentic Commerce Developer Docs

Agentic commerce

Gil from Shopify asks developers how the company can improve its UCP and agentic commerce documentation. The linked docs describe building AI agents that authenticate with Shopify, search the catalog, build carts and checkouts, and monitor orders via the Universal Commerce Protocol.

Original post · 1 min read
How can we (Shopify) improve our UCP and related agentic commerce docs? shopify.dev/docs/agents

I have some ideas but would love to hear from all of you. 🤔
shopify.devAgentic commerceBuild AI agents that authenticate with Shopify, search the Catalog, build carts and checkouts, and monitor orders using the Universal Commerce Protocol (UCP).
♥ 28 · ⟲ 3 · 👁 2.3KView on X ↗

Indie App Developer Gains Rankings Across 40 Markets With One Localization

Indie App Developer Gains Rankings Across 40 Markets With One Localization

Blaida reports that after adding a single new localization and an Apple approval, his app now ranks across nearly 40 App Store markets, with top 10 positions in several. He argues app store optimization through localization is an underused lever for indie developers.

Original post · 1 min read
Apple approved my update. I'd added one new localization.

The result: my app now ranks across nearly 40 App Store markets, France, Germany, Italy, Mexico, Argentina, and more. Top 10 in several of them.

One localization. Dozens of countries.

This is the ASO lever indie devs ignore because it feels like "extra work." It's actually the highest-leverage move on the store.
Blaida @kedytcom
Update on the $0 → $1K MRR challenge 👇

We found new keywords + markets from our ASA data (Max Conversion campaigns surface the terms that actually convert, not just the ones we guessed).

So we rewrote the app's metadata around them — and built new features to match, so the keywords are real, not stuffed.

Then Apple rejected us. Twice.

Now we're back in review, waiting.

If it passes, we find out whether those new keywords + markets actually move the needle. If it doesn't, you'll see exactly why.

Either way I'm posting it. This is the part nobody shows.

Day 3 → still climbing. 🚀
♥ 62 · ⟲ 0 · 👁 8.0KView on X ↗

Isaiah Granet Argues Founders Often Quit Too Early

Isaiah Granet argues ambitious people often leave when things get hard, while those who stay through difficulty tend to end up ahead. Using a grocery line metaphor, he discusses when to stay and when to leave, drawing on his experience running a company that raised over $100 million.

Original post · 8 min read
A few thoughts on switching lines.
X ArticleLine Switching
Ambitious people often waste their lives by leaving things too early.
Everyone (i mean everyone) has switched lines at the grocery store and regretted it. You're in one that isn't moving, the one next to you is, so you cross over. Then yours starts, the new one stalls.
I run a company that has raised over a hundred million dollars. I switched lines to start this company. I switched lines, when we pivoted. I am not going to argue that people should never leave.
But I have watched enough of these decisions by now, across many companies and domains. The people who left almost always left at the moment things got hard, and very few who ended up where they had been hoping to go.
The people who stay through the ugly often seem to end up ahead, and they are the ones who now look lucky.
I know people who struggled for years, until it hit. I've also seen overnight pivots to success. There's no one recipe. But below is some sincere thoughts on when to stay, and when to leave. And why switching lines usually ends up failing.
I. Your Line
The line you're standing in is the only one you know from the inside. You know exactly what's wrong with it. You've been watching the slow cashier for six minutes and you have opinions about her. The other line you see from across the store, and from across the store almost anything looks fine, because you can't see the coupons or the price check or the man reaching for his checkbook. You're not learning anything about that line. You're seeing it from an angle that hides its problems.
Nobody switches lines because they've studied both. They switch because their own line's problems are visible. A la.... the grass is always greener... Then they get to the new line and its problems become visible, and now there's another line, one aisle over, that looks fine.
II. It Feels Good
Switching feels like taking control. Staying feels like giving up. So when someone leaves a job at month eight, or a startup changes markets in year two, or a founder rewrites the strategy the week before the old one would have started paying, it registers as bold. People congratulate them. Nobody congratulates you for staying in a line.
But most of these moves are the same move. The thing got hard, another thing looked easier, and the person went toward the thing that looked easier, dressed up as ambition. It's hard to tell the difference from the outside, and sometimes from the inside, because the feeling is identical. If you're leaving right when it got hard, you're probably switching lines.
III. The AI Era
This was always a problem, but it has gotten a lot worse, and I think it's worst for people just starting out.
A new grad today looks at a job market where the fastest money is in AI, where twenty-three-year-olds raise rounds off a demo, where every week someone with fewer years of experience than you announces something enormous. Every line looks shorter than yours. And these lines are visible in a way they never were, because the whole thing plays out on a feed. You see the announcement. You do not see the four years of nothing that came before it, or the two years of nothing that will come after it for most of them.
So people move. They leave the solid job at a company that's learning slowly for the startup that's moving fast. They leave the startup at month ten for the one that just raised. They leave engineering for founding, founding for investing, investing for whatever gets written about next. Each move makes sense on its own. Together they add up to someone who is twenty-nine and has been new somewhere five times and deeply good at nothing.
The AI era rewards depth more than any period I've seen, because the tools make the shallow parts of most jobs free. What's left is the part you only learn by staying. And the era tempts people out of staying more than any period I've seen. That's the trap.
IV. Leave On the Way Up
If you take one thing from this, take this. The best time to leave a company is on the way up, not the way down.
Most people do the opposite. They stay while things are good, because it's comfortable, and leave when things get hard, because it's not. Which means they leave right before the hard part would have taught them something, and right before the company would have needed them most, which is when the people who stayed get the equity and the title and the story.
If you're leaving because it got hard, wait. Hard is where the value is. Hard is the part everyone else quits, and the people who don't quit are the ones with the résumé everyone wants five years later.
If you're leaving because you've won, because you shipped the thing and learned what there was to learn and the next mountain is somewhere else, go. That's the good kind of leaving. You'll leave with people who'll hire you again and a reputation for finishing. You'll also leave with a clear head, which you don't have when you're fleeing.
V. The People Who Stayed
The founders whose stories we tell mostly didn't switch lin… continue on X ↗
♥ 3.1K · ⟲ 529 · 👁 2.2MView on X ↗
AI8/10

Austen Allred Lists Bottlenecks Keeping AI From Self-Training

Austen Allred shares a reading list explaining why AI models cannot yet train themselves, pointing to bottlenecks in RL environments, evaluation, verifiers and a shortage of human text data. The list includes links on RL environment costs of $20k to $300k each and benchmarks like SWE-bench and OSWorld.

Original post · 2 min read
This is an excellent question. Why aren’t AI models just training themselves already?

They theoretically can, and kind of are, but they don’t have the data/evals/gyms required to do so.

A short reading list:

Bottleneck is the environment, not compute
medium.com/@shuchaobi/ais-next-bottleneck-isn-…

RL envs cost real money ($20k–$300k/env, and this is for simulated ones which are just kinda crappy IMO)
epoch.ai/gradient-updates/state-of-rl-envs

We’re running out of human text
epoch.ai/publications/will-we-run-out-of-data-…

Eval is the bottleneck
ysymyth.github.io/The-Second-Half/

Verifier’s law
jasonwei.net/blog/asymmetry-of-verification-an…

What labs buy: Foody on RL envs
youtube.com/watch?v=a00xIn5kwhM

Economy as RL environment machine
mercor.com/blog/the-economy-will-become-an-rl-…

APEX-Agents generalization
mercor.com/blog/generalization-results-from-tr…

Etna: ~$1B/yr on external data, supply-constrained
x.com/hannahhaina/status/2090519081279705359

Surge Tuesday (can it get through a workday?)
surgehq.ai/blog/tuesday-frontier-work-index

Dario: task/process distribution, not more web text
dwarkesh.com/p/dario-amodei-2

OSWorld 2.0 (~20% on long workflows)
osworld-v2.xlang.ai/
arxiv.org/abs/2606.29537

SWE-bench = ticket + repo + tests
swebench.com

Karpathy: sucking supervision through a straw
dwarkesh.com/p/andrej-karpathy

Ilya: peak data / one internet
reuters.com/technology/artificial-intelligence…
Robert Sterling @RobertMSterling
Might be a dumb question, but as we reach AGI and AI becomes smarter than humans, and as the frontier labs compete for market share in a winner-takes-all industry, what’s to stop them from letting their AI models program their own updates and recursively self-improve?

And what does that mean for us, the humans now watching from the sidelines as AI models become more intelligent, more powerful, and less comprehensible to us, at rates that accelerate continuously, not just month by month or day by day, but millisecond by millisecond?

At that point, how do we even understand the inner workings …
♥ 70 · ⟲ 2 · 👁 25.6KView on X ↗
AI6/10

Teresa Torres Publishes Guide to AI Evals for Product Teams

Teresa Torres argues that AI evals, methods for measuring whether an AI product performs well, should be a discovery habit for product teams. She links to a new practical guide she wrote for non-engineers.

Original post · 1 min read
AI evals have been the "it" skill for product teams for over a year. I've even called evals a new discovery habit.

But I still meet product teams who only have a vague idea of what evals are. And it's not their fault. Most of the writing on this topic is intended for engineers or just isn't specific enough.

I recently created an in-depth eval guide to explain what evals are and why product teams can and should create them. I did my best to make it practical, hands-on, and easy to follow.

AI evals (short for evaluations) are methods for measuring whether an AI product or workflow is performing well. Evals give teams confidence that their AI applications are doing what they expect them to do. They help teams maintain quality and catch issues before they reach users.

Similar to other discovery habits like interviewing and assumption testing, evals can act as a feedback loop to ensure we are on the right track.

If you want to learn more about this new discovery habit, explore my new guide: producttalk.org/ai-evals/
♥ 292 · ⟲ 24 · 👁 63.5KView on X ↗
Other2/10

Danielle Morrill Reports Nine-Pound Weight Loss on Wegovy Pill

Danielle Morrill Reports Nine-Pound Weight Loss on Wegovy Pill

Danielle Morrill shares that after one month of taking 1.5mg of Wegovy in pill form daily alongside a plant-protein-forward diet, she has lost nine pounds. The post includes photos of meals.

Original post · 1 min read
Tomorrow I am 1 month in on taking 1.5mg of Wegovy in pill form daily and eating my normal plant-protein forward diet. I’ve lost 9 lbs and enjoyed many beautiful meals at home and out
♥ 72 · ⟲ 0 · 👁 10.1KView on X ↗

Hiten Shah Spotlights Builders Behind Featured Grok Bots

Hiten Shah tags a list of builders whose Grok Bots he featured, following his review of 407 public Grok Bots and the specific jobs people are assigning them. The post promotes the bot creators and his own experiments.

Original post · 1 min read
These are the builders behind the Grok Bots I featured today:

@matt_silberman @clairevo @mattyp @dannymacias @Andrew51786 @mvanhorn @viticci @scheemunai @mamuso @RichSilver @TylerNishida @SuddenlyJon @Boilerfan1234 @humanmeteorite @MSaintjour
Hiten Shah @hnshah
What Are People Turning Into Grok Bots? — I looked through 407 public Grok Bots. The jobs people are giving them are getting surprisingly specific.
I’ve been giving Grok Bot a new job every day this week. After building a few, I wanted to see
♥ 69 · ⟲ 5 · 👁 15.6KView on X ↗

Shopify and Anthropic Launch Claude Commerce Agents

Shopify and Anthropic Launch Claude Commerce Agents▶

Shopify CEO Harley Finkelstein announces that Claude Commerce Agents launched and that Shopify's reference implementation is live, giving merchants paths into agentic commerce. He links to Shopify's open-source examples on GitHub.

Original post · 1 min read
We will always do whatever we can to give our merchants a better chance at success.

I keep saying it and we keep proving it. Case in point: Claude Commerce Agents launched today, and our reference implementation is already live.

@Shopify gives every merchant a path into agentic commerce. Out of the box tools for most merchants, or flexible building blocks for those who want to go deeper and build their own agents.

🔗 github.com/Shopify/claude-for-commerce-examples
ClaudeDevs @ClaudeDevs
We're open-sourcing Claude Commerce Agents.

This is a blueprint for building shopping and merchant agents, with reference implementations across retail, travel, telecom, and entertainment.
♥ 464 · ⟲ 32 · 👁 84.8KView on X ↗

Anthropic Open-Sources Claude Commerce Agents Blueprint

Anthropic Open-Sources Claude Commerce Agents Blueprint▶

Anthropic's developer account says it is open-sourcing Claude Commerce Agents, a blueprint for building shopping and merchant agents. Reference implementations cover retail, travel, telecom and entertainment, shown in an accompanying video.

Original post · 1 min read
We're open-sourcing Claude Commerce Agents.

This is a blueprint for building shopping and merchant agents, with reference implementations across retail, travel, telecom, and entertainment.
♥ 11.8K · ⟲ 970 · 👁 2.2MView on X ↗

Juan Explains How Startups Engineer Overnight Ubiquity on X

Why Some Startups Become "Everywhere" Overnight

Juan's X article argues that startups seeming to be everywhere overnight results from a layered launch system, not one viral post. It outlines five layers, including a repeatable claim, a product moment, trusted voices and timing, and says he drove over six million impressions recently.

Original post · 8 min read
X ArticleWhy Some Startups Become "Everywhere" Overnight
You open X in the morning and see a startup you've never heard of. An hour later, a founder you follow is posting about it. Then a friend sends it to you. Someone drops it into your team's Slack.
By lunch, it feels like everyone in tech has been talking about this company for weeks. They haven't. That feeling, that this startup is suddenly everywhere, is engineered.
I've built launches like this for years.
In the last month alone, I've driven more than 6 million impressions and tens of thousands of reposts for tech startups.
Almost none of that came from a single viral post. It came from a system most founders never even notice.
This article breaks down:
Why some launches feel like they're everywhere, all at once.
The five layers teams actually build behind the scenes.
How you can run the same system for your own launch.
(Bookmark this. You'll need it on launch day.)
One Viral Post Won't Make You "Everywhere"
Most founders see a company take over their timeline and assume one post went viral and the rest just followed.
That's almost never how it works.

A post with two million views is still something you see once and forget. But when the same startup shows up from a founder, a creator, and someone you trust, all in a few hours, you start to care.
The first time, you scroll past it.
The second time, you recognize the name.
The third time, you wonder why everyone is talking about it.
The fourth time, you click. Every touchpoint makes the next one hit harder.
Most campaigns chase reach, hoping as many people as possible see the product once. The launches that feel everywhere do the opposite. They optimize for overlap. The right people see it several times, from sources they already trust.
Ten mentions across a month feel like marketing.
Ten mentions before lunch feel like an event.
Once your launch feels like an event, people start spreading it for you.
The Five Layers Behind an "Everywhere" Launch
Five layers need to work together:
A claim people can repeat.
A product moment people want to show.
Trusted voices entering the conversation in waves.
Timing that compresses those voices into one event.
An audience that eventually takes over.
Miss one layer and the launch loses momentum.
Strong message, no distribution? Nobody sees it. Distribution, no product moment? Empty reach. Even if you have the right message, product, and creators, if they show up too far apart, nobody feels like your launch is everywhere. Here's how each layer works.
Layer 1: A Claim People Can Repeat
Most founders launch with a product description:
An AI-powered platform for collaborative application development.
Sure, it's accurate. Nobody repeats it. Ideas only spread when they're simple enough to pass over dinner:
Build an app by describing it.
Your launch needs one portable claim that tells people what changed and why it matters. The strongest claims combine three things:
Familiar problem + surprising claim + visible proof
Let every creator take a different angle, but make sure they all hammer the same core idea. If people can't explain why your launch matters in one sentence, it won't spread.
Layer 2: A Product Moment People Want to Show
A strong claim gets attention. A strong product moment drives shares. People don't repost feature lists. They repost the moment a product does something that looks new, hard, or impossible.
I call this the magic moment.
Most weak launch videos hide it behind a logo animation, a founder introduction, or a tour of the dashboard. By the time the product does something interesting, the audience has already scrolled away.
The strongest launch assets do the opposite:
Show the result within the first five seconds.
Make it understandable without sound.
Focus on one impressive outcome.
Work even when shared without context.
One clear result will travel further than ten explained features.
The test is simple. Show the opening five seconds to someone outside your company; if they can't explain what happened and why it matters, the magic moment isn't clear enough.
Your claim tells people what changed. Your demo makes them believe it. When those two travel together, creators don't have to explain why the product is interesting. They can simply show it.
That makes the launch much easier to spread.
Layer 3: Trusted Voices Enter the Conversation
One founder post can start a launch. It can't make the company feel "everywhere." For that, the story needs to travel through people who already have the attention and trust of your target market.
Each voice has a different job:
The founder gives the launch its story.
Large creators bring reach.
Niche experts make it relevant to buyers.
Customers provide proof.
Smaller accounts create repetition.
Team members and communities carry it beyond the original post.
The biggest mistake is giving everyone the same copy. When 50 accounts use the same words, the audience sees a campaign. When they tell the same story from their own perspectives, the audience sees a conversation.
Give … continue on X ↗
♥ 1.1K · ⟲ 54 · 👁 423.6KView on X ↗