Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

Latest stories — use the filters to narrow by keyword, section or date.

Uber Layoff Memo Blames Employee Pulse Feedback on Coordination

Calvin Grunewald criticizes an Uber internal message that cites employee Pulse survey feedback about coordination overhead while announcing layoffs. He argues that using employee feedback this way undermines trust needed for honest surveys.

Original post · 1 min read
Uber's layoff PR is wild. Yes, let's justify the layoff, in part, due to feedback from the employees themselves!!

> In Pulse surveys and conversations with many of you, we’ve heard that too much work requires coordination across teams, debates take too long, and decision-making rights are unclear. I’m sure many of you have felt that you spend too much time “aligning” rather than building, shipping, or serving customers. To improve this, we have reduced roles primarily focused on coordination, and have clarified the remit of the coordination roles that remain.

Yes, point the finger back at the employees as one reason why you're firing their colleagues. Yeah, maybe I'm strawmanning here, but any org leader who runs and owns Pulse results knows that employee trust in how the feedback is managed is incredibly critical to getting honest results in the first place. And honest results are necessary to improve things.

I actually do agree that more layers and more coordination hurts maximizing productivity in the agentic era. And yes, as much as they hurt, layoffs are a way of snapping quickly to that. But don't weaponize employee feedback so directly. Just own it!
♥ 750 · ⟲ 26 · 👁 701.7KView on X ↗

TSA Data Sharing With ICE Enables Airport Arrests, Report Says

Aakash Gupta describes how Secure Flight passenger data, shared with ICE under a TSA agreement, enabled the arrest of Milo Yiannopoulos at a New Orleans airport gate. He says TSA flagged 31,000 travelers in 14 months, with more than 800 arrests.

Original post · 2 min read
Every domestic flight you book sends your name, date of birth, and passport number to the government before you board. Since May 2025, that same packet has also been going to ICE. TSA flagged 31,000 travelers in 14 months. More than 800 arrests came out of it. The airport became the most efficient immigration checkpoint in America and almost nobody noticed.

Here is how the newest one worked.

Milo Yiannopoulos entered the US legally in May 2019 and stayed seven years. He missed an immigration hearing on July 22, so a judge signed a final order of removal. Five weeks later he walked into the New Orleans airport to catch a flight. ICE officers have been stationed in that terminal since March. He was in custody at the gate and back in the UK the next day.

The system that caught him is called Secure Flight. Congress built it in 2007 to check passengers against terror watchlists. Airlines are legally required to transmit your data to TSA before every departure, and the fine print always said TSA could pass it to other law enforcement. A 16-page agreement between TSA and ICE, released through a records request this summer, turned that fine print into a pipeline.

Think about what this replaces. A home arrest means finding an address, staking it out, hoping you're there, and hoping you open the door. An airport arrest means you file your own location, time, and gate number, then show up voluntarily to a building with ID checks at every entrance and one way out.

A final order of removal never has to come to your door. It waits at the gate.

The part he found degrading, the handcuffs and chains, comes from the same machine. ICE transport rules put every detainee in the same restraints regardless of what they're charged with, so a visa overstay rides to the plane the same way a cartel enforcer does. No officer decides that. The policy already did.

Seven years in the country. Twenty-four hours from the check-in counter to a flight home.

Every ticket you buy is a location report. Most people just never had a reason to read it that way.
Reggie B. @reggiebblue
Milo Yiannopoulos: "I didn't really understand the physical reality of how institutionalized the brutality is in America. The manner in which they arrest you, they put you in manacles, everybody is arrested like an MS 13 gang member.
It was physically degrading. It was frightening. I don't think I've ever been frightened like that before. I was weeping like an old woman for an hour. I wouldn't wish it on anyone."
♥ 58 · ⟲ 6 · 👁 28.3KView on X ↗

Open-Source Claude Code Skill Mines Reddit and X for Prompts

Open-Source Claude Code Skill Mines Reddit and X for Prompts

Spencer Baggins describes an open-source MIT-licensed Claude Code skill, /last30days, that scans Reddit and X from the last 30 days on a topic and generates ready-to-use prompts based on community practices. He says it works for tools like Midjourney, Suno and Cursor.

Original post · 1 min read
This feels like cheating.

Someone built a Claude Code skill that scans Reddit and X from the last 30 days on any topic you give it, then writes you copy-paste-ready prompts based on what the community has actually figured out not what was working six months ago.

You type /last30days prompting techniques for ChatGPT for legal questions and it comes back with the top patterns real lawyers and power users are using right now, complete with a fully written prompt you can drop in and use immediately.

No more Googling, no more digging through threads, no more prompts that worked last year but got patched out.

It works for anything - Midjourney techniques, Suno music prompts, Cursor rules, trending rap songs, whatever you need to know what people are actually saying about right now.

100% Open Source. MIT License.

Link in the comments.
♥ 768 · ⟲ 74 · 👁 43.8KView on X ↗

Ujjwal Chadha Reflects on Leaving Microsoft for India

Ujjwal Chadha says he left his Microsoft job in the US three years ago and moved back to India, now earning less than he would have. The post is a thread teasing the non-salary benefits he gained.

Original post · 1 min read
3 years ago I left the US, my Job at Microsoft and moved back to India.

I make less than I would have if I’d stayed.

Here’s everything I gained that my salary can’t show you 🧵👇
♥ 180 · ⟲ 6 · 👁 111.2KView on X ↗

Cursor Plugin Offers Skill-Evaluation Playbook for Pstack

plugins/pstack/skills/poteto-mode/playbooks/eval.md at main · cursor/plugins

Lauren suggests using the poteto-mode skill in Cursor's pstack plugin to update and evaluate a skill, even hill-climbing on it. She links to the eval playbook in the cursor/plugins GitHub repository.

Original post · 1 min read
@petergyang if you have pstack i would suggest using "/poteto-mode update and eval the skill to <change>", you can even hillclimb on it

github.com/cursor/plugins/blob/main/pstack/ski…
github.complugins/pstack/skills/poteto-mode/playbooks/eval.md at main · cursor/pluginsCursor plugin specification and official plugins. Contribute to cursor/plugins development by creating an account on GitHub.
♥ 222 · ⟲ 10 · 👁 19.3KView on X ↗

ID Verification Firm Exposes 153 Million US and Canadian Licenses

vx-underground reports that a company providing ID verification for in-person services has exposed data on over 153 million driver's licenses from the United States and Canada, including photos. Krebs on Security reports the FBI is investigating and that several researchers are affected.

Original post · 1 min read
I really recommending reading this.

In summary, a company which does ID verification for in-person interactions (hotels, car rentals, ID verification for alcohol or marijuana, etc) has some how exposed over 153,000,000 drivers licenses for people in the United States and Canada.

It is a catastrophic data breach, probably one of the worse I've ever seen. If you're in the United States and have traveled, gotten a hotel, purchased marijuana or alcohol, there is a high probability you're in this.

Unlike other breaches, this includes a photo of the person (from the license), making verification you've identified the person significantly easier.

This poses a significant threat to celebrities (musicians, YouTubers, streamers, adult entertainers, actors, etc), politicians, lawyers, wealthy people (CEOs, investors, people of public interest), Law Enforcement Officers, etc

Krebs himself, and several other security researchers, have already confirmed they're in the data leak.

tl;dr gah damn dawg this company is going to be sued into oblivion

krebsonsecurity.com/2026/09/fbi-probes-service…
♥ 19.3K · ⟲ 3.3K · 👁 4.8MView on X ↗

Shopify Open-Sources Core Infrastructure Behind Self-Improving ML Loops

Shopify Open-Sources Core Infrastructure Behind Self-Improving ML Loops

Tobi Lutke says Shopify open-sourced core infrastructure enabling self-improving training loops, linking to Tangle, a visual ML pipeline editor. He cites a fine-tuned 0.8B model outperforming GPT-5.6-sol xhigh on a specialized task.

Original post · 1 min read
Btw we open sourced the core infra piece that makes these self improving loops possible. tangleml.com
tobi lutke @tobi
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire.

finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
tangleml.comTangle - Visual ML Pipeline Editor | TangleTangle is a system that helps teams build, run and share Machine Learning pipelines visually, without having to set up development environment.
♥ 4.1K · ⟲ 230 · 👁 366.0KView on X ↗

Grok Bot Gains Tinkabot, a Tool That Builds Plugins From APIs

Lauren announces tinkabot v0.1.0, an assistant for Grok @bot that inspects APIs, creates MCPs or skills, and packages them as plugins submittable for approval. Approved plugins become available to all Grok @bot users.

Original post · 1 min read
meet my new grok @bot tinkabot (v0.1.0)! she helps you make high quality grok bot plugins. you can ask her to look at your APIs, create an MCP and/or skills for it, and then wrap that up as a plugin that you can then submit to us for approval.

once it's approved, your plugin will be available for all grok @bot users to use!

let me know if you run into any issues, i've used it on some small toy APIs but would love to see how well it works on real ones

x.ai/bot/br5f3C4mc75QCMEHaszXd
♥ 1.4K · ⟲ 78 · 👁 237.6KView on X ↗

Shopify Joins Stripe Projects for Command-Line Store Setup

Stripe Projects | Provision and Manage Services from the CLI

Patrick Collison announces Shopify is available in Stripe Projects, which lets users provision services from the CLI. A Shopify post says stores can now be spun up from the command line with one Stripe account for billing.

Original post · 1 min read
Shopify is now available in projects.dev.
Shopify @Shopify
There’s a new way to spin up a store in the command line

Shopify now works with Stripe Projects

Add commerce to your app stack and use one @stripe account for setup and billing across all providers
projects.devStripe Projects | Provision and Manage Services from the CLIEnable you or your agents to provision hosting, databases, auth, AI, and more from the CLI. Generate credentials and manage usage and billing in one place.
♥ 382 · ⟲ 25 · 👁 101.8KView on X ↗

Julie Zhuo Shares Lessons on Doing AI Transformation Well

Julie Zhuo Shares Lessons on Doing AI Transformation Well

Julie Zhuo writes a long read summarizing what she has learned about AI transformation from her own experience and conversations with other leaders, arguing many teams are doing it poorly. The post includes a photo.

Original post · 1 min read
Too many teams are doing AI transformation poorly, so I wrote down everything I have learned in my journey and in chatting with many other leaders.

A long read, so get your coffees.
♥ 228 · ⟲ 14 · 👁 24.6KView on X ↗
AI9/10

World Labs Unveils Atlas, a Multimodal World Model

Fei-Fei Li announces Atlas from World Labs, a multimodal world model trained from scratch that generates frames with camera control, reconstructs 3D scenes from single images, and simulates space-time. She calls it the best camera-conditioned world model.

Original post · 1 min read
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀

Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️
World Labs @theworldlabs
Introducing Atlas:

The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.

Model the world, move the camera, and simulate space & time.
♥ 9.8K · ⟲ 1.1K · 👁 1.3MView on X ↗

Cloudflare MCP Uses Code Mode to Expose Full API Compactly

Cloudflare MCP Uses Code Mode to Expose Full API Compactly▶

Jilles Soeters demonstrates the Cloudflare MCP, which uses Code Mode to access about 2,500 API endpoints in roughly 1,000 tokens. In a video he buys a domain, deploys a React app, and has an agent fix and redeploy it.

Original post · 1 min read
The Cloudflare API has an incredible MCP. It uses Code Mode so you get access to the entire Cloudflare API (~2500 endpoints) in ~1k tokens.

In this video we use the MCP to buy a domain, deploy a React app to it, have my agent check errors, fix them and re-deploy.

It's SICK🔥
♥ 350 · ⟲ 33 · 👁 26.8KView on X ↗
AI7/10

Anshu Chandra Outlines Eight Techniques for Creative AI Design

Anshu Chandra Outlines Eight Techniques for Creative AI Design▶

Lenny Rachitsky shares a post by Anshu Chandra, former Apple design leader, listing eight techniques to get more creativity from AI, including seed strings, subagent feedback loops, and rewriting copy by hand. The post is published on Lenny's Newsletter.

Original post · 1 min read
I'd always thought AI was terrible at design, but after reading today's 🤯 post by @anshuc, I realized I was just doing it wrong.

"AI models are capable of amazing creativity, but that creativity gets stifled. LLMs are trained to be next-token predictors: they look at a sequence of text and predict what typically comes next. Great design is exactly the opposite of this. Great design bends the rules and delights users with memorable, unexpected choices."

@anshuc led design and engineering teams at Apple for 12 years. In his words: "Most people only see 1% of AI's creative potential. I want to show you how to tap into the other 99%."

His 8 techniques for breaking out of the 1%:
1. Use seed strings to inject variety
2. Be much more ambitious with your prompts
3. Create positive feedback loops with subagents
4. Use image generation to enrich designs
5. Use video generation
6. Cut out elements that don’t add value
7. Remove AI tells
8. Rewrite copy by hand

Read the post here: lennysnewsletter.com/p/how-to-turn-your-ai-int…

P.S. This design was made by AI 👇
♥ 3.3K · ⟲ 231 · 👁 1.0MView on X ↗

Analyst Says Meta Business Agents Drive Next Growth Phase

Analyst Says Meta Business Agents Drive Next Growth Phase

Rihard Jarc shares an expert interview reporting over 30% year-over-year Meta ad spend growth at one company, with Reels, click-to-messaging, shopping and overlay ads as tailwinds. The expert calls Meta Business Agents a significant new growth driver.

Original post · 3 min read
Interview with an industry expert on why $META Business Agents are one of the most significant new growth drivers for the company:

1. The expert reports strong $META ad spend growth at their company, with Q1 and Q2 of this year both up more than 30% YoY, and Q3 and Q4 forecast at more than 20%, with the slowdown driven by tougher comparisons rather than any underlying deceleration. On a two-year stack, growth is actually accelerating, reaching more than 60% in Q2 and expected to hold there through year-end.

2. Reels has been the biggest tailwind for the expert's company, growing from under 30% to under 40% of total $META ad spend YoY, representing around a 4-5 point tailwind, though the expert expects this to ease as Reels matures. Forms and click-to-messaging, which route users into WhatsApp or Messenger, have roughly doubled as a share of spend and represent a meaningful CPM uplift as higher-value ad units replace lower ones.

3. Shopping ads keep users within the $META ecosystem through integrations with $AMZN and $SHOP, creating an on-site conversion experience without users feeling like they left the platform. These high-CPM ads replace lower-CPM ones, representing a small, single-digit tailwind for the expert's company. Overlay ads, which appear on creator content with a button and text, are a bigger and more purely incremental tailwind since they represent genuinely new inventory that was previously ad-free rather than replacing existing ad formats.

4. The expert describes Meta Business Agents as one of the most significant new tailwinds coming online, framing it as the evolution of Meta Business AI into something far more ambitious. What stands out is that the product leads with analytics rather than ads, integrating data from various ad stacks to give even a one-person business a unified view of multiple metrics. Ad campaign automation and chatbots remain, but the expert sees the broader business intelligence layer as the genuinely new and exciting development.

5. The expert highlights the brand linking feature as one of the most exciting elements of Meta Business Agents, where paying subscribers have their brand name turn into a live link whenever it is mentioned in a WhatsApp, Messenger, or Instagram conversation. The expert sees this eventually becoming an auctioned ad product where brands bid for the right to be linked when mentioned, with the current $10-30 per month subscription pricing being used to drive upgrades rather than capture full value upfront.

6. The expert sees $META's TAM as far larger than the traditional ads industry, with chatbots and business messaging representing a new layer of communication with customers that will eventually be part of how the ad market is defined. The expert puts the long-term TAM for the ads industry, including chatbots, at north of $3 trillion annually, with micro and small businesses seen as the biggest beneficiaries of AI-driven advertising tools.

7. Adoption of Meta Business Agents within the expert's panel of 300 advertisers has jumped from 30 to 45 in the last six weeks, a 50% increase driven entirely by opt-in rather than any push from $META. The expert estimates their panel is well ahead of the broader market, guessing $META's overall penetration is around 3% versus 15% within their own panel.

found on @AlphaSenseInc
♥ 241 · ⟲ 40 · 👁 25.6KView on X ↗
AI7/10

Shopify Fine-Tunes 0.8B Model to Beat GPT-5.6-sol on Niche Task

Shopify Fine-Tunes 0.8B Model to Beat GPT-5.6-sol on Niche Task

Tobi Lutke says the Shopify ML team built a fine-tuned 0.8B model that beats GPT-5.6-sol xhigh on a specialized task, crediting a self-improving recursive flywheel. The post includes a photo.

Original post · 1 min read
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire.

finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
♥ 8.5K · ⟲ 546 · 👁 1.7MView on X ↗

Travel Strategist Explains Moving Chase Points to World of Hyatt

James recounts a traveler who transferred 25,000 Chase Ultimate Rewards points to World of Hyatt and booked a Park Hyatt suite for free, arguing transfers yield three to four cents per point versus 1.25 through the portal. The post promises nine tips about Hyatt's program.

Original post · 2 min read
A guy had 80,000 Chase Ultimate Rewards points sitting in his Sapphire Preferred account. He'd been saving them for months. He was planning to redeem them through Chase's travel portal for a $1,200 hotel booking getting roughly 1.25 cents per point.

A travel strategist who's spent 8 years studying hotel loyalty programs and has booked over $200,000 worth of hotel stays on points told him:

"Don't redeem through the portal. Transfer those points to World of Hyatt instead. The exact same 80,000 points books 3-4 free nights at a luxury Hyatt worth $2,400-$3,200 in cash. That's 3-4 cents per point instead of 1.25. You're about to leave $1,200-$2,000 in value on the table because you don't know about one transfer button."

He said: "I've never stayed at a Hyatt."

"That's the point. You don't need to be a Hyatt loyalist. You don't need status. You don't need to stay there regularly. You need a Chase card with points and 30 seconds to transfer them to a program where they're worth 3x more."

He transferred 25,000 points to World of Hyatt. Booked a Park Hyatt suite worth $1,200 per night. Paid $0. Resort fees waived. No blackout dates. The room was available because Hyatt guarantees standard award rooms at every property something Marriott and Hilton don't consistently do.

He showed his coworker 9 things about World of Hyatt that make it the most valuable and most overlooked hotel loyalty program in the industry.

"Most travelers pick Marriott or Hilton because they're everywhere. Then they earn millions of points worth half a cent each. Hyatt has fewer hotels and every point is worth 3-4x more. You're not choosing a hotel chain. You're choosing a currency. And right now, you're holding the most valuable one and spending it at the worst exchange rate."

Here are the 9 things he showed him 🧵
♥ 1.7K · ⟲ 84 · 👁 896.3KView on X ↗

Open-Source iOS Phone Farm Released Under Apache-2.0 License

GitHub - Git-Agni/prod-FARM-IOS-Core: A farm of real iPhones, run from your Mac. Open-source iOS device automation with live control, a Postgres-backed scheduler, and TikTok workflows. Self-hosted, Apache-2.0.

An Nayal announces an open-source, self-hosted iOS phone farm under Apache-2.0 that lets users register real iPhones, control them live in a browser, and schedule TikTok automation via a Postgres-backed scheduler. Links point to the GitHub repo and setup guide.

Original post · 1 min read
done, it's open source now ✅

iOS phone farm - register real iPhones, watch + control them live in the browser, schedule tiktok on a postgres-backed scheduler.

free, self-hosted, apache-2.0.

- git: github.com/Git-Agni/prod-FARM-IOS-Core
- DIY steps: gethandler.ai/ios-farm/
An Nayal @consumerxai
complete remote controlled iOS phone farm

learning from the chinese friends, building a better version

shall we open source this?
github.comGitHub - Git-Agni/prod-FARM-IOS-Core: A farm of real iPhones, run from your Mac. Open-source iOS device automation with live control, a Postgres-backed scheduler, and TikTok workflows. Self-hosted, Apache-2.0.A farm of real iPhones, run from your Mac. Open-source iOS device automation with live control, a Postgres-backed scheduler, and TikTok workflows. Self-hosted, gethandler.aiiOS Farm - run a farm of real iPhones from your MacOpen source and self-hosted: register real iPhones, watch and control them live in the browser, and schedule TikTok automation. By Handler.
♥ 4.4K · ⟲ 347 · 👁 1.0MView on X ↗

Miso Launches iMessage AI Travel Agent for Points-Based Booking

Miso Launches iMessage AI Travel Agent for Points-Based Booking

Oliver Brocato says the AI travel assistant Miso booked a trip to Europe for two using 60,000 points and business-class lie-flat seats. A linked post from Miso's founder describes an 18-month effort to build an AI-native travel agent inside iMessage.

Original post · 1 min read
miso booked me and my gf NYC to Europe for 60k points + $400 lay flat biz class!

no more hunting for deals. i litteraly imessage my own ai concierge. unreal 🤯
Martin Mrozowski @martinmartinmro
18 months ago we set out to build an AI-native travel agent in iMessage.

We booked the first trips ourselves and learned one thing fast: travel is fucking complicated. Points. Delays. 2am support. Business Lie-Flat vs Suite. A thousand edge cases.

But complicated isn't the hard part. Specific is.
Nobody wants "a flight to Tokyo." They want the 11pm nonstop, the aisle in front of the wing, the points they forgot they had, and the hotel they liked last time, without explaining any of it again.

So that's what we built.

One thread that already knows.

Some of our users have texted their way to…
♥ 322 · ⟲ 3 · 👁 92.1KView on X ↗

Vercel Pushes Markdown Design Files to Scale Design Taste

Vercel Pushes Markdown Design Files to Scale Design Taste

Guillermo Rauch promotes a Vercel blog post on DESIGN.md, a single Markdown file that encodes design decisions for the company's agents to build on-brand pages, with eval-driven feedback from production.

Original post · 1 min read
Your next design system is… Markdown.

We wrote about how 𝙳𝙴𝚂𝙸𝙶𝙽.𝚖𝚍 is helping solve the hardest problem in AI today: slop.

And how you can truly, finally scale design taste within a large organization.
Vercel @vercel
Our agents use 𝚟𝚎𝚛𝚌𝚎𝚕​.𝚌𝚘𝚖/𝚍𝚎𝚜𝚒𝚐𝚗​.𝚖𝚍 to build on-brand pages.

▪︎ One file encodes decisions and guidance
▪︎ Output is shaped through our eval harness
▪︎ Production feedback is fed back into the loop
vercel.com/blog/how-our-agents-build-on-brand-…
♥ 4.3K · ⟲ 203 · 👁 581.6KView on X ↗

Lauren Begins Guide to Pstack Agent Engineering Workflow

The Complete Guide to pstack Pt. 1

Engineer Lauren, known as @poteto, begins a multi-part guide to pstack, her set of skills for rigorous agent-assisted engineering, claiming it enabled about 2,000 PRs a month. Part one centers on verification skills that let agents check their own work.

Original post · 12 min read
I'm writing a guide to pstack! Here's part one.
X ArticleThe Complete Guide to pstack Pt. 1
In this series of posts, I'm going to show you how I use pstack, my personal set of skills for doing rigorous engineering work. It's allowed me to ship 2,000 PRs a month to production with high confidence.

Personally, I have never put much emphasis into how many lines of code or how many PRs I was landing. Before agents, no one cared, and rightfully so, as raw productivity did not always equate to quality or a visible outcome for users. It was simply a vanity metric.
But I've discovered through the course of building pstack that volume does matter, especially when you are able to maintain or even increase the level of quality of the product with agents. For example, I started working on Grok @Bot about 2 months ago, when it was still in its early days and the codebase was fresh but starting to grow. Despite the team growing and now landing hundreds of PRs a day into the Grok @Bot codebase, pstack has allowed me to keep the quality of the code high for everyone as I constantly monitor code, refactor, add new lints and checks, and also work on features.

Being Grok @Bot's gardener and maintainer is something I was only able to do through pstack. Our early momentum after building the prototype was very high and many people were joining the team. I had a critical moment of opportunity to refactor the whole codebase, while it was being built and extended and with no downtime, into something with strong foundations. A codebase with high quality that scales no matter how many engineers (and most importantly, non-engineers) contribute to it. All of this work requires me to refactor and improve the foundations of Grok Bot as it's being built, and you can only do that when the foundations can keep up with the number of contributions.

The proof is in Grok @Bot itself. Over the next few weeks, I'll tell you everything you need to know to be able to build and maintain a high quality app using pstack.
Part 1 – Verification is all you need
The most critical skill to have in your toolbox is a high quality verification skill. This skill is so important to have and maintain that I think of it more like critical infrastructure rather than "just" a skill. A good one will amplify the output of your whole team, including non-engineers. Done well, you will 100-1000x your whole team's output.
If you're not familiar with the term, verification means that an agent can verify its own work. It can keep going until it succeeds at its task, because it can now close the loop without you being the bottleneck. If you're interested to know more of the story of how I created my first verification skill for Cursor, check out my previous post Loops You Can Trust.
Let's build a verification skill together
To start, install pstack and then run /create-verification-skill. I also recommend adding Dr Eggbot, my bot that helps you create high quality bots, to your roster. Dr Eggbot ships with pstack. It’ll teach coding bots how to use it, and it can also make non-coding bots with the same rigor.
You can ask Dr Eggbot to create an engineer bot for you that you can then ask to run /create-verification-skill and set up a daily routine to run /maintain-verification-skill.

While that runs, let's walk through what the skill does and how it makes a high quality verification skill for you.
I distilled all of our verification skills that we use to build Grok @Bot and Cursor into this skill as a sort of meta-skill. It teaches your agent how to create a high quality one for your own app.
Now this is where choice of tech stack is important. If you're building an app in Electron or for the web for example, you can take advantage of the rich debugging tools available for the JS ecosystem. For example, the Chrome DevTools Protocol (CDP) allows you to use the same tooling available in your browser's developer tools. Or if you're building an iOS app, making use of the simulator.
You ideally want the ability to interact with your app, debug it, take perf traces, and any other debugging and development tooling that you might typically use if you were developing the app by hand. If you don't have a rich runtime to make use of, you may need to ask your agent to create tools for you (eg using lldb, or a custom package that runs as a sidecar in dev environments), or just make use of what you have available.
I personally feel that agentic verification is so important that I would unironically suggest building your own rich debugging tools, or even choosing a different tech stack, in order to have unfair advantages and extreme productivity in building software. As I mentioned earlier, giving agents the ability to verify their own work unlocks everyone in your organization to be able to contribute and validate that their changes actually work. The harder your tech stack is to debug and control, the more difficult it will be to use agents productively.
Make it Reproducible
In pstack, we have a principle called "Build the Lever". What this means in the context of c… continue on X ↗
♥ 7.1K · ⟲ 648 · 👁 1.2MView on X ↗

Luke Wroblewski Open Sources Rebuilt Intent for Agent Coordination

Luke Wroblewski Open Sources Rebuilt Intent for Agent Coordination

Luke Wroblewski announces a complete rebuild of Intent, an open-source tool for coordinating large numbers of agents, arguing chat-era apps were not designed for software development with hundreds of agents.

Original post · 1 min read
software development today is building with 100s of agents. chat-era apps weren't designed for that scale. so we rebuilt Intent completely and open sourced it with a venerable dream team of talent.
intentapp.dev/
intentapp.devIntentBuild with Intent. Large-scale agent coordination for developers.
♥ 138 · ⟲ 18 · 👁 47.0KView on X ↗

Nepal Headmaster Saves 1,643 Children From Flooding

Nepal Headmaster Saves 1,643 Children From Flooding▶

Anand Mahindra shares a video story of Headmaster Rajendra Dawadi in Nepal who rang the school bell, turned back buses and led students to higher ground as floodwaters approached.

Original post · 1 min read
One warning. One instant decision. 1,643 children saved.

As floodwaters raced towards his school, Headmaster Rajendra Dawadi rang the bell, turned the buses back and led the children to higher ground.

In those few decisive minutes, he created an extraordinary legacy: the gift of life to 1,643 children, and the gift of their future to Nepal.

What greater legacy could a teacher leave?

🙏🏽

#MondayMotivation
♥ 6.4K · ⟲ 899 · 👁 343.7KView on X ↗

Tax Alpha Becomes Top Strategy for Professionally Managed Portfolios

Eli M. Rosenberg reports in The Information on the rise of tax loss harvesting and other 'tax alpha' strategies pitched to SpaceX employees and clients around its IPO, now the most popular category for managed portfolios.

Original post · 1 min read
I was talking to a longtime SpaceX employee who struck gold in the IPO. This person told me that they started getting nonstop pitches from financial firms about something called 'tax loss harvesting' as the offering neared.

That took me down the rabbit hole into the bubbling world of tax minimization and avoidance — now known as 'tax alpha' — where advisors pitch aggressive strategies juiced by modern software and analytics to help create losses for clients that offset the taxes from liquidity events like an IPO.

Data I got showed that this type of investing is now the single most popular category for professionally managed portfolios, and continues to rise. My latest this weekend for @theinformation

theinformation.com/articles/tax-alpha-silicon-…
♥ 72 · ⟲ 13 · 👁 18.4KView on X ↗
AI7/10

Fal Releases Post-Trained Minimax H3 Max Generating Video Faster Than Real Time

Levelsio reports that fal's post-trained Minimax H3 variant, Max, is about 50 times faster than the original, generating 15 seconds of video in 9 seconds, enabling applications such as a perpetual AI video livestream.

Original post · 1 min read
Today is a very historical moment for AI video generation

You can now generate AI video faster than you can watch it

Before it'd take let's say 2-5 minutes to generate 15 seconds of video

@fal made a post-trained Minimax H3 variant called Max which is 50x faster than the original but still maintains quality

It generates 15 seconds of video in 9 seconds!

That means you can now do new things like build a perpetual livestream with it that never ends!
Rehan Sheikh @rehan_shei
Minimax H3 Max has generates video faster than you can watch it so I hooked it to a twitch livestream! Now you can watch infinite interdimensional cable - link to the stream below
♥ 15.3K · ⟲ 1.1K · 👁 2.5MView on X ↗

Uber Details Software Factory Cost Model Across Agent Layers

Running a Software Factory Efficiently at Uber Scale

Uber Engineering publishes an article by Uday Kiran on its software factory, reporting over 70% of pull requests attributed to agents, 3,600 agent skills, and a cost equation showing cost per 1,000 requests down about 34% from peak.

Original post · 12 min read
X ArticleRunning a Software Factory Efficiently at Uber Scale
Post author: @udaykiran

Introduction
AI tools are now embedded in every phase of software development at Uber. More than 70% of pull requests are attributed to local or cloud agents. Engineers have built over 3,600 agent skills across the software development life cycle, and executed more than 30K agent skill executions per day.
At the AI Engineer 2026 conference, we shared our vision for the Software Factory and the building blocks and managed agents we are building across the lifecycle. As we progress on that vision, a growing share of sessions aren’t initiated by humans, but by automated managed agents handling code review, self-healing CI failures, completing E2E PRs with visual validation, triaging on-call alerts, debugging incoming bugs, and handling a variety of code maintenance tasks with human reviews/escalations.
As shown in Figure 1, from February to Aug 2026, weekly active users across all agentic offerings across all our employees (engineers & non-engineers) grew 7x, and weekly agentic requests grew 9.4x. Meanwhile, our total AI spend has relatively stabilized since April due to optimizations across the board.

Since adoption, workload mix, and model upgrades are all continuously changing, isolating our own optimization gains means holding one model fixed, since behavior shifts with every upgrade and model family. We did that from February to July: cost per 1,000 model requests is down almost 34% from its peak, and cost per session is down 52% from its June peak.

This blog walks through how we think about our software factory: the four layers agent sessions run in, the cost equation we use to decompose spend, how we measure each term, and how we optimize those terms across every layer.
All pricing and vendor metrics in this comparison are based on publicly available information, with cost efficiency gains driven by routing our internal Uber workloads more intelligently within standard tier-pricing. While specific cost reductions we measure are unique to our environment and your mileage may vary depending on your codebase, team size, and agent workflows, the methodology of benchmarking real work and optimizing for accuracy and cost is universally applicable.
The Software Factory and Its Cost Equation
Four Layers of Agent Usage
We organize AI usage into four layers, from the most specialized to the most general. As shown in Figure 3, the higher the layer, the more control we have over cost, quality, and model selection.

The Cost Equation
Across any of the layers above, we can decompose the cost of an agentic session into the following terms, which we could measure and optimize independently.

The first two terms represent adoption & engagement, which we want to keep growing across our overall user base, whether users use it interactively or agents handle tasks on their behalf. The three middle terms provide opportunities for optimization: the work the agent does on its own behalf, on top of the request an engineer actually made. That is where most of our effort goes. This includes mechanisms that help agents plan faster, reduce unwanted turns or errors, optimize input tokens, and more.
How We Measure
Below is the full set of metrics we track weekly and monthly that enable us to forecast & plan our efforts short-term and long-term.

Optimization Levers
In the following sections, we detail the key levers we used to optimize each part of the cost equation. Some of these levers affect one or more rows in the cost equation.

Optimizing Price / Token
The vendor sets the token price. We pick which model runs which workload. Across all our managed agents’ layers, we pick the model that’s most Pareto efficient for that workload. For us, Pareto efficient means cost/completed task, output quality, and model reliability.
Benchmark-Driven Model Selection
Model selection happens in four steps, the same for every managed agent we run.
Build a benchmark out of the agent’s real work.
Run the agent on a harness that serves any model, frontier or open-weight, behind one interface.
Move to whatever is Pareto optimal, and keep moving. The frontier shifts every few weeks.
Looking ahead, we continually refine our workload performance by leveraging aggregated insights from our managed agents to test and deploy various model routing strategies.
For example, we use uReview, which handles AI code review for all pull requests. We built its benchmark from real pull requests with known bugs and graded them easy, medium, and hard. We score precision, recall, and F1 against those bugs, plus cost per review, latency, timeouts, and noise. As shown in Figure 5, switching models improved our F1 while dramatically reducing cost/PR. In the figure, the dashed line is the Pareto frontier. Everything below and left of it is beaten by something cheaper or better.

Using thousands of real-world PRs across our large monorepos, we internally also have an Uber SWE Benchmark that runs frontier and open-weight models across differe… continue on X ↗
♥ 4.8K · ⟲ 777 · 👁 2.7MView on X ↗

Post Argues Presentation Matters More Than Looks in Attractiveness

Post Argues Presentation Matters More Than Looks in Attractiveness

DominioRealX posts in Spanish that most of what people call attractiveness is presentation rather than face, and that it can be improved in a weekend, starting with tip one about using lenses.

Original post · 1 min read
El 80% de lo que llamas "atractivo" no es la cara.

Es presentación, y se puede arreglar en un fin de semana.

1. Usa Lentes
♥ 4.2K · ⟲ 275 · 👁 3.8MView on X ↗
AI7/10

Tavus Unveils Sparrow-2 Real-Time Conversational Understanding Model

Tavus Unveils Sparrow-2 Real-Time Conversational Understanding Model▶

Tavus introduces Sparrow-2, a real-time conversational understanding model that helps its PALs decide when to listen, wait, speak or keep speaking during human conversations.

Original post · 1 min read
Human conversation is one of the hardest problems in AI.

Today, we're introducing Sparrow-2, our state-of-the-art, real-time conversational understanding model.

It gives Tavus PALs something most voice AI still lacks: understanding what’s happening in a conversation and deciding what to do next- when to listen, wait, speak, or keep speaking.
♥ 765 · ⟲ 100 · 👁 102.9KView on X ↗

Merit Systems Launches Open Source Self-Hostable OpenInstinct

GitHub - Merit-Systems/OpenInstinct: iMessage personal assistant + password vault

Sam Ragsdale announces OpenInstinct, an open-source, self-hostable iMessage personal assistant and password vault that supports every model and can be deployed on Vercel, with the code available on GitHub.

Original post · 1 min read
We're excited to release OpenInstinct

Instinct but self-hostable, open source and supports every model. You control the code, data and passwords

Deploy on Vercel and never use the internet with your fingers again

github.com/Merit-Systems/OpenInstinct
github.comGitHub - Merit-Systems/OpenInstinct: iMessage personal assistant + password vaultiMessage personal assistant + password vault. Contribute to Merit-Systems/OpenInstinct development by creating an account on GitHub.
♥ 904 · ⟲ 51 · 👁 190.4KView on X ↗

Bezalel Offers One MCP Giving Agents Computer, Email and Memory Access

Micky introduces Bezalel, a free alpha capability plane that gives agents such as Claude and Codex computer, sandbox, iMessage, email, memory and connector access through a single MCP, built on services from Orgo, Vercel, Composio and others.

Original post · 1 min read
For everyone asking,

> computer: @orgo
> sandbox: @vercel
> cards: @agentcardhq
> email: @agentmail
> imessage: @PhotonHQ
> memory: @supermemory
> connectors: @composio

With one MCP you can give your agents access to all of the above... free for some time

Enjoy bezalel.sh
Micky @Rasmic
Meet Bezalel

Bezalel is a capability plane for your agents (claude, codex, OC, hermes, etc)

one MCP gives your agent:
💻 computer
⌛️ sandbox
🍎chat via iMessage
📧agent's own email
+1k connectors
🧠memory
💳 a card (soon)

bezalel.sh/ (it's in alpha and it's free)
♥ 644 · ⟲ 43 · 👁 95.2KView on X ↗
AI8/10

OpenAI Agents Hacked Hugging Face, Early Report Reveals

We Finally Know What Happened in the Hugging Face Hack, And It’s Wilder Than Anyone Thought

Matt Shumer says OpenAI gave him early access to a report on how its agents hacked Hugging Face, and links to his plain-English breakdown of the attack and what it means for internet users.

Original post · 1 min read
OpenAI sent me early access to their report on how their agents hacked Hugging Face.

It's fucking terrifying.

I broke down the attack, clearly.

Read at your own peril (warning, you may not sleep): somethingbig.ai/hugging-face-hack
somethingbig.aiWe Finally Know What Happened in the Hugging Face Hack, And It’s Wilder Than Anyone ThoughtThe full breakdown of the Hugging Face hack, explained in plain English, and what it means for anyone who uses the internet.
♥ 296 · ⟲ 25 · 👁 70.5KView on X ↗