Friday, October 9, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

Latest stories — use the filters to narrow by keyword, section or date.

AI5/10

George Pu Finds Open-Source Qwen Model Can Replace Costly ElevenLabs Subscription

George Pu says he nearly paid 330 dollars a month for ElevenLabs to narrate his blog, but found the open-source Qwen 3.5 14B model runs well on his laptop for the cost of electricity. He argues many AI subscriptions are a UI over free models.

Original post · 1 min read
Almost signed up for ElevenLabs to narrate my blog. $330/month.

Then I tried running an open-source model on my own laptop. Qwen 3.5 14B.

Sounds fine. 200 posts a month. Costs me electricity.

I almost paid $4,000 a year to rent a model I can run myself.

Most AI subscriptions right now are just a nice UI on top of something free.
♥ 2.6K · ⟲ 91 · 👁 186.1KView on X ↗

Greg Isenberg Outlines Seven Lessons for Running a Startup With AI Agents

Greg Isenberg Outlines Seven Lessons for Running a Startup With AI Agents▶

Greg Isenberg shares a guide to Paperclip, an open-source project for managing teams of AI agents such as CEO, engineer and QA roles, covering persona prompts, memory via heartbeat checklists, skills, quality review and token tracking.

Original post · 3 min read
I met the guy behind Paperclip. he won't show his face, but he just built one of the FASTEST growing open-source projects in AI.

how to use Paperclip to hire AI agents to ACTUALLY run a startup with 0 employees:

1. with paperclip, you hire a team of AI agents like CEO, engineer, QA, video editor, content strategist and manage them from one dashboard.

it works with Claude Code, Codex, OpenCode, or any model on OpenRouter. you're not locked into one provider.

2. your AI agents wake up capable but with zero memory. they don't know who they are, where they are, or what they're supposed to be doing. kinda like that movie memento from back in the day

you need to leave them Polaroids like heartbeat checklists, persona prompts, written context. that's how you keep them on track.

3. when an agent makes a mistake, you don't rewrite everything. you add one rule to their persona prompt.

"always define a success condition for every task."

"always pass work to QA before closing." you're training them like you'd train a junior hire. one correction at a time.

4. skills extend what your agents can do. want a video editor who can produce animated content? install the Remotion skill. want security reviews? there's a skill for that.

5. the biggest lever for quality is encoding your own taste. AI can do everything except know your values. design sensibility, brand voice, success criteria but you have to write it down.

6. don't one-shot your startup. agentic design patterns matter. the simplest one: after the engineer builds something, QA reviews it. structure prevents compounding errors. one-shotting an entire app is fun for 30 minutes, then it falls apart.

7. Paperclip tracks every token spent and every task completed. you can use your existing subscriptions (Claude, Codex) so spend shows as $0, or hook into API credits for real dollar tracking.

8. importable companies are coming. Gary Tan's G-Stack, a full game studio, 300+ agent repos... you can "acqui-hire" a proven agent team into your Paperclip instance instead of building from scratch. the future is downloading a tested org that actually works.

9. routines let you automate recurring work. "every day at 10am, read what was merged into the main branch and write a Discord update celebrating community contributors." it runs, you review, you improve. every task is traceable.

10. maximizer mode is next. you tell the CEO "build this game" and it does whatever it takes and hires who it needs, keeps pressing until it's done. no token anxiety. just outcomes.

use @ideabrowser for startup ideas/trends to get started

thank you for @dotta for doing this podcast and breaking down exactly how people can hire ai agent teams with paperclip

you won't find an episode like this anywhere else

episode is live on @startupideaspod on your fav platforms (follow for more)

is this not the greatest time in history to be building?

im rooting for you

now go watch my frien
♥ 2.6K · ⟲ 259 · 👁 466.1KView on X ↗

Sendblue Releases CLI Giving AI Agents an iMessage Number

Sendblue Releases CLI Giving AI Agents an iMessage Number▶

Nikita introduces the Sendblue CLI, installable via npm, which provisions an iMessage number for an agent after running a setup command. A demonstration video accompanies the announcement.

Original post · 1 min read
Introducing Sendblue CLI 🟦🎉

iMessage numbers for your agents.

1️⃣ npm install -g @sendblue/cli
2️⃣ sendblue setup

Done. Your agent has an iMessage number
♥ 2.1K · ⟲ 113 · 👁 453.6KView on X ↗

Claude Code Paired With Google Stitch 2.0 Redesigns Vibe-Coded Apps

Claude Code + NEW Stitch 2.0 just changed how I design apps

Prajwal Tomar argues that generic AI-built app UIs are a workflow problem and describes using Google Stitch 2.0 with Claude Code via MCP to generate a design system and redesign a client app in under an hour.

Original post · 11 min read
X ArticleClaude Code + NEW Stitch 2.0 just changed how I design apps
Google Stitch 2.0 + Claude Code via MCP is the workflow I’ve been testing… and the results are genuinely insane.
I've been saying this for a while now. If your app looks like AI slop, that's not an AI problem. That's a workflow problem.
Most builders are stuck in the same cycle. You build something functional with AI. The features work. The logic is solid. But the UI looks generic and you know it. Your users know it too. And the moment someone lands on your product, they make a judgment in about two seconds.
The old fix was to hire a designer. Spend $3,000 to $10,000. Wait weeks for Figma files. Then spend even more time getting someone to actually implement those designs into your codebase. By the time the design was live, you'd already lost momentum and burned through cash you didn't need to spend.
That entire process is now optional.
I was able to take a client project that looked like every other vibe coded app and completely redesign it in under an hour. Professional typography. Consistent color system. A design language that actually holds together across every screen.
This is the workflow I'm running now. And I think every builder shipping with AI needs to understand how it works.
Why Stitch 2.0 Is Actually Different
I've tested a lot of AI design tools over the past year. Most of them generate something that looks decent on the first screen and then falls apart the moment you try to build a second page. Consistency is the problem. It always has been.
Stitch 2.0 solves this in a way nothing else has.
It's an AI native canvas. You can feed it screenshots of your existing app, drop in inspiration images from places like Dribbble or 21st.dev, or even paste a URL and let it reverse engineer the design of any website you like. It takes those creative seeds and generates full UI designs with multiple variants you can pick from.
But the real unlock is not the designs themselves. It's what Stitch builds in the background while it's designing.
A complete design system.
Typography scales covering display, headline, label, title, and body fonts. A primary, secondary, and tertiary color palette that's auto generated to be complementary. Color scales for every shade. Component rules and patterns. Elevation and depth specs. Even dos and don'ts for your design language.
All of this gets documented automatically in a file called design.md. This is a plain markdown file that captures every single design rule in one place. And this is the file that changes everything when you bring Claude Code into the picture.
How I Actually Use This on Real Projects
Let me walk you through exactly what I do. No theory. Just the workflow.
I start by taking screenshots of whatever I'm redesigning. If it's a client project or something I've been building myself, I screenshot the main screens and drop them directly onto the Stitch canvas. If I'm starting fresh, I'll grab two or three inspiration images from Dribbble instead. You don't need more than that. You're not looking for something to copy. You're looking for a direction.
Then I write one focused prompt. I tell Stitch what the app does, which screens I want redesigned, and the design direction I'm going for. Dark mode, minimal, editorial, whatever fits the product. I also specify font preferences because fonts are honestly the single fastest way to elevate how an app feels. Serif for headings, clean sans serif for body is a combination that works really well for most SaaS products.
Stitch generates multiple variants from that one prompt. This is important. Don't just accept the first output. Look at each variant and pull the best elements from each. The typography from one, the layout from another, the color energy from a third. You're curating, not just accepting.
One thing most people miss about why Stitch produces such strong output is that it generates images first before any code. That means it's not constrained by what HTML and CSS can do. It can imagine anything visually. Then you work backwards from that reference to build it. That's why the designs feel so much more polished than what you get from prompting a coding tool directly.
You can also talk to Stitch via voice now instead of typing. It transcribes your words into prompts automatically. Small feature but genuinely useful when you're deep in a session and want to keep moving fast.
Design.md Is the Real Game Changer
Once you're happy with your designs, go to the right hand panel in Stitch and click on Design Systems. You'll see that Stitch has already created one for you automatically based on everything you've been designing.
Click into it and you'll find the full design system I described earlier. Typography, colors, components, rules, everything documented and organized.
Now click on design.md.
Copy the entire file. Go to your project. Create a new file called design.md in the root directory. Paste it in. Save it.
That file is now the single source of truth for your entire design language.
Here's why this matters … continue on X ↗
♥ 1.0K · ⟲ 103 · 👁 805.5KView on X ↗

Indian Developer's Prompting Framework Tops GitHub, Author Shares Patterns

Indian Developer's Prompting Framework Tops GitHub, Author Shares Patterns

Brady Long promotes a GitHub prompting framework built by an Indian developer over 14 months and shares 11 prompt patterns from the repo that he says he has used for three weeks. The post is largely promotional and the claimed benchmark results are unverified.

Original post · 1 min read
🚨BREAKING: An Indian developer just hit #1 on GitHub with a prompting framework that outperforms every major benchmark.

No VC money. No research lab. Just a laptop and 14 months of testing.

Here are the 11 prompt patterns from his repo that I've been using for 3 weeks:
♥ 1.8K · ⟲ 286 · 👁 292.1KView on X ↗

Aiden Bai Launches Expect, Open-Source Browser Testing for Coding Agents

Alex Reibman endorses a post from Aiden Bai introducing Expect, an open-source tool that lets coding agents like Claude Code or Codex QA apps in a real browser, record videos of bugs, and fix them iteratively. It runs as a CLI or agent skill.

Original post · 1 min read
Ok this is exactly what I was looking for
Aiden Bai @aidenybai
Introducing Expect

Let agents test your code in a real browser

1. Run Claude Code / Codex to QA your app
2. Watch a video of every bug found
3. Fix and repeat until passing

Run as a CLI or agent skill. Fully open source
♥ 314 · ⟲ 8 · 👁 99.0KView on X ↗

Analysis Estimates Original 2008 IPL Team Bid Values in Rupees

Analysis Estimates Original 2008 IPL Team Bid Values in Rupees

Lalit Kumar Modi recounts the original IPL franchise bids awarded in January 2008, converting them to rupees at the exchange rate of that day, and argues that the value appreciation of teams is greater than media suggests. He notes actual figures would appear in company registrar filings.

Original post · 1 min read
This is original bids for @IPL that were awarded on 24th January 2008 in mumbai at cricket center at 12:00 pm. The bids were converted into rupees on that day. One dollar was 40 rupees on that morning. So one can just judge the true value appreciation today. 🙏🏽 further the amount bid was spread out to be paid evenly over 10 years. Many teams were in profit post year one itself. So the amount came out of their cash flow. So value appreciation is far greater than what the media conceives it to be. Actual numbers each team spent to buy the team will show up only in company filings with registrar of the company. So check that for accuracy and true value appreciation.
♥ 2.5K · ⟲ 317 · 👁 761.2KView on X ↗

Analyst Questions Harvey's $11 Billion Valuation Against Lexis and Westlaw

Matt Janiga questions whether legal AI startup Harvey's $11 billion valuation is justified given its competition with Lexis and Westlaw, which have large legacy data businesses. He argues Harvey lacks those datasets and would struggle to match their fee revenue.

Original post · 3 min read
The Harvey fundraise at an $11 Billion valuation is really interesting, and on the verge of head scratching.

Harvey feels like it competes with Lexis Nexis and Westlaw, which are the other two legal tools every major law firm has.

I used Lexis's AI tool a lot in my prior role. It was decent and seemed to improve over time. I assume Lexis will continue to improve it. It honestly competes with ChatGPT and Gemini more than Harvey.

The law firm lawyers I know who use Harvey like it, but it's not their sole AI tool. Like every AI tool on the market, it also has limitations and pain points.

Lexis is owned by RELX PLC and that conglomerate has a market cap of ~$65B. Westlaw is owned by Thompson Reuters and that conglomerate has a market cap of ~$55B (has been swinging, in part due to news about AI advancements and competitors like Harvey).

The interesting thing is that Lexis and Westlaw have legacy businesses built on datasets of legal precedents and carefully curated regulatory materials like opinion letters and legislative history. They also offer other products that drive material revenue, like Lexis's identity verification databases and value-added services.

Harvey doesn't have those things. And unless it can displace Lexis or Westlaw, it doesn't seem like it can earn the fees that those providers currently take from law firms on an annual basis. Legal revenue is an estimated 25% of Lexis's business — is Harvey really already on par with Lexis in the legal space vis-a-vis its $11B valuation? Westlaw drives closer to 40% of Thompson Reuters revenue, so maybe Harvey does still have room to double its valuation off of fee revenue. But that feels like a tough mountain to climb.

I'm also skeptical that Harvey can survive the thousands of paper cuts of lawyers opting for more general use AI tooling from the likes of Anthropic, Gemini and OpenAI. Anthropic has made amazing strides in general business work product, and all three are useful tools in developing memos and contracts.

There's also a last issue facing Harvey. If it replaces too many associates or associate hours, law firms aren't replacing costs — they're ripping out revenue generators. As someone who hires law firms, I'm not paying Cravath or MoFo $1,000 an hour for a partner to use Harvey. I'm paying those rates to get an associate, counsel or partner who has specific knowledge and skills to advance my project faster. It's great for me if Harvey usage shaves 5 hours off my bill on a project. But not good for the law firms, because I don't have some magic increase in projects to help them make up the lost revenue.

Law firms who adopt Harvey more will have to change their billing models. And I'm not sure you can teach that many old dogs the necessary number of new tricks to keep pumping up Harvey's valuation.

Okay. Rant over. Going to touch grass for 20 minutes.
♥ 334 · ⟲ 10 · 👁 72.8KView on X ↗

Sierra Releases Ghostwriter, an Agent That Builds Customer Service Agents

Sierra Releases Ghostwriter, an Agent That Builds Customer Service Agents▶

Bret Taylor announces that Sierra is releasing Ghostwriter, which lets enterprises create customer experience AI agents through conversation rather than forms, with voice, multilingual support, system actions and guardrails. He argues every software UI will eventually be an agent.

Original post · 1 min read
Today, Sierra is releasing Ghostwriter, our agent for building agents. With Ghostwriter, you can create an AI agent for your customer experience — one that can chat, pick up the phone, speak dozens of languages, take action on your systems of record, and be protected with industry-leading guardrails — simply by having a conversation. No clicking, no forms, no menus.

Codex and Claude Code have transformed how we build software, making it possible for software engineers to orchestrate and review the work rather than doing all the work themselves. We think the same transformation will happen for all software. Rather than every enterprise app having a web app for humans and an API for automation, every software platform’s UI will be an agent that can do the work on your behalf.

I recorded a demo of my building and optimizing an agent with Ghostwriter so you can see how powerful and easy it is to use. It’s completely changed the way our early adopters build agents, and it’s changed the way I think about the software industry. Let me know what you think, and, if you’re interested in trying it out at your business, please reach out directly.
♥ 3.2K · ⟲ 293 · 👁 1.1MView on X ↗

Dev-Browser CLI Uses Google WebMCP to Control Real Chrome Instance

Dev-Browser CLI Uses Google WebMCP to Control Real Chrome Instance▶

am.will praises dev-browser, which uses Google's WebMCP to drive the user's main Chrome instance with its existing sessions and cookies rather than a sandboxed Playwright browser, and describes how to enable remote debugging. The post is enthusiastic and offers little technical detail.

Original post · 1 min read
OMG you guys, this is incredible! This is using Google's new WebMCP function to control your browser, but not only is it lightning fast, but its unique because it is using your main Chrome instance.

Not some sandboxxed Playwright instance that doesn't want to remember your sessions, cookies, or passwords.

Your real Chrome instance. It's incredible.

You need to enable:

chrome://inspect/#remote-debugging

Also, it doesn't even require a skill to use. It just works. I'm thinking about making one anyway.

I'm telling you, download this and try it. This is my new daily for sure.
Sawyer Hood @sawyerhood
Introducing the new dev-browser cli.

The fastest way for an agent to use a browser is to let it write code.

Just `npm i -g dev-browser` and tell your agent to "use dev-browser"
♥ 2.7K · ⟲ 193 · 👁 458.8KView on X ↗
AI6/10

Developer Runs 35-Billion Parameter Model on $600 Mac Mini

Developer Runs 35-Billion Parameter Model on $600 Mac Mini▶

thestreamingdev reports running a 35-billion parameter AI agent on a 16GB M4 Mac mini by paging the model from SSD at about 30 tokens per second, claiming 18.6 times the speed of the same approach on NVIDIA hardware. The claims are shared in a thread with a demo video.

Original post · 1 min read
I ran a 35-billion parameter AI agent on a $600 Mac mini.
Specs: M4 Mac-Mini 16GB RAM

The model doesn't fit in RAM. It pages from the SSD at 30 tokens/second.

On NVIDIA, the same paging gives you 1.6 tok/s. Apple Silicon gives you 30. That's 18.6x faster.

No cloud. No API keys. $0/month.

Here's what it can do 🧵
♥ 3.2K · ⟲ 209 · 👁 732.5KView on X ↗

Sawyer Hood Launches Dev-Browser CLI for Agent Browser Automation

Sawyer Hood Launches Dev-Browser CLI for Agent Browser Automation▶

Sawyer Hood introduces the dev-browser CLI, which lets agents use a browser by writing code, installable via npm with a single command. The post includes a demo video.

Original post · 1 min read
Introducing the new dev-browser cli.

The fastest way for an agent to use a browser is to let it write code.

Just `npm i -g dev-browser` and tell your agent to "use dev-browser"
♥ 3.0K · ⟲ 281 · 👁 871.0KView on X ↗
AI7/10

Google Makes Lyria 3 Music Models Available in Public Preview

Google Makes Lyria 3 Music Models Available in Public Preview▶

Google for Developers announces that Lyria 3 and Lyria 3 Pro are in public preview via the Gemini API and Google AI Studio, offering two variants, tempo and song structure control, and image-to-music input. A demo video accompanies the post.

Original post · 1 min read
🎵 Lyria 3 and Lyria 3 Pro are now available in public preview via the Gemini API and in @GoogleAIStudio — and it’s music to our ears 🎵

🎼 Choose from two distinct variants to match production & latency needs (Lyria 3 Pro and Lyria 3 Clip)
📢 Direct the model with more precision & control (set specific tempos and song progression in your prompts)
🖼 Create projects using multimodal input support (such as image-to-music input)
♥ 182 · ⟲ 30 · 👁 83.2KView on X ↗

Ramp Data Shows Top AI Spenders Doubling Revenue Since 2023

Ramp Data Shows Top AI Spenders Doubling Revenue Since 2023

Eric Glyman reports that the top quartile of AI spenders on Ramp have more than doubled revenue since 2023 while the bottom quartile is flat, citing examples like a Texas roofing company and a Florida construction firm. He frames it as a widening gap that most businesses don't yet see.

Original post · 1 min read
Since 2023, the top quartile of AI spenders on @tryramp have more than doubled their revenue. Bottom quartile? Flat

A roofing company in Texas. A window installer in Utah. A construction firm in Florida that grew 65%

The gap is accelerating and most companies don't feel it yet
Eric Glyman @eglyman
Getting on the right side of the ice — The ice has cracked
If you want to understand what's about to happen to American businesses, picture one night in the Antarctic over a century ago.
Ernest Shackleton and 27 men were camped on ice
♥ 598 · ⟲ 68 · 👁 421.1KView on X ↗

Google's TurboQuant Compresses LLM KV Caches to Three Bits

Google's TurboQuant Compresses LLM KV Caches to Three Bits

Jen Zhu describes Google Research's TurboQuant, which compresses LLM key-value caches to 3 bits per value using random rotation and PolarQuant quantization, reporting at least 6x memory reduction and up to 8x faster attention with no measured accuracy loss. The post links to Google's research blog.

Original post · 1 min read
When I was consulting for @HBO Silicon Valley, zero-loss compression was the holy grail Richard Hendricks chases that perfect middle-out algo could shrink everything w/out breaking a single bit.

Google just did something even more practical for the AI era: TurboQuant compresses LLM key-value caches down to 3 bits per value using random orthogonal rotation + PolarQuant scalar quantization & optional 1-bit QJL residual correction.

=>> 6× memory reduction, up to 8× faster attention (on H100), & 0 degradation on LongBench, Needle-in-a-Haystack, and RULER for models like Gemma. No retraining, no calibration needed.

Fiction just got out-engineered by reality. 😅💚💚
Google Research @GoogleResearch
Introducing TurboQuant: Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency. Read the blog to learn how it achieves these results: research.google/blog/turboquant-redefining-ai-…
♥ 8.7K · ⟲ 671 · 👁 1.2MView on X ↗

Claire Vo Argues Design Teams Lag Behind in Corporate Influence

Claire Vo argues design culture is broken at many companies, with design teams resisting change, lacking political skill in campaigning for resources, and failing to make a quantified case. She responds to a Lenny's Newsletter observation that design hiring stalled as AI speeds up engineering.

Original post · 1 min read
I’ll say the thing no one is saying: design culture is broken in lots of companies.

Often design teams & designers are the most resistant to change org in the EPD triad, with highly vocal AI opponents, and little skill or interest in the art of campaigning for influence or resources. Won’t hold a number like a PM, not yelled at about timelines like engineering. While I have brought design topics to the board convo, not a single board has pressed me our design talent, strategy, or velocity. Most teams treat design like a tax they don’t want to pay, and those that *do* take a deep interest and want to invest in design get back big “get out of my figma” energy. And if you’re too precious about craft to dirty your hands with the dark art of corporate politics, good luck getting more headcount. If a PM or engineer can get 85% there with tailwind and a dream, you better come to the table with more than “I represent the user.”

Great designers are worth more than almost anyone on the team, and I’ve worked with lots of gems, but this is 0% surprising to me.
Lenny Rachitsky @lennysan
I don’t know exactly what’s going on here, but it does feel AI-related. Unlike PM and eng, which started growing in 2024 (two years post-ChatGPT), design didn’t. If I had to venture a theory, I’d say that because AI is allowing engineers to move so quickly, there’s less opportunity—and less desire—to involve the traditional design process.

That said, you’d think design would become a differentiator as more products compete for attention. Something to think about for your company! We’ll keep watching this trend and AI’s impact on org design more generally.

One interesting observation we made …
♥ 819 · ⟲ 62 · 👁 188.4KView on X ↗
AI7/10

Eric Schmidt Says Defining Success Now Matters More Than Execution

Eric Schmidt Says Defining Success Now Matters More Than Execution▶

In a video clip, Eric Schmidt argues that the key advantage in AI-driven work is precisely specifying problems and evaluation functions, after which systems can run and produce results overnight.

Original post · 1 min read
Eric Schmidt says the 10x advantage is no longer execution. It is defining what counts as success.

A programmer writes a spec and an evaluation function, runs it at 7pm, and wakes up to what was invented overnight.

The advantage now belongs to whoever can specify the problem precisely.

The rest will be automated.
♥ 2.7K · ⟲ 332 · 👁 393.0KView on X ↗

Aakash Gupta Says AI Product Management Careers Lie Deeper in the Stack

Aakash Gupta Says AI Product Management Careers Lie Deeper in the Stack▶

Aakash Gupta argues that application-layer AI PM roles are crowded and low-ceiling, while higher-paying roles require skills like probabilistic thinking and model evaluation. He promotes a podcast episode with a former AI PM at Netflix, Amazon and Meta.

Original post · 1 min read
The "easiest" path into AI product management is also the most crowded and lowest-ceiling.

There's a stack of AI PM roles. At the top: Application PMs. They own the user experience layer. How users interact with AI, how you build trust, how you make AI reliable for everyday use. This is the closest to traditional product management. And that's exactly the problem.

Every PM repositioning into AI right now is aiming at this layer. They shipped a chatbot feature. They designed an AI-powered search experience. They added "AI" to three bullet points on their resume. The application layer is where the conversion is easiest and the competition is most brutal.

She's been an AI PM at Netflix, Amazon, and Meta. Her breakdown of the full stack on this episode revealed something most career advice skips: the layers below the application tier require fundamentally different skills. Not UX intuition.

Probabilistic thinking. Model evaluation. Understanding why the AI is unreliable, not just managing the user's perception of reliability.

The $900K roles don't live at the layer everyone is rushing toward. They live deeper in the stack, where the supply of qualified PMs drops off sharply.

The roadmap isn't "get into AI PM." It's "get into the right layer."
Aakash Gupta @aakashgupta
AI PMs at Netflix get paid $900K+.

She's been an AI PM at not just Netflix, but also Amazon and Meta. And today, she broke down how you can too:

1:43 Types of AI PMs
7:11 - Technical Concepts Masterclass
58:57 - How to Job Search Well
♥ 28 · ⟲ 1 · 👁 9.1KView on X ↗

Manthan Gupta Applies Karpathy's Autoresearch Idea to LLM Inference

What Happened When I Applied Karpathy's Autoresearch Idea to LLM Inference

Manthan Gupta describes building Auto-Inference-Optimiser, an open repo where an AI coding agent hill-climbs on MLX inference speed on Apple Silicon under a locked evaluation harness. The article focuses on the failed experiments and fake wins behind benchmark gains.

Original post · 10 min read
X ArticleWhat Happened When I Applied Karpathy's Autoresearch Idea to LLM Inference
Most "AI optimization" demos are fun to watch for the same reason benchmark tweets are fun to watch: they show you the win, not the search.
You see the final graph. You see the +12% or the "runs 2x faster now" claim. What you usually do not see is the graveyard of bad ideas behind it. The settings that looked promising but were just noise. The optimizations that made throughput better by quietly making the model worse. The fake wins that only happened because the benchmark got easier.
So I built a small repo called Auto-Inference-Optimiser (star the repository!) to study exactly that.
The idea is simple: lock the evaluation, open one file for experimentation, and let an AI coding agent hill-climb on inference speed forever on Apple Silicon.
The most interesting part was not that it achieved a speedup. It was the kind of speedup it achieved, what it failed to improve, and what that indicates about inference engineering on real hardware.
Let's get into it.
Why I Built This
I care a lot about inference right now.
Not in the abstract "LLMs are cool" sense. I mean the actual production questions: where latency comes from, what batching buys you, what prompt processing costs, how KV cache decisions change throughput, and where the hardware wall starts pushing back.
There is a lot of content online about training. There is also a lot of content online about agents. But there is still not enough content that combines both instincts: build a tight experimental harness, let the agent search inside it, and use that process to learn something real about inference.
This repo was my way of doing that.
It is inspired by Karpathy's Autoresearch, but pointed at a different layer of the stack. Instead of searching over training code on a GPU box, this one searches over an MLX inference pipeline on a Mac because I am GPU poor (please sponsor a GPU).
What The Repo Actually Does
At a high level, the repo turns "make inference faster" into a bounded optimization problem.
The structure is intentionally small:

That boundary is doing most of the work.
prepare.py is read-only. It fixes the benchmark model, the prompts, the warmup behavior, the averaging logic, and the quality gates. The agent cannot "win" by quietly changing the test.
inference.py is the search surface. That is where the agent is allowed to touch sampling, prefill step size, prompt formatting, and the general generation path.
program.md tells the agent how to behave:

That is the core harness.
And I like this design a lot because it bakes in three things that most autonomous coding demos hand-wave away:
Reversibility - bad ideas are cheap to discard.
Observability - every run leaves behind metrics and logs.
Constraints - the agent is not allowed to optimize by moving the goalposts.
The Evaluation Is The Real Product
The truth is that the most important file in this repo is not inference.py. It is prepare.py.
That file fixes the benchmark around a small Apple Silicon friendly model, runs warmups, averages across multiple runs, and evaluates five different prompt types:
explanation
long-context summarization
reasoning
creative generation
code generation
That already makes the benchmark better than a lot of speed demos, because decode-heavy and prefill-heavy cases behave differently.
But the more important choice is the quality gate.
This repo does not let the agent optimize only for tokens/sec. It requires two checks to pass:
avg_perplexity has to stay below a threshold
sanity_check has to stay above a threshold
That second gate matters a lot.
Perplexity is useful, but it is still a model-internal metric. It can tell you that outputs are becoming unstable or degenerate, but it does not fully tell you whether the answer is still usable. So the repo also checks for concrete task-level correctness: did the train speed answer contain 48? Did the transformer explanation mention the right ideas? Did the LCS prompt actually return something that looks like Python code?
This is one of my favorite design choices in the whole project.
Because if you do not defend quality explicitly, an optimization harness will absolutely "improve" your system by making it worse.
What Actually Worked
After the optimization runs, the pattern was surprisingly clear.
Here is the short version:

But the more interesting part is where they came from.
1. Argmax sampling was the biggest win
On the Qwen run, setting sampling to greedy decoding gave the largest gain: about +10.8% generation throughput.
On the Llama run, it was also the best keep: about +2.6%.
That tells you something important: sampling overhead is not free. Top-p decoding is doing real work every token, and if your objective is pure throughput, removing that work can matter more than a lot of fancier ideas.
Of course, there is a trade-off.
You get deterministic output and lose diversity. So this is not a universal recommendation for every product. But as an inference lesson, it is very clean: sometimes the fastest path is just doin… continue on X ↗
♥ 516 · ⟲ 44 · 👁 216.9KView on X ↗
AI6/10

Shiv Lists Startups Building Infrastructure for AI Agent Economy

Shiv Lists Startups Building Infrastructure for AI Agent Economy

Shiv lists thirteen companies building primitives that let AI agents act as users, including email, phone numbers, browsers, sandboxes, memory, payments, voice and web search. He frames the trend as an economy of AI coworkers.

Original post · 1 min read
Lots of companies are now building primitives for an economy where AI agents are the primary users instead of humans.

They're betting on an economy of AI coworkers.

1. AgentMail (@agentmail): so agents can have email accounts

2. AgentPhone (@tryagentphone): so agents can have phone numbers

3. Kapso (@andresmatte): so agents can have WhatsApp phone numbers

4. Daytona (@daytonaio) / E2B (@e2b): so agents can have their own computers

5. Browserbase (@browserbase) / Browser Use (@browser_use) / Hyperbrowser (@hyperbrowser): so agents can use web browsers

6. Firecrawl (@firecrawl): so agents can crawl the web without a browser

7. Mem0 (@mem0ai): so agents can remember things

8. Kite (@GoKiteAI) / Sponge (@PayspongeLabs) : so agents can pay for things.

9. Composio (@composio): so agents can use your SaaS tools

10. Orthogonal (@orthogonal_sh) so agents can access APIs easily

11. ElevenLabs (@ElevenLabs) / Vapi (@Vapi_AI) so agents can have a voice

12. Sixtyfour (@sixtyfourai) so agents can search for people and companies.

13. Exa (@ExaAILabs): so agents can search the web (Google doesn’t work for agents)

If you stitch all of these together, you get a digital coworker that looks more human than AI.
♥ 2.2K · ⟲ 238 · 👁 279.1KView on X ↗

Insanely Fast Whisper Promoted as Free Local Transcription Tool

Insanely Fast Whisper Promoted as Free Local Transcription Tool

Nav Toor promotes Insanely Fast Whisper, an open-source MIT-licensed tool that runs Whisper transcription locally on NVIDIA GPUs or Apple Silicon. He claims 150 minutes of audio transcribes in 98 seconds, citing benchmarks against paid cloud and transcription services.

Original post · 1 min read
🚨 OpenAI charges $0.006/minute. Google charges $0.024. AWS charges $0.024.

Someone just open sourced a tool that does it for $0. And it's faster than all of them.

It's called Insanely Fast Whisper. And that's not hype. That's the benchmark.

150 minutes of audio. 98 seconds to transcribe. On your own machine. No API key. No cloud. No per-minute billing.

Here's what the numbers look like:

→ Whisper Large v3 + Flash Attention 2: 150 min of audio in 98 seconds
→ Distil Whisper + Flash Attention 2: 150 min in 78 seconds
→ Standard Whisper without optimization: 31 minutes for the same job
→ That's a 19x speedup. Same model. Same accuracy. Just faster.

Here's what it does:

→ One command to transcribe any audio file or URL
→ Speaker diarization — knows WHO said WHAT
→ Transcription AND translation to other languages
→ Runs on NVIDIA GPUs and Mac (Apple Silicon)
→ Flash Attention 2 for maximum speed
→ Clean JSON output with timestamps
→ Works with every Whisper model variant

Here's the wildest part:

Otter.ai charges $100/year. Rev charges $1.50/minute. Descript charges $24/month. Enterprise transcription contracts cost thousands.

Podcasters, journalists, researchers, lawyers, content creators — anyone still paying for transcription is lighting money on fire.

8.8K GitHub stars. 633 forks. MIT License.

100% Open Source.

(Link in the comments)
♥ 5.6K · ⟲ 489 · 👁 509.5KView on X ↗

Claire Vo Shares Practical Tips for Running OpenClaw Agents

Claire Vo lists practical tips for OpenClaw, including session resets at 4 a.m., remote screen sharing to a Mac mini, heartbeat versus cron configuration, browser profile colors, read-only Google auth and API key hygiene.

Original post · 1 min read
Random @openclaw tips that are super simple but almost no one realizes
- your sessions poof at 4 am, overnight amnesia is built into the system
- you don’t need a monitor for your Mac mini turn on screen share & remote in from your laptop
- it’s SOUL is promoted to not bother you overnight or talk to much, it tries to get you to go to bed
- you probably have your session dmscope or cron target sessions set wrong which is causing your claw to act like it has a tbi
- it’s probably stashed an API key somewhere it shouldn’t
- you can give browser profiles their own color so you can tell when it’s working in its profile
- you can auth gog with read only perms
- read the docs on heartbeat vs cron
- give your telegram bots cute emojis
♥ 451 · ⟲ 22 · 👁 61.6KView on X ↗

Pieter Levels Revives ThisHouseDoesNotExist With Newer Image Models

Pieter Levels Revives ThisHouseDoesNotExist With Newer Image Models▶

Pieter Levels announces he has revived ThisHouseDoesNotExist.org, his 2022 AI architecture project, migrating it to a Hetzner VPS and upgrading it to current image models. The site now generates about twelve new designs daily and lets users vote on them.

Original post · 1 min read
✨ I've brought back ThisHouseDoesNotExist.org from the dead

It was my first visual AI project in 2022, and it's this project that generated random @ArchDaily-style architecture designs that made me realize AI image models could do interior design

That led me to make InteriorAI.com, which then led me to finetune my first interior design model which then for fun I uploaded my own photos too, which led me to make AvatarAI.me and then pivoted that in to PhotoAI.com

So this project has a special place for me

It was still alive but wasn't generating new designs anymore because it ran on Stable Diffusion 1.5 which was outdated and everything stopped working about a year ago

I've now migrated it to its own Hetzner VPS now, which means I can run Claude Code on the server with it, and cleaned it up and upgraded it to the latest AI image models (including Nano Banana Pro)

It now generates about 12 new designs every day again, and you can up or downvote the ones you like or don't like!
@levelsio @levelsio
✨ My new project is now live:

thishousedoesnotexist.org/

🏡 It uses A.I. to let you generate @ArchDaily-style modern architecture houses on-the-fly

If you generate any nice ones, reply them here pls 😊
♥ 292 · ⟲ 13 · 👁 204.8KView on X ↗

Griffin Hilly Publishes Claude Code Setup Built From Saved Bookmarks

One Claude Code Setup to Rule Them All

Griffin Hilly presents a Claude Code workflow distilled from bookmarked posts, including a CLAUDE.md operating model with plan-first protocols, orchestrator-first delegation and dialectic reviews. The setup is published as a GitHub repository for others to clone.

Original post · 5 min read
X ArticleOne Claude Code Setup to Rule Them All
You've seen the X articles.
Maybe you've downloaded Claude Code but you weren't sure where to get started.
You saw @Karpathy's post about auto-research but you're not sure if you have anything interesting to research.
You saw a bunch of tweets about things to add to your CLAUDE.md, but then you saw another saying your CLAUDE.md was too long and should be cut down to basics.
You're lost and you don't know what to do.
This is the tweet for you.
I've read all those tweets for you. Or more accurately, I've read some of them and my Claude has read them all. And my Claude distilled all of that into a single, simplified workflow that you can copy for yourself.
Just clone this repository and you're good to go: github.com/griffinhilly/claude-code-synthesis
Alternatively you can have your Claude do it for you.
Every week I download all the bookmarks I've saved and have my Claude read them. We consider adding anything we don't have already in our workflow. Your Claude can do the same.
Here's what Claude and I have found:

The Operating Model (CLAUDE.md)
The single most important file. Copy it to ~/.claude/CLAUDE.md and it changes how Claude approaches every task. Here's what's in it:
Leverage Doctrine. You do the thinking. Claude does the doing. You ideate, decide, and steer. Claude researches, implements, and executes. When uncertain, it surfaces options with tradeoffs instead of deciding silently.
Plan-First Protocol. Every task starts with: what's the objective? How will we know it worked? What are the sub-tasks? Research agents plan, implementation agents execute. Never both.
Scope Discipline. Claude pushes back on ambitious plans. "This is a 3-session project. Want to start with just X?" A working smaller thing beats a half-finished grand vision.
Orchestrator-First. The session agent is a manager, not a worker. Before any task, it decides: handle directly, delegate to a subagent, or route to MCP? This is the single biggest lever for productivity.
Dialectic Reviews. For important decisions, don't ask "what should I do?" Spawn opposing agents — one argues FOR, one argues AGAINST — with a referee to synthesize. Dramatically better than asking one agent for pros and cons. (h/t @danpeguine and @systematicls for the Hunter/Skeptic/Referee pattern)
Anti-Sycophancy. If an approach has clear problems, Claude says so directly, proposes an alternative, and accepts override. Sycophancy is a failure mode.
Test-First Bug Fixing. When a bug is reported, write a test that reproduces it before trying to fix it. (h/t @tangming2005 — this was his "single biggest improvement to my CLAUDE.md")
Operationalize Every Fix. Don't just fix the bug. Write tests that catch the whole class of similar bugs. Check for other instances. If it reveals a gap in your instructions, update CLAUDE.md. Every bug is a learning opportunity. (h/t @doodlestein)
Evals Before Specs. Define how you'll evaluate success before writing the spec. The progression: evals → spec → plan → implement → verify. (h/t @synopsi, who now spends 90% of time on evals)
Prefer the Boring Solution. Can this be fewer lines? Are abstractions earning their complexity? Don't build 1,000 lines when 100 suffice. (h/t @karpathy — "agents bloat abstractions, have poor code aesthetics")
Progressive Disclosure. Don't dump everything into CLAUDE.md. Keep it lean with trigger rules ("when X happens, read guide Y"). Guides load on-demand. (h/t @toddsaunders and @mstockton)
Structured > Prose. For rules agents MUST follow, use XML tags and JSON, not markdown paragraphs. Claude processes tagged content differently. (h/t @ihtesham2005 and @Austen)
Workflow Evolution. The workflow is a living system. When a session reveals a new pattern, encode it. Use your tools to improve your tools. It's a flywheel, not a static config. (h/t @doodlestein's Agent Flywheel) ALSO SERIOUSLY, GO FOLLOW @doodlestein
Corrective Framing. "Remember to do X" doesn't work. Instead, present a possibly-wrong claim: "You should be doing X — are you still doing it?" Mismatches trigger natural correction. (h/t @yishan)

The COMP System
Every project gets 4 files:
- CLAUDE.md — how the AI should behave here
- ORIENT.md — what a human needs to know to work here
- MEMORY.md — accumulated decisions, gotchas, context
- PLAN.md — roadmap, progress, next steps
Separate behavioral instructions from human orientation from accumulated knowledge from direction. Each has a different audience and update frequency.

The Guides
Seven situational guides that Claude loads on-demand — delegation templates (7 agent types with prompt structures), shell rules, context efficiency, API-over-scraping, PostgreSQL batching, overnight autonomous runs, and a skills reference.
The delegation templates alone are worth the clone. Implementer, Researcher, Reviewer, Batch Worker, Session Reviewer, Explorer, Creative — each with a prompt template, model recommendation, and mandatory report format.

The Bookmark Pipeline
This is the meta-move. Every week… continue on X ↗
♥ 224 · ⟲ 18 · 👁 114.4KView on X ↗

United Unveils Relax Row Three-Seat Lie-Flat Economy Product

John Collison shares United Airlines' announcement of Relax Row, three adjacent economy seats with adjustable leg rests that form a lie-flat space, arriving next year on more than 200 787s and 777s. He comments that it is the product airline passengers have long wanted.

Original post · 1 min read
United built the product that everyone who has every been on an airplane has wanted!
United Airlines @united
The entire row is alllllll yours.

Welcome to United Relax Row, three adjacent United Economy seats with adjustable leg rests that can each be raised or lowered to create a cozy lie-flat space for stretching out...

You'll also get a mattress pad, blanket and two pillows. If you’re traveling with kids, a plushie too! United Relax Row will be available starting next year on more than 200 of our 787s and 777s, each with up to 12 of these brand-new rows.

united.com/Elevated
♥ 14.3K · ⟲ 376 · 👁 3.7MView on X ↗

Wade Foster Recounts Zapier's Early YC Advice to Launch and Grow

Wade Foster Recounts Zapier's Early YC Advice to Launch and Grow

Wade Foster recalls Zapier's 2012 Y Combinator office hours, where Garry Tan urged them to launch immediately and Paul Graham challenged them to grow revenue 10 percent week over week. He argues growth rate is the key signal for startups.

Original post · 1 min read
It was May 2012, and we hadn't yet launched @Zapier.

Two YC office hours changed everything. The first was with @GarryTan.

He had one question: "Have you launched?"

We said no. We had users. People had paid us. But in our minds, the product wasn't good enough.

Garry didn't care. He said launch now. We did and realized we were dumb to wait.

Our second office hours was with @PaulG. Same question: "Have you launched?"

This time we got to say yes. So he gave us an assignment: "Grow revenue 10% week over week. 10% is great. 20% is fantastic."

So we set off to grow 10% a week. And we did.

Last week PG posted this thread, and it brought all this back.

10% a week is 142x a year. Even at scale, growth rate is the signal.

Focus on growth rate and you'll find the future.
♥ 626 · ⟲ 29 · 👁 56.2KView on X ↗

Millie Marconi Highlights Open-Source Claude Code Skill for Recent Prompts

Millie Marconi Highlights Open-Source Claude Code Skill for Recent Prompts

Millie Marconi promotes an MIT-licensed Claude Code skill that scans Reddit and X over the last 30 days on a given topic and generates ready-to-use prompts based on community findings. She says it works for areas including ChatGPT, Midjourney, Suno and Cursor.

Original post · 1 min read
This feels like cheating.

Someone built a Claude Code skill that scans Reddit and X from the last 30 days on any topic you give it, then writes you copy-paste-ready prompts based on what the community has actually figured out not what was working six months ago.

You type /last30days prompting techniques for ChatGPT for legal questions and it comes back with the top patterns real lawyers and power users are using right now, complete with a fully written prompt you can drop in and use immediately.

No more Googling, no more digging through threads, no more prompts that worked last year but got patched out.

It works for anything - Midjourney techniques, Suno music prompts, Cursor rules, trending rap songs, whatever you need to know what people are actually saying about right now.

100% Open Source. MIT License.

Link in the comments.
♥ 5.3K · ⟲ 427 · 👁 487.7KView on X ↗

Anthropic Lets Claude Control Users' Computers in Research Preview

Anthropic Lets Claude Control Users' Computers in Research Preview▶

Anthropic's Claude account announces a research preview that lets Claude use a user's computer to open apps, navigate browsers and fill in spreadsheets. It is available in Claude Cowork and Claude Code on macOS only.

Original post · 1 min read
You can now enable Claude to use your computer to complete tasks.

It opens your apps, navigates your browser, fills in spreadsheets—anything you'd do sitting at your desk.

Research preview in Claude Cowork and Claude Code, macOS only.
♥ 136.6K · ⟲ 14.0K · 👁 78.4MView on X ↗

Lenny Rachitsky Explains How Executive Calendars Shape Product Decisions

Lenny Rachitsky Explains How Executive Calendars Shape Product Decisions▶

Lenny Rachitsky describes the fragmented, back-to-back nature of an executive's day and why product managers overestimate how much leaders remember of their pitches. He shares this via a video and a quoted post about Jessica Fain's chief-of-staff pitch to Slack's CPO.

Original post · 1 min read
People don't understand executive calendars.

I describe an executive's calendar as like a strobe light going off.

You wake up at 8AM, you've already got a huge list of urgent things going on.

You go from a meeting with finance on a budget, to an interview for another executive, to a people problem, to a legal problem, to a product review.

And the product manager coming to that product review, who's trying to make a pitch thinks I've been prepping for this meeting for two weeks.

But the executive coming into that session hasn't thought about you since.
Lenny Rachitsky @lennysan
Jessica Fain's best product ideas kept dying, and she couldn't figure out why.

So at eight and a half months pregnant, she pitched @SlackHQ's CPO @aunder on becoming her Chief of Staff. She wanted to see how executive decisions actually get made from the inside.

What she learned changed everything she knew about influencing execs.

People don't realize that an executive's calendar is like a strobe light going off. Budget meeting, a people problem, a legal issue—then your product review. You've been prepping for three weeks. They haven't thought about you since the last meeting. They may not …
♥ 338 · ⟲ 21 · 👁 237.5KView on X ↗
AI8/10

Leopold Aschenbrenner Publishes Essay Series on AGI Strategic Outlook

Leopold Aschenbrenner Publishes Essay Series on AGI Strategic Outlook

Leopold Aschenbrenner argues that few people are pricing in coming AI progress and shares a multi-part essay series titled Situational Awareness: The Decade Ahead. The series covers deep learning trendlines, compute scaling, the international situation and a hypothetical government-led project.

Original post · 1 min read
Virtually nobody is pricing in what's coming in AI.

I wrote an essay series on the AGI strategic picture: from the trendlines in deep learning and counting the OOMs, to the international situation and The Project.

SITUATIONAL AWARENESS: The Decade Ahead
♥ 13.3K · ⟲ 1.8K · 👁 8.9MView on X ↗