Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

Latest stories — use the filters to narrow by keyword, section or date.

Chamath Palihapitiya Recommends Jeffrey Katzenberg's Essay on AI for Creativity

Chamath Palihapitiya shares Jeffrey Katzenberg's essay 'The World is Changing: AI For Creativity' and calls it worth reading. Katzenberg describes watching a founder demo a stunning AI-generated animated scene and reflects on its impact on artists.

Original post · 1 min read
This is worth reading.
Jeffrey Katzenberg @jeffreykWNDR
The World is Changing: AI For Creativity

By Jeffrey Katzenberg

A few months ago, I sat in my office in Silicon Valley and watched as a tech founder showed me something extraordinary. On the screen was a fully realized, beautifully lit, well-composed animated scene. It was stunning and it made me feel exactly what I felt in 1986 watching Luxo Jr. That was the first time I watched a computer-animated 3D character take a breath and seem, against all reason, to have life. It left me in awe.

Later that day, I received a text from an artist I've known for thirty years, 350 miles to the south, in …
♥ 1.4K · ⟲ 75 · 👁 464.8KView on X ↗

Amazon Launches Boomerang Rehiring Push for Former Employees

Gergely Orosz comments on a Business Insider report that Amazon is recruiting former employees, including some it laid off, under a 'Boomerang Reengagement Initiative.' He notes that Amazon's hiring and firing policy was built on not rehiring most departing staff.

Original post · 1 min read
This is strange to see, because Amazon's complete hiring and firing policy has been built on the notion that they do NOT want to re-hire the majority of the people leaving (who are marked as non-regretted attrition, even if they are regretted - incentives!) but hire new instead
Eugene Kim @eugenekim222
New: Amazon wants former employees back — including some it laid off. In one email, a recruiter called the push the “Boomerang Reengagement Initiative.” In another, a recruiter asked whether Amazon’s RTO policy had driven the ex-employee away.

businessinsider.com/amazon-boomerang-hiring-re…
♥ 1.1K · ⟲ 37 · 👁 172.5KView on X ↗

Anthropic Details How It Made claude.ai Three Times Faster

Claude

Boris Cherny highlights an Anthropic blog post describing how the team made claude.ai three times faster in two weeks using Claude to measure, debug and improve performance. The post includes prompts and methods for engineers optimizing their own apps.

Original post · 1 min read
If you've noticed how fast claude.ai/login and the Desktop app have become in the last few weeks, here's how we did it.

Lots of juicy learnings & techniques in the blog post for engineers working on speeding up your own apps.
ClaudeDevs @ClaudeDevs
We made claude​.ai 3x faster in two weeks.

Here’s how we use Claude to measure, debug and improve performance. Prompts and methods included.

claude.dev/blog/how-we-made-claude-ai-faster/
claude.aiClaudeClaude is Anthropic's AI, built for problem solvers. Tackle complex challenges, analyze data, write code, and think through your hardest work.
♥ 5.3K · ⟲ 168 · 👁 847.5KView on X ↗

Veed Open-Sources OpenEdit for Agent-Driven Video Editing

Veed Open-Sources OpenEdit for Agent-Driven Video Editing▶

Sabba Keynejad introduces OpenEdit, an open-source agent-driven pipeline for creating subtitles, motion graphics, slides and rendered videos. He argues the challenge is native editing with fonts, branding and repeatable templates, not generating video from code. A linked post by Deedy Das describes producing a launch video for about $2.

Original post · 1 min read
Generating a good video from code is not the problem.

The problem is how you edit it natively.

How you use use fonts, branding and build repeatable templates.

And that’s why we built OpenEdit.

github.com/veedstudio/open-edit
Deedy @deedydas
Opus 5.5 is incredible at instructional video generation.

I made this launch video for a inference startup in 1min for ~$2. Videos like these used to take weeks if not months and a lot of coordination with agencies and 1000x the costs.

Humans broadly prefer video to text. This changes the substrate of communication. These videos actually help communicate technical ideas in seconds (photorealistic video gen like Seedance is not very useful here).
- changes how often marketing should be talking about products and launches
- change how sales people can talk about technical products to their cus…
github.comGitHub - veedstudio/open-edit: Open-source, agent-driven editing pipeline: create subtitles, motion graphics, slides, edit and render videos.Open-source, agent-driven editing pipeline: create subtitles, motion graphics, slides, edit and render videos. - veedstudio/open-edit
♥ 297 · ⟲ 5 · 👁 50.6KView on X ↗

Cursor Engineer Shares Prompt for Improving Agent Token Efficiency

Eric Zakariasson shares a detailed prompt based on Cursor's experience for optimizing an LLM agent harness's token efficiency without hurting task quality. It covers measuring cost per completed task, weighting token billing types, and avoiding prompt patterns that make models reluctant to work.

Original post · 15 min read
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy

# Improve this agent harness's token efficiency

You're working on an LLM agent harness: the system prompt, tool definitions, request assembly, context caching, compaction, and retrieval, and how work is split across agents. Make the agent's runs cheaper without making it worse at its job.

- Objective: lower price-weighted token cost per completed task.
- Constraint: no measurable drop in task quality.

Measure per task, not per request. Every turn resends the prefix (tools, instructions, setup, and the conversation so far), so a change that shrinks each request but adds turns can cost more. Weight tokens by billing type: output, uncached input, and cached input are priced very differently.

Work in this order: map the harness and measure the baseline, rank the opportunities, make the changes that are safe to make directly, put the rest behind flags or in proposals, then report.

Figures below come from one team's production coding agent and its multi-agent experiments. Use them to gauge magnitude, not as targets. One round of these changes (prompt trimming, tool offloading, cache layout, sparse line numbers, subagent tuning) cut that team's overall token cost about 7% with no loss in quality. The larger percentages apply only to the part of the request each change touched.

## Principles

1. Change what the harness sends, not how hard the model tries. Don't ask the model to conserve tokens. A harness that told its model to "take care to preserve tokens and not be wasteful" found it grew reluctant to take on ambitious tasks and sometimes quit, saying it wasn't supposed to waste tokens.
2. Capable models need definitions, not commands. Lists of "DO NOT", "You must", and "Important", and guards against older models' habits, can usually be replaced with plain descriptions of what each tool does. One team cut about two-thirds of its system prompt this way, and the shorter prompt worked across model families. Instruct only on what the model can't know (the product, the environment, the user's processes) and on quirks you've seen in transcripts.
3. Static context is for what most turns need. Everything else should be discoverable when needed. Less up-front context also means less confusing or contradictory information.
4. Expect removals to win. Guardrails written for weaker models, coordination steps that became bottlenecks, and prompting for behavior the model now does on its own all cost tokens.
5. Real usage decides. Evals are a fast proxy, but they skew toward hard problems and miss the real mix of requests.

## 1. Map the harness and measure the baseline

Find:

- Where requests are assembled, the system prompt, and tool schemas. If a framework or SDK builds requests, find its hooks for message order, cache control, and tool loading.
- How tool results are formatted, and how history is kept, trimmed, or summarized.
- How subagents or parallel agents are spawned, if any.
- Which models and provider APIs are used. From the provider's docs, get the prompt caching behavior (automatic or explicit breakpoints, TTL, minimum cacheable length) and the prices for output, uncached input, and cached input.
- Existing logging, token accounting, and evals.

If the harness doesn't record per-request token usage by billing type and cache hits, add that first. Everything later depends on it.

Then render a few real requests (from logs, or by running representative tasks) and count tokens per section with the model's tokenizer or the API's usage fields. Produce:

- Cost share by source × billing type. Sources: system prompt, tool definitions, skill/rule/integration descriptions, user messages, file reads, search results, command and other tool output, history, summaries, subagents.
- Static tokens per request, cache hit rate, and turns per task.
- Per tool: the share of runs that call it at least once, and its error rate.

Read the rendered requests, not just the templates. Duplication, leaked volatile values, and misordered blocks only show up there.

Rank opportunities by share of spend × fraction removable ÷ quality risk.

## 2. System prompt and injected context

Label every instruction:

- Keep: product or environment knowledge the model can't infer, fixes for quirks seen in this model's transcripts, and rules a mode depends on.
- Rewrite: commands and emphasis into plain descriptions. Reminders into constraints: "No TODOs, no partial implementations" works better than "remember to finish implementations." Vague quantities into ranges: "generate 20–100 tasks" gets far more ambitious behavior than "generate many tasks."
- Delete: things capable models do by default, guards against behavior you haven't seen from this model, text that repeats tool descriptions, and lines that could contradict a user request. Models trained to rank system instructions above user messages will side with the system prompt.
- Move: anything per-user or per-request (date, environment, repo state, lists of skills or subagents, user rules) into a user-role setup message after the cache boundary.

Audit other injected context the same way. As models improved, the team behind these figures dropped directory trees, pre-retrieved snippets, compressed copies of attached files, lint errors injected after every edit, forced expansion of short file reads, and caps on tool calls per turn. They kept small, high-value facts: OS, repo status, and open or recently viewed files.

Skip checklists for open-ended work. The model optimizes the listed items and deprioritizes everything else.

## 3. Tool definitions

Tool schemas ride along on every request. Most tools beyond the core set were each needed in under 20% of conversations, and moving them out of static context cut tool-description tokens 60%. Doing the same for integration tools (such as MCP servers), with names in context and full schemas in one folder per server that the agent can search with grep or jq, cut total tokens 46.9% in sessions that used them.

- Keep in static context: high-frequency tools (for a coding agent: read, search, edit, shell), tools the model tries to call even when they're absent, and tools a mode depends on.
- Offload the rest: leave a name or one-line pointer and make the full schema discoverable on demand. Group related tools so they load together, and put status (such as "needs re-authentication") where the agent will see it.
- Tighten what remains: describe behavior and arguments, and drop usage lectures.
- Pick the split by testing a few configurations and tracking tokens, cost, latency, tool-call errors, and task success.

## 4. Cache layout

Order each request so the reusable prefix is as long as possible:

`tool definitions → system instructions → [breakpoint] → setup message (skills, subagents, rules, environment) → [breakpoint] → conversation`

- Keep the prefix byte-identical across turns. Use deterministic tool order and serialization, put timestamps and IDs after the boundary, and don't rewrite earlier messages except when compacting.
- Use explicit breakpoints if the provider supports them. Otherwise rely on automatic prefix caching with the stable part first. Respect TTL and minimum-length rules.
- Switching models mid-conversation throws away the cache (caches are per model and provider) and hands the new model a history it didn't write. When a different model is needed, run it as a subagent with fresh context.

Explicit breakpoints plus moving per-request setup after them cut cold cache misses 20%.

## 5. Tool results and other context added during a run

- Large outputs (commands, integrations, logs): write them to a file and return the path, size, and a short tail. The agent can tail, grep, or read ranges for more. Truncating loses data, and inlining bloats every later request. Treat long-running terminal sessions the same way.
- High-volume formats: look for overhead repeated on every line or item. Numbering every 10th line of a file read instead of every line cut cache-read tokens 1.6% without hurting citation accuracy. Each number costs 3–5 tokens, and agents read tens of thousands of lines per session. Also check repeated absolute paths, verbose JSON keys, ANSI codes, progress bars, and repeated headers.
- Good retrieval saves exploration turns. Adding semantic search alongside grep raised codebase question-answering accuracy 12.5% on average and cut the iterations users needed.
- Tool errors waste tokens and leave confusing debris in context. Classify expected errors (invalid arguments, unexpected environment, provider error, timeout, user abort), treat unknown errors as harness bugs, and track rates per tool and per model. One focused effort along these lines cut unexpected tool errors 10×.

## 6. Long runs: compaction, subagents, and model mix

- Compaction: keep the summarization prompt short and the summary compact, carry forward plan state and remaining tasks, and save the full history to a file the agent can search for details the summary dropped. A model trained to self-summarize from a one-line prompt wrote ~1k-token summaries with half the compaction error of a multi-thousand-token prompt that produced 5k+ token summaries. Untrained models may need more guidance, so test how short you can go. A more expensive summarization model made a negligible difference.
- Scratchpads and running notes: rewrite them instead of appending. For repeated work in one environment, a small agent-maintained notes file with a line budget, loaded at start, is a promising way to shorten later runs.
- Subagents: fresh context keeps the parent lean, but isolation adds coordination cost (duplicate or stale work). If the model already delegates on its own, remove prompting that pushes it to. Have subagents return short handoffs: what was done, findings, concerns, and deviations. A subagent should use a different model only when the user or harness says so.
- Model mix: in large multi-agent runs, workers used at least 69% of tokens, and over 90% in most runs. A frontier planner with cheap workers matched a frontier model doing everything at about one-eighth the cost. Planner choice still changes worker spend. One planner that cost less on its own saw its workers use several times more tokens, and the run cost more overall. Measure the whole tree.
- Routing and reasoning effort: send simple turns to a cheaper model or lower effort, and upgrade only when a stronger model is clearly better. A router built this way matched or beat single frontier models on user satisfaction at 41–68% lower cost.
- Reasoning continuity: if the API returns reasoning items (including encrypted ones), pass them back on later turns and alert when they go missing. Dropping them cost one reasoning model 30% on a coding benchmark, and it burned tokens reconstructing its plan.

## 7. Fit the harness to each model

Adapt to what each model was trained on instead of forcing one shape on all of them. If you've tuned the harness for a similar model, start from that version.

- Edit format: use the one the model was trained on (for example, patch-style or search-and-replace). An unfamiliar format costs extra reasoning tokens and causes more mistakes.
- Shell or tools: shell-first models fall back to `cat` or inline scripts. Name tools after their shell equivalents (such as `rg`), and if needed add: "If a tool exists for an action, prefer to use the tool instead of shell commands (e.g. read_file over `cat`)."
- Literalness: some model families follow instructions literally and others tolerate imprecision. Some spiral on emphasized wording. Strip caps and emphasis for literal models.
- Triggers: some models ignore a tool until told when to use it. A literal trigger works: "After substantive edits, use the <lint tool> to check recently edited files for linter errors. If you've introduced any, fix them if you can easily figure out how."
- Progress updates: if a model reports progress through reasoning summaries, keep them to 1–2 sentences that note new findings or a change of tactic, and remove instructions about messaging mid-turn.
- Quirks worth a targeted line: hedging or refusing as context fills ("context anxiety"), declaring completion early, stopping to ask permission, and calling tools that don't exist.

Tie each added instruction to the transcript behavior it fixes. Re-audit when models change, since guidance one version needed can be dead weight for the next.

## 8. Validate

- Offline: run a fixed set of realistic tasks before and after, ideally drawn from real usage and phrased the way users actually write (short and ambiguous). Compare task success, tokens, cost per task, turns, and tool errors. Don't ship a change that lowers success.
- Online, if you have users: A/B test each change or small bundle. The primary metric is cost per completed task. Guardrails are task success signals, tool-call errors, latency, turns per task, and cache hit rate. For a coding agent, a good success signal is how much agent-written code survives over time. In general, check whether the user's next message moves on or reports a problem.
- Ship only when cost drops and no guardrail regresses beyond noise. Record null results.

## What to change directly and what to propose

- Change directly, each in its own revertible commit: token and cache telemetry, deterministic serialization and tool order, moving volatile content out of the cached prefix, explicit cache breakpoints, writing large outputs to files instead of truncating, passing back reasoning items that are being dropped, and fixes for recurring tool errors.
- Change behind a flag so it can be tested: system prompt edits, tool offloading, output format changes, compaction changes, and subagent prompting.
- Propose only: changes to which models run, routing, reasoning-effort defaults, or how work is split across agents.

## Traps

- Asking the model to use fewer tokens or do less.
- Truncating tool output.
- Dropping reasoning items to save input tokens.
- Volatile content in the cached prefix, or tool order that changes between requests.
- Offloading a tool the model needs on the first turn or tries to call when it's missing.
- Emphasis-heavy prompts (MUST, NEVER, IMPORTANT, all caps), especially with literal models.
- Forcing a terser output format than the model was trained on. Fewer output tokens can mean less thinking and worse results.
- Optimizing raw token counts instead of cost, per request instead of per task, or evals instead of real usage.
- Switching models mid-conversation to save money.
- Adding coordination layers that become bottlenecks.

## Report back with

1. The harness map and baseline: cost by source × billing type, with the biggest sources called out.
2. A ranked list of changes: layer, what changes, estimated savings and how you estimated them, quality risk, how to validate, and how to roll back.
3. The changes you made, including a system prompt diff with a keep, rewrite, delete, or move reason for each line.
4. A test plan for the flagged changes.
5. Gaps: anything you couldn't find or measure.
♥ 3.3K · ⟲ 148 · 👁 353.1KView on X ↗

Sheel Mohnot Says Agents Threaten Expedia and Other OTAs

Sheel Mohnot argues that online travel agencies such as Expedia are especially exposed to agent disintermediation because they do not control inventory and their value lies in comparison and discovery. He responds to a post in which Muse found a direct hotel booking cheaper than OTA options.

Original post · 1 min read
Expedia / OTA's are particularly exposed to agent disintermediation

They don’t control the underlying inventory, the value is in comparison and discovery... but thats what agents do.
Vivek Goyal @Goyal_Vivek
Muse checked hotel website, Expedia, booking and found direct booking is 5-10% cheaper and has free cancellation and found me a $50 hotel credit.. @alexandr_wang
♥ 251 · ⟲ 18 · 👁 50.4KView on X ↗
Other3/10

Vijay Shekhar Sharma Shares Clip on British Slave Owner Payments and Indian Diaspora

Vijay Shekhar Sharma Shares Clip on British Slave Owner Payments and Indian Diaspora▶

Vijay Shekhar Sharma shares a video clip he calls fascinating, covering how Britain kept paying slave owners and how Indians came to be in Guyana, Jamaica and Mauritius. The post gives little detail beyond the clip itself.

Original post · 1 min read
You will find this clip super fascinating!
- how British kept paying slave owners till ….. and how come so many Indians in Guayana, Jamaica or Mauritius… and more
👑
♥ 931 · ⟲ 212 · 👁 70.4KView on X ↗

Anduril and Palmer Luckey Win Wildfire XPRIZE for Rapid Fire Detection

Palmer Luckey says Anduril won the Wildfire XPRIZE for detecting and suppressing a wildfire within 10 minutes of ignition. He argues that even small gains could save the U.S. roughly $500 billion a year lost to wildfires.

Original post · 1 min read
Thanks, Peter! This is the beginning of the end for destructive wildfires. In a country that loses ~$500B every year to wildfires, even tiny gains make for huge savings, to say nothing of the incalculable value of lives and homes.
Peter H. Diamandis, MD @PeterDiamandis
Congrats to @PalmerLuckey and Anduril for winning the Wildfire XPRIZE - for detecting and suppressing a wildfire within 10 minutes of ignition!
♥ 25.1K · ⟲ 1.7K · 👁 836.4KView on X ↗

Vercel Launches Drives for Sandbox in Public Beta

Guillermo Rauch argues successful AI agents need separated brain, hands and files components, and announces Vercel Drives, persistent storage for Vercel Sandbox now in public beta. Drives allow up to four mounts per sandbox and 16 TiB per Drive.

Original post · 1 min read
Muse, Instinct, OpenClaw, Claude Code…
All successful agents have 3 key components:

🧠 Brain → model, harness (logic)
👐 Hands → tools, computer, browser
🗃️ Files → memories, skills, repos

The 'easy' way is to throw all these in 1 stateful computer (a Mac Mini)

Like, you run 𝚌𝚕𝚊𝚞𝚍𝚎 or 𝚏𝚡 in your mac, you keep it running all day with 𝚌𝚊𝚏𝚏𝚎𝚒𝚗𝚊𝚝𝚎, it has storage, and CLIs and apps installed.

But if you want to cost-efficiently run agents in the cloud, you actually start breaking down these parts.

🧠 The harness can run in Fluid compute. To make it reliable across restarts, rollouts, crashes, you make its event log durable using Workflow.

👐 The hands can be a dedicated browser fleet like Browserbase/Kernel, a computer like Sandbox, and even more efficient lightweight tools like just-bash.

🗃️ 🆕 What was missing was a way to also decouple storage. Imagine you want to run a memory consolidation cron job every night ("dreaming"). You can read/write to the files directly without 'booting up' the agent's full computer.

Today we're introducing the perfect companion to Sandbox: Drives. We shipped the computer for agents, now we're giving you the 'external disk' you can attach at will. It's early, and we'll be expanding capabilities here quickly.

Btw, breaking apart the agent into these independent parts not only optimizes costs in a big way, it also *massively* improves security and auditability. I'd argue you can't even run a secure agent otherwise!
Vercel Developers @vercel_dev
Vercel Sandbox now has persistent storage with Drives, in public beta on every plan.

▪︎ Store agent workspaces, data, models, deps
▪︎ Read snapshots across parallel sandboxes
▪︎ Mount up to four Drives per sandbox
▪︎ Up to 16 TiB per Drive

vercel.com/changelog/drives-for-vercel-sandbox…
♥ 2.3K · ⟲ 138 · 👁 262.7KView on X ↗
AI5/10

Brad Gerstner Predicts Personal AI Assistants Will Reshape Internet Aggregators

Brad Gerstner says a personal AI assistant with memory of each user's preferences is inevitable and will disrupt aggregators in insurance, travel and shopping. He frames AI agents as shifting the fulcrum of value delivery.

Original post · 1 min read
As discussed w @bgurley two yrs ago - a personal, super assistant in every American’s pocket w perfect memory & understanding of my life - my likes & dislikes - able to work non stop & achieve almost anything is now inevitable. It is a huge gift to humanity. It will 10x all of us, give us more magic experiences & save time from drudgery to spend w family & friends. Many internet aggregators (insurance, travel, shopping) will resist but resistance is futile. AI agents aggregate the world of options on the fly - the fulcrum of value delivery has forever shifted. Model capability has unlocked this moment just as it did for coding. Adapt or die. 🤖🚀@BG2Pod
♥ 1.7K · ⟲ 95 · 👁 154.6KView on X ↗

Mark Suster Shares Jeffrey Katzenberg Essay on AI and Creativity

Mark Suster recommends an essay by Jeffrey Katzenberg on AI for creativity, which describes a founder's computer-animated scene and an artist's reaction. Suster presents it as a case for storytelling over AI slop and as an explanation of Jevons Paradox.

Original post · 1 min read
When people ask you to explain "Jevon's Paradox" send them this.

When people not as close to technical change as you are seem despondent, ask them to consider the arguments laid forward below for a better future

When people ask for proof of creativity, storytelling & soul over AI slop... show them the magical writing of @jeffreykWNDR
Jeffrey Katzenberg @jeffreykWNDR
The World is Changing: AI For Creativity

By Jeffrey Katzenberg

A few months ago, I sat in my office in Silicon Valley and watched as a tech founder showed me something extraordinary. On the screen was a fully realized, beautifully lit, well-composed animated scene. It was stunning and it made me feel exactly what I felt in 1986 watching Luxo Jr. That was the first time I watched a computer-animated 3D character take a breath and seem, against all reason, to have life. It left me in awe.

Later that day, I received a text from an artist I've known for thirty years, 350 miles to the south, in …
♥ 128 · ⟲ 8 · 👁 49.3KView on X ↗

Ben Thompson and Madhu Guru Debate How Consumers Want to Buy Through AI Agents

Ben Thompson says tech companies poorly craft marketing for agentic AI, while Madhu Guru argues consumers have differing relationships with tasks like shopping versus hiring a roofer. The exchange centers on which consumer activities AI agents should handle.

Original post · 1 min read
One of the all time insane clipping cycles. Few people are more agent-pilled than me. I’ve literally built my own. When Muse came out I said it was more important than solving math problems. And somehow, saying that tech companies don’t know how to craft marketing messages that resonate is being framed all over the Internet as being a Luddite.
Madhu Guru @realmadhuguru
I respectfully disagree with @benthompson on this. He’s missing a key point - Consumers don’t have one relationship with “doing things”.

I’ve learned it the hard way from building consumer and smb products for years at Google and now Meta.

Browsing clothes can be entertainment (for some). But hiring a roofer is miserable (for most).

Even with shopping, sometimes you want to browse for an hour. Sometimes you want the right thing just delivered asap.

His argument sounds a lot like the early 2000s:

“Why would I buy clothes online? I need to try them on.”

“Why would I buy a $500 TV online? I…
♥ 422 · ⟲ 11 · 👁 77.6KView on X ↗

Firecrawl Says Jev Analyzed Fortune 1000 Data in 3.6 Seconds

Firecrawl says its Jev tool made 5,000 decisions in 3.6 seconds using data from Firecrawl Alexandria. The quoted launch post says Alexandria pulls data on Fortune 1000 companies and offers 100+ data sources.

Original post · 1 min read
Jev made 5,000 decisions in 3.6 seconds using data from Alexandria!
Eric Ciarla (hiring) @ericciarla
Introducing Jev + Firecrawl Alexandria.

We pulled data on every Fortune 1000 company, then Jev analyzed it in 3.6 seconds.

With 100+ data sources, there's so much more to work with.

Alexandria is live now for anyone using Jev and beyond!
♥ 685 · ⟲ 38 · 👁 68.8KView on X ↗
AI6/10

Deedy Says Opus 5.5 Generates Startup Launch Videos for About $2

Deedy Says Opus 5.5 Generates Startup Launch Videos for About $2▶

Deedy reports that Opus 5.5 produced a launch video for an inference startup in about a minute for roughly $2. He argues instructional video generation will change marketing, sales and internal technical communication.

Original post · 1 min read
Opus 5.5 is incredible at instructional video generation.

I made this launch video for a inference startup in 1min for ~$2. Videos like these used to take weeks if not months and a lot of coordination with agencies and 1000x the costs.

Humans broadly prefer video to text. This changes the substrate of communication. These videos actually help communicate technical ideas in seconds (photorealistic video gen like Seedance is not very useful here).
- changes how often marketing should be talking about products and launches
- change how sales people can talk about technical products to their customers
- allow technical people to easily explain concepts internally without long docs
And thats just scratching the surface within startups.

Prompt: “make a modern slick and punchy video for a modern startup that works on inference”
♥ 3.3K · ⟲ 152 · 👁 347.8KView on X ↗

Austen Allred Reacts to Official Trailer for You Can See Everything

Austen Allred replies to a post by Nathan Fielder sharing the official trailer for the film You Can See Everything. The post is a short reaction with no further detail.

Original post · 1 min read
I have so many questions. This is so insane.
nathan fielder @nathanfielder
Here is official trailer, You Can See Everything
♥ 289 · ⟲ 2 · 👁 79.2KView on X ↗

Hiten Shah Highlights Muse Permission Controls for Network Protocols

Hiten Shah Highlights Muse Permission Controls for Network Protocols

Hiten Shah shows a permission screen from the consumer AI agent Muse that exposes controls for SSH, SMTP, DNS, TCP and other protocols. He notes users can block protocols or require approval per connection and is hosting an AI Permissions 101 session.

Original post · 1 min read
Muse is a consumer AI agent.

This is one of its permission screens.

It exposes controls for outbound SSH, SMTP, IMAP/POP3, database connections, FTP, DNS, TCP and UDP.

You can block a protocol entirely or let Muse ask before each connection.

Friday at 10 AM PT I’m doing AI Permissions 101.

hiten.com/ai-permissions-101
♥ 137 · ⟲ 9 · 👁 18.1KView on X ↗

Levelsio Reports Over $10M Annual Revenue and About 94.5% Profit Margin

Levelsio Reports Over $10M Annual Revenue and About 94.5% Profit Margin

Levelsio says his business passed $10M a year in revenue plus investment gains with roughly 94.5% profit margin, after replacing SaaS tools with vibecoded alternatives. He describes ETF, Nvidia and startup investments and stable monthly revenue of $200K to $250K.

Original post · 2 min read
💰 Passed $10M/y in revenue + investment gains this year

I've always been aggressively increasing my profit margins to as high as possible, especially the last few years, and especially this year where I've reduced my SaaS dependencies to just a handful of companies by vibecoding my own replacements

It's now about 94.5% profit!

Spend less -> lower costs -> more profit -> more money saved -> more money to invest -> more % returns on investments

And at some point that investment return starts to pass by your business revenue which it has for me for about a year now

Most of my investments have been simple ETFs like Vanguard S&P500, but also a few lucky stock and startup bets that worked out:

Back in 2022, I couldn't get any available GPUs, so I thought it might be smart to buy Nvidia if there's such a shortage and that was a good bet too

Then back in 2023 @dannypostma recommended me to invest in the companies we used for our AI inference for our photo AI apps, so I did that too. First I invested in @replicate and then also @FAL

From then I just asked every app/service I used if I could invest, like @cursor_ai in 2025 which then got acquired by @spacex recently!

My business revenue has remained very stable at about $200K-$250K/mo, but even with that investment returns have now surpassed that

Investments fluctuate a lot so some years it's up some years it might be down a lot (who isn't expecting a big crash at some point?), and most investment gains are unrealized (I never sell!), so keep that with a big grain of salt!

I keep my revenue/gains together with my profit margin and my sleep score in my situation monitor dashboard, as those are equally important :D

Many people in the last decade told me to spend all my money, either personally or with my business to grow it, we don't know how that would have ended up, maybe my companies would be bigger, but I'm happy I did it in my own lean profitable healthy way and it ended up great I think! 🙏😊
♥ 8.1K · ⟲ 163 · 👁 805.7KView on X ↗

Google Gemma Team Releases DiffusionGemma-Jev Endpoint on Cloud Run

Linus Ekenstam notes that frontier labs are rapidly releasing JEV forks, citing Google's Gemma team's djev. A linked Google post and GitHub repo describe deploying a Jev API-compatible endpoint on Cloud Run with one command.

Original post · 1 min read
I love seeing how fast the frontier labs are jumping in on creating forks.

Just today we’ve seen omni-jev and now djev from Google Gemma team.

It will be clear in a few weeks, just how powerful JEV is going to be inside harnesses and applications.
Google Gemma @googlegemma
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command.

Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec.

It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle.

Get the code and instructions here: github.com/taeold/djev-run
♥ 67 · ⟲ 7 · 👁 18.2KView on X ↗

Danny Postma Rebuilds Landing Page With Nine AI Agent Skills

Danny Postma Rebuilds Landing Page With Nine AI Agent Skills

Danny Postma says he rebuilt his landing page using AI agents by creating nine reusable skills rather than one-shot prompting. He reports a 34% higher conversion rate and is considering turning the skills into a course.

Original post · 1 min read
A few weeks ago I rebuilt my landing page with AI agents.

No one-shot. I created 9 skills instead from all my years of knowledge to speed up my time.

Test just finished w/ 34% higher conversion rate 🚀

Wondering if I should turn these skills into a course for your AI agents 🤔
♥ 449 · ⟲ 9 · 👁 30.0KView on X ↗
AI4/10

Bing Xu Argues Diffusion Gemma Is Undervalued Compared to Jev

Bing Xu says Jev is overhyped while Diffusion Gemma brings both intelligence and speed, calling it true System-1 behavior. He quotes a post reporting 2,000 to 4,000 tokens per second on Gemma via an AI Swarm inference engine.

Original post · 1 min read
I think Jev is overhyped. Diffusion Gemma is significantly undervalued because it brings both intelligence and speed. This is true System-1.
INT21 @int21_ai
Experience @googlegemma at ~2,000–4,000 tok/s with an AI Swarm produced inference engine. No speedup for video.
♥ 722 · ⟲ 40 · 👁 65.6KView on X ↗

Noah Smith Cites Columbia Data on Admissions Discrimination Against Indian Americans

Noah Smith states that Indians face more discrimination than other groups at elite American universities. He shares a quoted post citing Columbia admissions data showing South Asian applicants had markedly lower admission odds than similar white applicants.

Original post · 1 min read
Indians face far more discrimination than other groups at elite American universities.
Werner Zagrebbi @zagrebbi
I'll note again: the modern analog of the Jewish quota applies chiefly to Indian Americans.

In Columbia's admissions data, East Asians applicants scoring 1550 had 18% lower admission odds than similar whites. Similar South Asians, however, had odds 57% lower. twitter.com/AAGDhillon/status/2102470668243632…
♥ 1.4K · ⟲ 154 · 👁 60.2KView on X ↗

a16z Speedrun Talent Lead Reports More Than 100 Portfolio Placements

Jordan, who runs talent for a16z speedrun, says the role has resulted in over 100 hiring placements at portfolio companies this year at no cost to either side. He invites people to reach out to get on his radar.

Original post · 1 min read
i run talent for a16z speedrun - my whole job is getting people hired at our portfolio companies

this year that's been 100+ placements, free for both sides

takes 1 minute to get on my radar :)
♥ 355 · ⟲ 9 · 👁 594.5KView on X ↗

Tom Funds Entire Dev Team Before Landing First Customer on Website Widget

Tom Funds Entire Dev Team Before Landing First Customer on Website Widget▶

In a video, Tom, who reportedly earns $50K per month from a website widget, says he funded his dev team before having any revenue. He credits his marketing agency's roughly 70 hosting clients for financing the early build.

Original post · 1 min read
This dude Tom who makes $50K/month from a website widget says: he funded his entire dev team before he had a single customer...

"Yet we had zero revenue. So on paper, terrible idea. Don't recommend it at all. But with that gave us the confidence to sell into these bigger companies that we kind of needed to get on board to pay us."

"I did run a marketing agency for 10 years, which still had like 70 plus clients paying like website hosting. So I used my like very precious income that I did have to live to basically fund the start of this."
♥ 64 · ⟲ 5 · 👁 7.7KView on X ↗

Hamel Husain and Shreya Shankar Release Evals Skill for AI Coding Agents

Hamel Husain and Shreya Shankar Release Evals Skill for AI Coding Agents

Lenny Rachitsky recommends installing a new evals skill from Hamel Husain and Shreya Shankar that guides AI coding agents in building product-specific AI evals. The linked GitHub repo collects these skills, and his post cites examples of evals improving results at Ramp, Shopify, Harvey and Cursor.

Original post · 1 min read
Pro tip: Install this new evals skill from @HamelHusain and @sh_reya, it'll save you many hours and a lot of mistakes

github.com/ai-evals-course/evals-skills
Lenny Rachitsky @lennysan
Evals have been coming up more and more in my conversations with podcast guests and PM friends.

Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them:
— @tryramp took its automatic receipt collection from 35% to 83% accuracy.
— @Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced.
— @harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score.
— @cursor_ai tuned its Auto Balance routing, …
github.comGitHub - ai-evals-course/evals-skills: Skills that guide AI coding agents to help you build product-specific AI evals.Skills that guide AI coding agents to help you build product-specific AI evals. - ai-evals-course/evals-skills
♥ 1.4K · ⟲ 93 · 👁 230.1KView on X ↗
AI9/10

OpenAI Launches GPT-6 Sol and Luna With 50% Lower API Prices

OpenAI Launches GPT-6 Sol and Luna With 50% Lower API Prices▶

OpenAI Developers announced that GPT-6 Sol and Luna are launching today, with API prices 50% lower than GPT-5.6. The post positions Sol for building and Luna for scaling to production.

Original post · 1 min read
GPT-6 Sol and Luna just landed in Astra’s orbit.

Both launch today with API prices 50% lower than GPT-5.6.

Build with Sol. Scale with Luna. To production and beyond.
♥ 9.7K · ⟲ 752 · 👁 982.0KView on X ↗

Viral Post Contrasts Early Hopes With Current Results in Simple Comparison

Viral Post Contrasts Early Hopes With Current Results in Simple Comparison

Madeline Summerville shares a lighthearted photo comparing how something started with how it is going, quoting a post about New York City mayor Zohran Mamdani and a $130 million DoorDash wage settlement. The post offers little original content beyond the quoted news item.

Original post · 1 min read
how it started: how it’s going:
Democrats Deliver @DemzDeliver
🚨 BREAKING: Zohran Mamdani forces DoorDash to pay $130 million in stolen wages to NYC workers.
♥ 49.1K · ⟲ 5.1K · 👁 1.9MView on X ↗

X Launches Numbers, an Optional Way to Be Contacted Without Following

X Launches Numbers, an Optional Way to Be Contacted Without Following▶

XChat announces X Numbers, a feature letting users share a number so others can message or call them without following or accepting requests. The post is a short video announcement of the product.

Original post · 1 min read
X Numbers are here.

Share yours with anyone you want to contact you, even if you don't follow them. It's an optional way for people to message or call you, without you needing to accept requests or follow them back.
♥ 9.7K · ⟲ 1.2K · 👁 2.9MView on X ↗

Zach Klein Jokes That Daniel Lurie Tested the Waters First

Zach Klein Jokes That Daniel Lurie Tested the Waters First▶

Zach Klein posts a short video with a playful remark that Daniel Lurie went first and proved the water is warm, after which everyone followed in. The post is light commentary with no substantive detail.

Original post · 1 min read
Daniel Lurie jumped in first, proved the water is warm, now everyone is cannonballing into the deep end.
♥ 168 · ⟲ 1 · 👁 20.2KView on X ↗
AI9/10

Anthropic Introduces Claude Opus 5.5 in New Claude 5.5 Model Family

Anthropic Introduces Claude Opus 5.5 in New Claude 5.5 Model Family▶

Claude announces Claude Opus 5.5, the first model in its new Claude 5.5 family. Anthropic says it performs at the level of Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.

Original post · 1 min read
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.

It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
♥ 97.1K · ⟲ 9.0K · 👁 28.1MView on X ↗