Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Agents & Dev Tools

Coding agents, developer tools, workflows, open source

Developer Runs Codex Across Mac Mini and MacBook From Phone

Developer Runs Codex Across Mac Mini and MacBook From Phone

Nick describes running the Codex app on a always-connected Mac mini and a MacBook, linking both devices so threads can be started and resumed from either, with mutual SSH for file access.

Original post · 1 min read
My laptop has become a “satellite device” since I started using Codex from my phone. And my Mac mini has become the “home.” It’s clunky, but the end state feels more like how we’re going to be working in the near future:

I’m currently running the Codex app on 2 devices:
1. my MacBook
2. my Mac mini

My laptop isn’t reliably connected to Wi-Fi enough, so I keep a Mac mini on my desk that is always connected.

When I kick off new threads from my phone, I start them on the Mac mini. When I’m working from my desk, I run them there too.

The cool part is that I’ve added my MacBook and Mac mini as connected devices to each other. That means I can start and resume threads from either device. So if I’m in a meeting but want to continue a thread on my laptop that was started on my Mac mini, I can do that.

I’ve also set up mutual SSH for Mac mini <> MacBook, so files are easy to access from either side. It’s not fully seamless yet, but the model works.

What this means:

- I have an always-on Codex that is accessible from my phone, with its own dev environment
- All threads are always accessible from any of the 3 devices
- I can run heartbeat threads that stay on 24/7

It’s a little makeshift today, but the shape of it feels very real to me: Codex is no longer tied to whichever computer happens to be open in front of me. It starts to feel like something I can stay connected to across whatever device I’m using.
♥ 1.8K · ⟲ 106 · 👁 424.8KView on X ↗

Peter Steinberger Uses Codex in Ephemeral Crabbox Sandboxes for Bugs

Peter Steinberger Uses Codex in Ephemeral Crabbox Sandboxes for Bugs

Peter Steinberger says he has Codex recreate bug states in disposable Crabbox environments to reproduce, fix and verify issues, running about ten sessions in parallel. He links to Crabbox, a CLI for running repository commands in local sandboxes, cloud VMs and hosted agent sandboxes.

Original post · 1 min read
Whenever I investigate a bug, I let codex recreate the exact state in an emphemeral crabbox, verify the bug, fix it, verify the fix.

No messy state because local system might be polluted, and no slowdown because I run 10 sessions in parallel. crabbox.sh/
crabbox.shCrabbox — Run Any Repository Command in the Right BoxRun repository commands in local sandboxes, cloud VMs, SSH hosts, Windows and WSL2, macOS, or hosted agent sandboxes through one CLI.
♥ 3.0K · ⟲ 152 · 👁 322.0KView on X ↗

Ten Open-Source GitHub Repos Pitched as Business Opportunities

Ten Open-Source GitHub Repos Pitched as Business Opportunities

Nav Toor lists ten open-source projects, including Cal.com, Plausible, Ghost, n8n, Supabase, Medusa, AppFlowy, Coolify, Listmonk and Penpot, and suggests reselling or self-hosting each as a business. Several claimed revenue and funding figures are presented without sourcing, and the post links to the repositories.

Original post · 2 min read
Here are 10 GitHub repos that quietly print money while you sleep.

1. Cal. com
Open-source Calendly. Fork it, white-label it, sell to dentists and lawyers for $200/month. The founders hit $5M ARR in 3 years doing exactly this.
Repo → github.com/calcom/cal.com

2. Plausible Analytics
Privacy-first Google Analytics. Self-host it, resell to agencies for $50/month per client. Two founders bootstrapped this to 7 figures.
Repo → github.com/plausible/analytics

3. Ghost
Open-source Substack with 100% margin. 1,000 readers at $5/month equals $60,000 a year. Forever.
Repo → github.com/TryGhost/Ghost

4. n8n
Open-source Zapier. Sell automation services for $500-$2,000 per setup. n8n raised $14M because the agency model behind it works.
Repo → github.com/n8n-io/n8n

5. Supabase
Free Firebase replacement. Build a SaaS in a weekend, charge $29-$99/month. They raised $116M for a reason.
Repo → github.com/supabase/supabase

6. Medusa
Open-source Shopify. Take 5% on every sale forever. Zero rev share to Shopify.
Repo → github.com/medusajs/medusa

7. AppFlowy
Open-source Notion. Sell self-hosted to enterprises worried about data privacy. They raised $30M because this market is massive.
Repo → github.com/AppFlowy-IO/AppFlowy

8. Coolify
Open-source Vercel and Heroku. Charge developers $20/month to manage their deployments. Replace their $200 Vercel bill.
Repo → github.com/coollabsio/coolify

9. Listmonk
Open-source Mailchimp. Send unlimited emails for the cost of an AWS bill. Resell to agencies at 10x markup.
Repo → github.com/knadh/listmonk

10. Penpot
Open-source Figma. Sell self-hosted design tools to agencies who refuse to upload client files to the cloud.
Repo → github.com/penpot/penpot

The difference between developers who build features and developers who build businesses is one decision.

Pick one of these. Fork it this weekend. Ship it next week.

The founders behind these repos already proved the model.

Save this. Share it with the developer in your life who deserves to break free.

100% free. 100% open source.
github.comGitHub - calcom/cal.diy: Scheduling infrastructure for absolutely everyone.Scheduling infrastructure for absolutely everyone. - calcom/cal.diy
♥ 6.0K · ⟲ 769 · 👁 450.8KView on X ↗

Printing Press Launches Agent-Native CLI Library and Generator Factory

Printing Press Launches Agent-Native CLI Library and Generator Factory▶

Matt Van Horn, with Trevin, introduces Printing Press, a library of agent-optimized command-line tools for services like Linear, ESPN and Google Flights, plus a generator that creates new CLIs for any product via a slash command. The tools are SQLite-backed and work with Claude Code, Codex, OpenClaw and Hermes.

Original post · 1 min read
Introducing the Printing Press, a CLI-factory and a CLI-library. Built with @trevin. 🏭🖨📚

Most APIs suck for agents. Most MCPs suck for agents. Most official CLIs suck for agents. They waste tokens and time. @steipete started making his own because of this.

📚 A Library of agent-native CLIs you install today (Linear, ESPN, Flight GOAT (Google Flights + Kayak nonstop), Contact Goat (LinkedIn + Happenstance + Deepline more) +30+ more)
🏭 A factory that prints new ones for any service - just type /printing-press <product name>

CLIs are fast, local, SQLite-backed. Work in Claude Code, Codex, OpenClaw, Hermes.

🌐 printingpress.dev
♥ 3.6K · ⟲ 280 · 👁 1.4MView on X ↗

Mobbin MCP Connects 600,000 App Screens to Claude and Cursor

Mobbin MCP Connects 600,000 App Screens to Claude and Cursor▶

Ihor describes an MCP from Mobbin that gives Claude and Cursor access to 600,000 screens from real apps. He demonstrates pulling 43 paywall examples from apps like Revolut, Uber, and Duolingo to inform builds.

Original post · 1 min read
first MCP that actually changed how i work, not just another connector

@mobbin plugged 600k screens from real apps into claude/cursor. ask for a paywall — it pulls 43 examples from revolut, uber, duolingo and builds from what works, not from imagination
♥ 1.1K · ⟲ 40 · 👁 119.0KView on X ↗

Open-Source Project Make-Comics Generates Full Comic Books From Prompts

Open-Source Project Make-Comics Generates Full Comic Books From Prompts

Tom Dörr shares a GitHub repository, Nutlope/make-comics, that creates complete comic books with AI from a single prompt. The post links to the repo and includes an image.

Original post · 1 min read
Generates full comic books from a single prompt

github.com/Nutlope/make-comics
github.comGitHub - Nutlope/make-comics: Create full comics with AI in secondsCreate full comics with AI in seconds. Contribute to Nutlope/make-comics development by creating an account on GitHub.
♥ 361 · ⟲ 39 · 👁 12.1KView on X ↗

Aaron Levie Argues Headless Software Is the Future

Aaron Levie posts that headless software is the future, quoting a Grok announcement that its assistant now connects to Microsoft Teams, Salesforce and Box through three new connectors. The post is brief and offers no further explanation of the thesis.

Original post · 1 min read
Headless software is the future
Grok @grok
Grok now works where you work.

Message your coworkers in Microsoft Teams, manage customers in Salesforce, pull up files in Box. Three new connectors are now live.
♥ 325 · ⟲ 20 · 👁 79.8KView on X ↗

Vercel Launches Open Source deepsec Coding Security Scanner

Vercel Developers announces deepsec, an open source, CLI-first security harness that uses pluggable coding agents and sandboxed scaling to find and fix vulnerabilities in large repositories. It can run on Vercel's AI Gateway or a user's own subscription.

Original post · 1 min read
Introducing deepsec, an open source coding security harness.

• CLI-first
• Sandbox-based scaling
• Pluggable coding agents
• Designed for large-scale repos
• Use AI Gateway or your own subscription

After months of successful internal use, we put it to the test on some of the largest open source codebases.
vercel.com/blog/introducing-deepsec-find-and-f…
♥ 2.6K · ⟲ 268 · 👁 934.7KView on X ↗

Open-Source Bot Automates Trading on Polymarket BTC Markets

Open-Source Bot Automates Trading on Polymarket BTC Markets

Tom Dörr shares a GitHub repository for an algorithmic trading bot that automates 15-minute Bitcoin price prediction trades on Polymarket, combining multiple signal sources and risk management.

Original post · 1 min read
Automates 15-minute BTC trades on Polymarket

github.com/aulekator/Polymarket-BTC-15-Minute-…
github.comGitHub - aulekator/Polymarket-BTC-15-Minute-Trading-Bot: A production-grade algorithmic trading bot for Polymarket's 15-minute BTC price prediction markets. Built with a 7-phase architecture combining multiple signal sources, professional risk management, and self-learning capabilities.A production-grade algorithmic trading bot for Polymarket's 15-minute BTC price prediction markets. Built with a 7-phase architecture combining multiple signal
♥ 450 · ⟲ 60 · 👁 30.2KView on X ↗

Developer Tracks 430 Hours of Claude Code Use, Finds 73% Waste

I tracked 430 hours of Claude Code usage. 73% was wasted on these 9 patterns.

Mnimiy reports logging 430 hours and $1,340 in API spend across 90 days of Claude Code sessions, concluding that 73% of tokens went to nine overhead patterns. The article follows Anthropic's acknowledged usage-limit problems and offers a fix for each pattern.

Original post · 12 min read
X ArticleI tracked 430 hours of Claude Code usage. 73% was wasted on these 9 patterns.
For 90 days I logged every single Claude Code session: timestamp, prompt, response, token count, model, exit reason.
430 hours of work. 6 million input tokens. $1,340 in API spend.
Then I sat down with the data and asked the only question that mattered: how much of this was actually doing my work, and how much was overhead?
The answer was uncomfortable: 73% of my tokens went to nine invisible patterns that I'd been doing on autopilot.
Not bad prompts, I write decent prompts.
Not big models when I needed small ones, I knew that one already.
Patterns deeper than that. Patterns nobody talks about because they're invisible until you instrument them.
If you're hitting Claude Max usage limits more than once a week, you have at least 4 of these. Probably 7.
Below: each pattern, exactly how much it costs, and the 30-second fix.

Why this matters now
Anthropic admitted the problem in late March 2026: "people are hitting usage limits in Claude Code way faster than expected."
Max 5 subscribers reported quota exhaustion in 19 minutes instead of the expected 5 hours.
The Pro $200/year users said the limit maxed out every Monday and didn't reset until Saturday - "out of 30 days I get to use Claude 12."
Anthropic acknowledged the issue. The peak-hours quota change in late March explained part of it. The rest was a prompt-caching bug — two independent bugs in the cache layer that silently inflated costs by 10-20x for some sessions.
A user reverse-engineered the Claude Code binary to find it (GitHub issue #40524). Downgrading to v2.1.34 helped some people.
Some of it is real. Most of it is you.
I'm not going to tell you to "use Haiku for simple tasks" or "start a new chat every 15 messages."
Below: the patterns that the obvious advice misses.
The methodology
I ran an HTTP proxy between Claude Code and the Anthropic API. Logged every request: full payload, response, token counts (input/output/cache), latency, model. 90 days. 430 hours of active work.
For each request, I categorized the tokens into:
Productive — content that directly informed my actual question
Cache hit (free) — system prompt + CLAUDE.md cached, no marginal cost
Cache miss (paid) — same content, recomputed because cache expired
Conversation history re-read — re-tokenizing previous messages on every turn
Hook injection — pre-pended context from PreToolUse / UserPromptSubmit hooks
Skill loading — skill SKILL.md content loaded into context for invocations
Tool use overhead — JSON schemas, tool definitions, tool result blocks
Extended thinking — <thinking> blocks
CLAUDE.md — project rules loaded every turn

Productive tokens: 27%. The other 73% is the 9 patterns below, ranked by how much they cost me.
Pattern 1. CLAUDE.md bloat (~14% of total tokens)
The pattern: my CLAUDE.md grew to 4,800 tokens over 6 months. Every turn loaded all 4,800 tokens. Every session loaded them again on each new request. Most of the rules were never relevant to the task at hand.
A 5,000-token CLAUDE.md costs you 5,000 tokens before you've typed a word. Every turn. Every session. A constant baseline tax. Multiply by 200 turns per week and it's 1 million tokens a week of your CLAUDE.md alone.
The 30-second fix:

If you're over 1,500 tokens combined, refactor:
Move framework-specific rules to project-level CLAUDE.md (only loads in that project)
Extract repeated patterns into skills (loaded only when invoked)
Delete anything you can't remember writing
Convert "explain why" verbose rules into 3-word imperatives
I cut mine from 4,800 to 900 tokens. Same behavior. 31% reduction in baseline cost, instantly.
Pattern 2. Conversation history re-reads (~13% of total tokens)
The pattern: every follow-up message re-tokenizes the entire conversation history. By message 30 in a chat, each turn is paying for messages 1-29 to be read again. The math: at ~500 tokens per exchange, message 30 costs 30× message 1.
I had sessions with 60+ messages. The last message was costing 60× the first. Tokens spent re-reading old context: catastrophic.

The 30-second fix:
Edit the prior message instead of follow-up. Up-arrow -> edit -> re-send. The bad exchange gets replaced, not stacked.
Hard cap conversations at 20 messages. When you cross 20, ask Claude to summarize what's been done and start a fresh chat with that summary as the first message.
Use /compact instead of /clear when you need continuity. /compact summarizes and restarts. /clear nukes everything.
I went from 60-message sessions to 15-message average. 40% drop in conversation re-read cost.
Pattern 3. Hook injection waste (~11% of total tokens)
The pattern: I had 4 plugins installed. Three of them registered UserPromptSubmit hooks that injected context. Combined: 6,200 tokens of hook injection on every prompt I submitted, before Claude even read what I asked.
These hooks are designed to be helpful. They inject branch names, recent file changes, instinct summaries, memory snippets. Each one is small. Together they're a wall.
The 30-second fix:

Aud… continue on X ↗
♥ 955 · ⟲ 120 · 👁 1.6MView on X ↗

Cursor Releases SDK for Building Agents on Its Own Runtime

Cursor Releases SDK for Building Agents on Its Own Runtime▶

Cursor announces the Cursor SDK, which lets developers build agents using the same runtime, harness and models that power Cursor. Agents can run in CI/CD pipelines, power automations or be embedded in products, shown in a demo video.

Original post · 1 min read
We’re introducing the Cursor SDK so you can build agents with the same runtime, harness, and models that power Cursor.

Run agents from CI/CD pipelines, create automations for end-to-end workflows, or embed agents directly inside your products.
♥ 8.7K · ⟲ 803 · 👁 3.1MView on X ↗

Claude Code's Head of Product Explains Anthropic's Faster Shipping

Lenny Rachitsky summarizes an interview with Cat Wu, Head of Product for Claude Code at Anthropic, covering shorter product cycles, the merging of PM and engineering roles, building ahead of model capability, and using model introspection.

Original post · 5 min read
My biggest takeaways from Claude Code's Head of Product @_catwu:

1. Anthropic’s product development timelines have gone from six months to one month, sometimes one week, sometimes one day. Part of this acceleration is access to the latest models (i.e. Mythos). Another is shipping new products into “research preview,” making clear it's early, experimental, and might not be supported forever. Another is an evergreen "launch room "where engineers post ready features and marketing turns around announcements the next day.

2. The PM role is shifting from coordinating multi-month roadmaps to enabling teams to ship daily. As Cat puts it, “There should be less emphasis on making sure you are aligning your multi-quarter roadmaps with your partner teams and more emphasis on, OK, how can we figure out the fastest way to get something out the door?”

3. The most efficient shipping unit is an engineer with great product taste. On Cat’s team, many engineers go end-to-end—from seeing user feedback on Twitter to shipping a product by the end of the week—without a PM involved. Also, almost all the PMs on the Claude Code team have either been engineers or ship code themselves, and the designers have been front-end engineers. The roles are merging, and the most valuable skill is product taste, not job title.

4. Build products that are on the edge of working. Claude Code’s code review product failed multiple times because earlier models weren’t accurate enough. But because the prototype was already built, they could swap in Opus 4.5 and 4.6 and immediately test whether the gap was closed. Teams that wait for the model to be ready will always be a cycle behind.

5. The most underrated skill for building AI products is asking the model to introspect on its own mistakes. Cat regularly asks the model why it made an unexpected decision. The model will explain that something in the system prompt was confusing, or that it delegated verification to a subagent that didn’t check its work. This reveals what misled the model so the team can fix the harness.

6. Every model release forces their team to revisit existing products and audit their system prompt to remove features the model no longer needs. Claude Code’s to-do list was a crutch for earlier models that couldn’t track their own work. With Opus 4, the model handles it natively. Features built as scaffolding for weaker models become debt when the model catches up—so the team actively strips them.

7. Anthropic employees build custom internal tools instead of buying SaaS products. A sales team member built a web app that pulls from Salesforce, Gong, and call notes to auto-customize pitch decks—work that used to take 20 to 30 minutes now takes seconds. Their core stack is Claude Code, Cowork, and Slack. No Notion, no Linear, no Figma.

8. People underestimate how much Claude’s personality contributes to its success. As Cat describes it, “When you reflect on everyone you’ve worked with, there’s just some people where you’re like, I really like their energy, their vibe.” Claude is designed to be low-ego, positive, competent, and earnest—qualities that make it feel like a great coworker, not just a tool. This isn’t cosmetic; it’s what makes people want to use Claude for hours every day. The team has a dedicated person, Amanda, who “molds Claude’s character,” and it’s one of the hardest roles at the company because success is so subjective.

9. The future of work is managing fleets of AI agents, not doing the work yourself. Cat sees a clear progression: first, individual tasks become successful. Then people start running multiple tasks at the same time (multi-Clauding). Next, people will run 50 or 100 tasks simultaneously, which will require new infrastructure—remote execution, better interfaces for managing tasks, agents that fully verify their work, and self-improving systems that incorporate feedback. The human role shifts from doing the work to knowing which tasks to look into, verifying outputs, and giving feedback that makes the system better over time.

10. Hire people who lean into chaos and face every challenge with a smile. At Anthropic, there are weeks when a P0 on Sunday becomes a P00 by Monday and a P000 by Monday afternoon. If you get too stressed about any one thing, you’ll burn out. Their team looks for people who can look at a hard challenge and say, “Wow, that’s gonna be hard. But I’m excited to tackle it and I’m gonna do the best that I possibly can.” This mindset—optimism, resilience, and comfort with constant change—is increasingly essential as the pace of AI development accelerates.

Don't miss the full conversation: youtube.com/watch?v=PplmzlgE0kg
Lenny Rachitsky @lennysan
How Anthropic’s product team moves faster than anyone else

I sat down with @_catwu, Head of Product for Claude Code at @AnthropicAI, to get a peek into their unprecedented shipping pace, how AI is changing the PM role, and how to be the right amount of AGI-pilled.

We discuss:
🔸 How Anthropic’s shipping cadence went from months to weeks to days
🔸 The emerging skills PMs need to develop right now
🔸 Why you should build products that don't work yet—then wait for the model to catch up
🔸 Why a 95% automation isn't really an automation
🔸 Cat’s most underrated AI skill (introspection)
🔸 What …
♥ 2.8K · ⟲ 293 · 👁 847.8KView on X ↗

Academic Research Skills Package Covers Full Research Pipeline for Claude Code

Academic Research Skills Package Covers Full Research Pipeline for Claude Code

Tom Dörr shares a GitHub repository called academic-research-skills, a set of Claude Code skills covering research, writing, review, revision and finalization of academic work.

Original post · 1 min read
Covers full academic research pipeline for Claude Code

github.com/Imbad0202/academic-research-skills
github.comGitHub - Imbad0202/academic-research-skills: Academic Research Skills for Claude Code: research → write → review → revise → finalizeAcademic Research Skills for Claude Code: research → write → review → revise → finalize - Imbad0202/academic-research-skills
♥ 2.0K · ⟲ 255 · 👁 142.0KView on X ↗

Developer Builds Offline AI Skin Journal Running on iPhone

Locally This, Locally That

Aman describes Ambrosia, an offline AI journaling app for tracking skin responses, built around a fine-tuned Qwen3-VL-2B model that runs entirely on device. He covers local finetuning on a MacBook, LoRA training on MedQA, and quantized GGUF export.

Original post · 4 min read
X ArticleLocally This, Locally That
I built an AI app that runs completely on device. No servers, API requests, or data leaving the device. Local models are finally getting good. Just look at the recent releases from @Alibaba_Qwen and @googlegemma to see that the gap is shrinking.
x.com/Alibaba_Qwen/status/2046939764428009914
I am excited about what this unlocks for hardware, consumer, and privacy.
Robots and edge devices can't rely on perfect signal for their user experience. A robot working in a farm, a construction site, or even a living room can't just freeze up the moment it loses connection. Smart glasses and similar devices running local models have lower latency and much better UX.
For consumer apps, local LLMs make free tiers economically viable. Growth at all costs becomes very expensive when every user burns GPUs on your dime. Local inference flips that math.
@signulll is clearly facing this now:
x.com/signulll/status/2044097057825124430
The third vector is privacy. I have mixed feelings here. Anecdotally, most of my friends don't care where their data goes. My guess is, in the future, privacy-sensitive fields like law and medicine will require on-device or on-prem models to stay compliant. Users won't care, but regulated professionals will.
The hardware is already here. We're all walking around with phones that can comfortably run 2B-parameter models. So I wanted to see how far I could push that with a real app.
Ambrosia
I built Ambrosia, an offline AI journaling app to track how my skin responds to different diets and products. I picked skin specifically for two reasons: I've always wanted a better way to watch conditions and progress over time, and photos are a natural unit of a journal entry.
Two things I love about the app:
1. AI-generated labels and trend tracking, all local
2. Everything including the images stays on device
The Base Model
I went with Qwen3-VL-2B: multimodal, small enough to run on my iPhone, and already well-optimized for llama.cpp. I used the Q4_K_M quantized version from @huggingface, and ran everything through @RunAnywhereAI because their on-device SDKs are the best I've used.
Finetuning the Model Locally
I wanted to prove that I could take a base model and finetune it end-to-end locally on my MacBook (M2 Max).
With Codex, I trained a small LoRA adapter on MedQA, then merged and exported it back into the same Q4_K_M GGUF format so that it would work on my iPhone.
I also experimented with autoresearch (@karpathy’s autonomous framework for model training) to maximize performance. I gave Codex the source code and let it sweep configurations. The sweep landed on a lower learning rate (1e-4) as the winner, beating baseline by ~5 points on the micro-run. Scaled up to full validation, it held at 47.88%.

Finetuned Model: huggingface.co/amankishore/qwen3-vl-2b-medqa-gguf
Constraining the Task
Initially I built Ambrosia as a chat first experience. It was bad. I would document a new skin issue, and the model would respond with three long paragraphs about consulting a dermatologist.
I constrained the task. Qwen was much better at generating labels from images and analyzing trends across journal entries. Ambrosia went from a worse medical ChatGPT to a journal with ambient AI trend analysis.
This is why small models are underrated. They should not be compared to general assistants like ChatGPT. They're much better as narrow specialists constrained to a few tasks.
Since this task is disjoint from MedQA, I built a small eval set of journal entries a user might track (acne, texture, redness, etc) and compared the finetuned model against the base. The finetune gave a small but real lift on label quality, while trend analysis was unchanged. For a 2B model doing a task it wasn't trained for, I'll take it.

Edge Intelligence
Ambrosia was a small experiment, but it maps to the shape of the future: take a small base model, aggressively constrain the domain, tune for your specific task, ship it. Repeat.
Chips keep getting faster. Small models keep getting smarter. Eventually these lines will converge. When they do, you’ll have the power of AGI, in the palm of your hand.
♥ 913 · ⟲ 52 · 👁 798.2KView on X ↗

Heycounsel Hosts Claude Cowork Hackathon for Lawyers Building Legal Skills

Heycounsel Hosts Claude Cowork Hackathon for Lawyers Building Legal Skills▶

Brian Scherer reports a five-hour hackathon on April 21 in which more than 200 lawyers across 10 countries built production-ready legal AI skills with Claude Cowork, producing 13 completed projects. The skills are published on the Heycounsel showcase site.

Original post · 1 min read
On April 21, we challenged lawyers from all over the world to spend an afternoon building real, production-ready skills with Claude Cowork.

5 hours. 200+ lawyers. 10+ countries. 13 completed projects producing new legal AI skills.

Every skill is live on our showcase site at go.heycounsel.com/hackathons
♥ 461 · ⟲ 38 · 👁 60.2KView on X ↗

Greg Brockman Recommends 28-Minute Tutorial on OpenAI's Codex

OpenAI president Greg Brockman endorses a tutorial by Riley Brown covering seven capabilities of Codex, including file access, persistent memory, plugins, skills, image generation, computer use and automations.

Original post · 1 min read
a great codex tutorial:
Riley Brown @rileybrown
Learn 95% of Codex in 28 minutes

These are the 7 knowledge work capabilities...
inside Codex, the super-app

00:00 Intro
02:19 Capability 1 - Full File Access
07:41 Capability 2 - Persistent Memory
10:46 Capability 3 - Plugins
13:52 Capability 4 - Skills
19:22 Capability 5 - GPT Image Access
21:03 Capability 6 - Browser and Computer Use
23:58 Capability 7 - Automations
25:31 Bonus Feature - Chronicle
27:21 Summary
♥ 3.3K · ⟲ 289 · 👁 495.4KView on X ↗

Steve Yegge Says Google Has Two-Tier System for Claude Access

Steve Yegge follows up on his earlier tweet about Google's AI adoption, citing anonymous Googlers who describe DeepMind engineers using Claude daily while most other teams are pushed onto internal Gemini variants. He says he has not verified each account.

Original post · 3 min read
My tweet last week about Google's AI adoption drew a lot of pushback, to say the least.

Since then, Googlers from multiple orgs have reached out to me independently and anonymously. They've expressed fear of being doxxed, concern about what they saw as bullying of me, and general corroboration of my original tweet. I haven't verified each person's story, but the picture these Googlers paint is consistent across sources. It is more specific than what I originally wrote, and somewhat bleaker.

What they describe is a two-tier system. DeepMind engineers use Claude as a daily tool. Most of the rest of Google does not. When the question of equalizing access came up internally, the proposed response was to remove Claude for everyone — which DeepMind objected to so strongly that several engineers reportedly threatened to leave.

Non-DeepMind engineers get pushed onto internal Gemini variants behind router-style names that obscure which underlying model is actually serving a request. Multiple engineers describe regressions and reliability problems severe enough that some senior people have stopped using the tools. A senior manager on a major product line reportedly flagged attrition concerns over exactly this issue.

Googlers say leadership knows the gap is real. The response has been to mandate AI usage in OKRs and individual expectations, and to stand up an internal token-usage leaderboard. Unfortunately, managers have been told both that the leaderboard won't be used for performance reviews and, separately, that it absolutely will. And I hear other stories that Google's culture is not adapted properly yet for high-volume coding.

Addy Osmani's reply on behalf of Google said over 40,000 SWEs use agentic coding weekly. I don't doubt the number. But weekly use of a thin tool is precisely the box-checking I described in the original post. Volume of opens isn't adoption — and "weekly" is a low bar that includes a lot of people who tried it once and went back to writing code by hand.

The clearest thing I'm hearing is that Googlers do want to use high-quality agentic tools. They are asking repeatedly for better ones. But overall, this is not a picture of an engineering org that is fine.

My goal in the first tweet, and now, is always the same — get more people using AI and agentic coding. Nobody is as far ahead as they might look from the outside, and none of you are as far behind as you might be worried you are.

To all the Googlers who've reached out: thank you. You took a real risk and I appreciate you. Be safe. And good luck getting good models!
♥ 3.1K · ⟲ 203 · 👁 1.0MView on X ↗

Guide Explains How to Triage Compromised Google Workspace OAuth App

Omar shares steps for Google Workspace admins to check for a compromised third-party OAuth app tied to the Vercel incident. The instructions cover navigating admin API controls and revoking access by client ID.

Original post · 1 min read
Here's how to triage:

1. Go to admin.google.com

2. Security → Access and data control → API controls → App access control → Manage Third-Party App Access

3. Search for client ID:
110671459871-30f1spbu0hptbs60cb4vsmv79i7bbvqj

if found → revoke / block
Vercel @vercel
Our investigation has revealed that the incident originated from a third-party AI tool with hundreds of users whose Google Workspace OAuth app was compromised.

We recommend that Google Workspace Administrators check for usage of this app immediately. vercel.com/kb/bulletin/vercel-april-2026-secur…
♥ 2.2K · ⟲ 275 · 👁 571.3KView on X ↗

Louise de Sadeleer Says Claude Can Replace Opus Clip for Short Videos

Louise de Sadeleer Says Claude Can Replace Opus Clip for Short Videos▶

Louise de Sadeleer argues users can cancel their OpusClip subscription and instead create 9:16 short-form clips using Anthropic's Claude, sharing a short demo video.

Original post · 1 min read
It's time to cancel your @OpusClip subscription. Use @claudeai to create 9:16 clips instead.

Here's the super fast demo, while I get ready 💋
Louise de Sadeleer @LouiseDSadeleer
So who's gonna tell Opus Clips users they can create clips in Claude Code?
♥ 2.2K · ⟲ 99 · 👁 352.9KView on X ↗

Google Open-Sources osv-scanner for Dependency Vulnerability Checks

A post introduces osv-scanner, Google's open-source tool that scans lockfiles, containers and vendored code against the osv.dev vulnerability database. It highlights guided remediation, call analysis, support for 11+ ecosystems, and offline scanning.

Original post · 1 min read
GOOGLE BUILT A VULNERABILITY SCANNER AND OPEN-SOURCED IT

most devs ship code without knowing half their dependencies are ticking time bombs

osv-scanner fixes that

it scans your entire project lockfiles, containers, even vendored c/c++ code and maps every dependency against the osv.dev database

supports 11+ ecosystems. npm, pip, cargo, maven, go modules, gem. all of it.

the guided remediation feature is the real unlock... it doesn't just tell you what's broken.... it tells you exactly which version upgrades fix the most issues with the least risk

call analysis built in. so you only get alerts for vulnerable functions your code actually calls. no noise

works offline too. download the db once, scan without internet

one command to scan your whole directory:
osv-scanner scan source -r ./

github.com/google/osv-scanner
♥ 1.2K · ⟲ 187 · 👁 128.5KView on X ↗

Guide Offers Ways to Stretch Claude Code Usage Limits

Aakash Gupta promotes a guide to getting more out of Claude Code without hitting usage limits, quoting a post by Pawel Huryn that identifies four root causes of harness problems and provides copy-paste templates.

Original post · 1 min read
If you want to get more out of Clade without hitting limits, read this.
Paweł Huryn @PawelHuryn
Claude Code's Limits Are Generous. The Problem Is Your Harness. — Same PM workflow as my $750/mo era. Anthropic fixed 3 bugs. The 4 root causes still on your side, with copy-paste templates.

I'm on Claude Code Max (20x). 3 days in. 12% used. Same workflow that
♥ 175 · ⟲ 20 · 👁 68.3KView on X ↗

Steven Tey Urges Admins to Restrict Unconfigured Google OAuth Apps

Steven Tey Urges Admins to Restrict Unconfigured Google OAuth Apps

Steven Tey warns that third-party Google OAuth apps requesting scopes beyond basic profile data are a dangerous attack vector. He recommends Workspace admins restrict unconfigured third-party apps, linking to the Google admin settings page and crediting a tip.

Original post · 1 min read
Biggest takeaway from this: 3rd-party Google OAuth Apps that request scopes beyond the basic info (name/user/profile pic) is a dangerous attack vector.

To safeguard your org from attacks like this, highly recommend asking your Google workspace admin to restrict "unconfigured third-party apps" to only be able to request basic info needed 👇

Here's the direct link to access that settings page: admin.google.com/ac/owl/settings

h/t @matid for the pro-tip!
Guillermo Rauch @rauchg
Here's my update to the broader community about the ongoing incident investigation. I want to give you the rundown of the situation directly.

A Vercel employee got compromised via the breach of an AI platform customer called Context.ai that he was using. The details are being fully investigated.

Through a series of maneuvers that escalated from our colleague’s compromised Vercel Google Workspace account, the attacker got further access to Vercel environments.

Vercel stores all customer environment variables fully encrypted at rest. We have numerous defense-in-depth mechanisms to prot…
♥ 1.7K · ⟲ 201 · 👁 331.6KView on X ↗

Free-Claude-Code Proxy Runs Claude Code on NVIDIA Free Tier

Free-Claude-Code Proxy Runs Claude Code on NVIDIA Free Tier

Hasan Toor promotes free-claude-code, an open-source proxy that routes Claude Code to NVIDIA NIM models using a free API key. It supports Kimi K2, GLM 4.7, MiniMax M2 and Devstral, and includes a Telegram bot for remote control.

Original post · 1 min read
Goodbye Claude Code subscription fees.

Someone just built a proxy that runs Claude Code completely free... and it's wild.

You literally plug in a free NVIDIA API key and point Claude Code at localhost.

That's it.

It handles everything:
- Converts Anthropic API calls to NVIDIA NIM format
- Unlocks 40 requests/min for free
- Supports Kimi K2, GLM 4.7, MiniMax M2, Devstral and more
- Streams thinking tokens and tool calls live
- Even includes a Telegram bot so you can run Claude Code from your phone

No API bill. No rate limit panic. No vendor lock-in.

Honestly, this goes beyond router tools like OpenRouter.

It doesn't just swap the model... it turns Claude Code into a free agent you can control remotely.

The project is open-source on GitHub.

It's called free-claude-code.
♥ 6.2K · ⟲ 902 · 👁 652.0KView on X ↗

Ryan Wiggins Details Building a Local Second Brain With Claude Code

Creating a Second Brain with Claude Code

Mercury VP of Product Ryan Wiggins published a long article describing a locally run personal knowledge system built with Claude Code, indexing about 15,000 documents using QMD vector search, with hooks, orchestrators and MCP/CLI tools. He says it doubled his productivity and includes the workflow and prompt.

Original post · 11 min read
X ArticleCreating a Second Brain with Claude Code
I've 2x’d my productivity as a VP of Product @mercury by creating a "Second Brain" using 5 years of work history, 15k docs with 3.5 million words, and every tool in my stack. It runs locally, is a core part of my every use of LLM, and gets better everyday.
Today, I want to share the stack, the workflow, and the prompt to build it:
Background
I am a VP of Product for @mercury, which is a long way of saying I'm in a lot of meetings, consuming a lot of content across different tools (linear, slack, notion, data analyses), and trying to make sure I actually get stuff done. Working at a company for 5 years and being an information addict, I am essentially a walking encyclopedia for Mercury post 2021-today -- but I've recently found that my scope + workload means I can't keep every plate spinning.
One day, I was scrolling X and came across a series of posts that caught my attention, starting with @tobi's QMD. QMD is a local vector search, and then a few other posts started to show up that connected a few dots for me:
Claude Code launched hooks (per-event prompt injections)
GasTown / OpenClaw launched with the power of orchestrators writing memory + delegating to sub-agents (among many other patterns)
MCPs/CLIs hit a critical mass, and enough of my core tools were available without having to ask admins to give me API keys
@tylercowen did an interview and talked extensively about "writing for AI" in a way that struck a chord - how much output of work already exists that I'm not using?
I decided that it was time to build
Prep work (~1-2 hours end to end)
To start, I needed a library of all the content I could know about... so I downloaded every document I've ever created for my job at Mercury + any relevant product strategy, analysis, retro, reflection on execution, etc. This netted out to over 15k documents and 3.5 million words. Maybe I've read them all, but I've forgotten most. These became a folder that I just called "raw data", and I ran QMD to index this on my computer.
To see if this worked, I used Claude Code to ask about random memories and surprising insights from this knowledge base - the amount of delight/surprise I experienced in seeing how much more capable vector search was than text-based search gave me the confidence to keep going. I asked one questions about books that it would think I like, and it was spooky how good of recommendations it gave me. I think this is my best advice in this journey: test every step of the way! Easy to get caught in hill climbing a local maxima
Train my brain and connect it to my tools (~2 hours)
With all the raw data, I needed to help it make sense of me + what my goals are + the tools I used, so pursued three paths:
Explain myself - to be able to create a second brain, it needed to know what mine was doing. I wrote up a me.md explaining who I am (work + life), gave it my goals + performance reviews for the last 5 years + set of personal priorities. The most humbling part was the system pointing out that I've been making the same strategic mistake for years, according to my own performance reviews, and was making it that week as I was setting up the system
"Distill" the data - I spun up an agent team to use the me.md + the knowledge base to create a set of docs between me <> raw knowledge base. This idea largely came from the idea that LLMs regularly distill down smaller models to take tasks, and I had no idea if it would help me in this, but Agent Teams had just launched and so I had a swarm of them find the main "themes" we've worked on from the knowledge, give sourced histories of this, and summarize key lessons. These created a context.md folder
Tools - I use a few tools (Google Docs, Linear, Notion, Metabase) , and luckily most have connectors on Claude Code or these companies are actively launching MCPs/CLIs. A few didn't, but I spun up specific skills that crafted direct API calls to be able to complete tasks like "run a query for XYZ".
Claude had access to all the information about me + the tools I used + had a massive library of all my work, but did it really know anything? Does anyone?
Wire it up (<1 hour)
At this point, I had so many words + documents that it was time to actually find use or abandon ship. But I didn't want to have to go search this every time and that's when "hooks" caught my attention.
Hooks from Claude Code let you insert content into your prompt without needing to ask (or when a session starts, after a tool use, or when a session stops). Using the UserPromptSubmit hook, I enabled my Claude Code to use qmd to find names + topics + specific documents related to my prompt.
This is a nerd-out moment, but when searching for files in Finder, it is mostly a name + raw text search.... but QMD can help bring context into searches. My system is tuned to figure out a query, then returns results using one of two techniques:
vsearch (semantic/vector) — understands meaning of my question. "How's the funnel performing?" finds … continue on X ↗
♥ 922 · ⟲ 120 · 👁 313.4KView on X ↗