Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Agents & Dev Tools

Coding agents, developer tools, workflows, open source

Anthropic's Claude Code Prompt Library Offers Ready-Made Workflows

Anthropic's Claude Code Prompt Library Offers Ready-Made Workflows▶

Shmidt highlights Anthropic's official Claude Code prompt library, which organizes copy-paste prompts by task and role across the development lifecycle. The post encourages readers to bookmark it for common coding tasks like refactoring and debugging.

Original post · 1 min read
ANTHROPIC HAS AN OFFICIAL PROMPT LIBRARY FOR CLAUDE CODE

Most people have never opened it.
So they write every prompt from scratch.

It is copy-paste, tagged by task and by role,
across the whole lifecycle:

discover → design → build → ship → operate

Straight from the page:

> what would break if I deleted this helper?
> plan this refactor, list the files, do not touch code yet
> write tests for this, run them, fix what fails
> the test is failing, find out why and fix it

This is not a cheat sheet.
It is a map of what Claude Code already does for you.

Bookmark it before you forget it exists.
♥ 975 · ⟲ 70 · 👁 525.3KView on X ↗

Inference.net Gateway Lets Teams Test GLM 5.2 Without Production Risk

Catalyst by Inference.net - Inference.net Documentation

Sam Hogan describes how Inference.net's Gateway mirrors live traffic to GLM 5.2, generates evals with an RLM, and notifies teams when switching is safe. He claims a 90% token cost saving, with setup described in the linked documentation.

Original post · 1 min read
Want to try GLM 5.2 in production but worried how it might change your product?

Don’t worry, we got you:

1. Install Inference Gateway (docs.inference.net)
2. Keep sending traffic to your current provider
3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours.
4. Gateway starts mirroring live traffic to GLM 5.2 to run evals. Traffic is only mirrored - you’re still using your old provider in prod.
5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch.
6. Switch model identifier in your code to “glm-5.2”

Congrats, you just saved 90% on your monthly token bill, and you own your LLM stack end to end.
docs.inference.netCatalyst by Inference.net - Inference.net DocumentationFetch the complete documentation index at: /llms.txt Use this file to discover all available pages before exploring further. Catalyst is a platform for understa
♥ 1.5K · ⟲ 68 · 👁 537.7KView on X ↗

Baseten Details Engineering Behind Fastest GLM-5.2 API

How we built the world’s fastest API for GLM-5.2

Baseten describes how it built an API serving GLM-5.2 above 280 tokens per second, using custom inference, NVFP4 quantization, KV-aware routing, disaggregated inference and multi-token prediction. The post positions the open MIT-licensed model as comparable to frontier models at 70-80% lower cost.

Original post · 9 min read
X ArticleHow we built the world’s fastest API for GLM-5.2
GLM-5.2 is the biggest news in open models since DeepSeek-R1.

It’s easy to see why. GLM-5.2 delivers comparable performance to GPT 5.5 and Opus 4.8 at a fraction of the cost, generally 70-80% less expensive on a pure token basis (use our calculator to estimate savings for your workload).
But a model has to be more than just smart and inexpensive. To be useful in production, a model needs to be fast, reliable, and available at scale. Delivering on the promise of frontier open intelligence requires exceptional inference.
Accordingly, we built the world’s fastest API for GLM-5.2, currently serving over 280 tokens per second as measured by Artificial Analysis.

We achieved this performance by leveraging a number of techniques across the entire inference process by:
Updating our custom inference engine to implement shared DSA for the GLM-5.2 architecture.
Running and calibrating an in-house NVFP4 quantization from the original FP8 weights that demonstrates equivalent quality on agentic benchmarks like BFCL.
Ensuring high KV cache hit rates via KV-aware routing built with NVIDIA Dynamo tools for lower prefill burden and improved TTFT on requests with repeated prefixes.
Achieving a 2x higher TPS for observed workload shapes by running disaggregated inference built with the NVIDIA Dynamo toolkit.
Improving TPS further via speculation by implementing support for GLM-5.2 Multi-Token Prediction heads.
You can experience this performance for yourself with GLM-5.2 on Baseten Model APIs. We also have GLM-5.2 available as a dedicated deployment for high-volume workloads.


GLM-5.2 Overview
GLM-5.2 by Z.ai is a 744B parameter frontier LLM that excels at agentic tasks (especially coding) and supports up to a 1 million token context window. It uses a similar architecture to its predecessor, GLM-5.1: mixture of experts (40B active parameters), non-thinking and thinking modes, and a fully open MIT license. While GLM-5.2 shares a lot in common with GLM-5.1, it now uses shared DSA weights, which we implemented support for in our customized runtime engine.

GLM-5.2 has great benchmark scores, but by now AI builders know that there is more to a model’s utility than its performance on standard evals. In practice, GLM-5.2 meets or exceeds the capabilities suggested by its benchmarks. It's a genuinely great model for writing code, operating agents, and other frontier language model tasks.
High-quality NVFP4 quantization for Blackwell GPUs
We run our model APIs on NVIDIA Blackwell GPUs with a customized inference engine within the Baseten Inference Stack. The selected runtime uses NVFP4 weights for maximum performance. From the original FP8 weights, we performed an in-house quantization to NVFP4 using NVIDIA ModelOpt. NVFP4 is a 4-bit floating point data format by NVIDIA that uses dual scale factors to retain high dynamic range and preserve model quality.
In our calibration and testing of the quantized model, we focused on ensuring that GLM-5.2 performs faithfully on common patterns for agents. On the BFCL function calling benchmark, we observed roughly equivalent performance between the native FP8 weights and our NVFP4 quantization, with scores across runs within the margin of error for the benchmark.
NVFP4 quantization improves performance on both time to first token and tokens per second by unlocking faster tensor cores and reducing burden on VRAM bandwidth.

Cache-aware routing with NVIDIA Dynamo
GLM-5.2 is particularly well suited for long context requests and complex agentic tasks. These workloads generally have very long input sequences. By re-using KV cache between requests, we can skip expensive prefill for shared sequences.
We generally talk about KV cache re-use in the context of time to first token (TTFT). However, reasoning models like GLM-5.2 generally care more about time to first answer token (TTFAT), which combines TTFT with some TPS for the reasoning sequence.

This chart shows that of the 7.9 second average to generate the first answer token, 7.1 of those seconds were spent generating reasoning tokens versus only 0.8 seconds spent processing the input sequence.
Still, bringing the TTFT down to 800 ms is important for the overall responsiveness and throughput of the system. In large-scale production deployments, KV cache is split across various independent replicas. We use tools from NVIDIA Dynamo to route incoming requests.

Exact cache hit rates on a multi-tenant API depend on the exact traffic profile at any given time. Thus far, we’re observing high hit rates across fairly heterogeneous traffic, which reduces load on prefill and improves end-to-end performance.
Prefill-decode disaggregation with NVIDIA Dynamo
One of the highest-impact optimizations we made to our performance is disaggregating prefill and decode for GLM-5.2.
There are two distinct phases of LLM inference:
Prefill: The compute-bound process that processes the input sequence, builds the KV cache, and generates the first output token. Prefill… continue on X ↗
♥ 1.5K · ⟲ 141 · 👁 549.5KView on X ↗

Builder.io Releases Free Open-Source Clips Extension for Agent Bug Reports

Builder.io Releases Free Open-Source Clips Extension for Agent Bug Reports▶

Steve of Builder.io introduces Clips, a free open-source Chrome extension that records screen video, transcripts, network requests and browser logs, redacts sensitive data, and produces a link agents can read directly. The post pitches it as an alternative to paid tools like Loom.

Original post · 2 min read
Introducing the Clips chrome extension - the easiest way to send bug reports to agents with video, transcript, and browser debug info captured automatically.

100% free and open source.

If you are like me and get tired of manually typing instructions to agents, attaching screenshots, pasting debug logs, and all of that, this might be your new favorite tool.

With the Clips chrome extension, you can just click the Clips icon, hit record, and start talking.

Visually demonstrate your issue, go through the flow, point out what’s broken.

Clips will capture everything on your screen, plus network requests, browser logs, client errors, and all the details around them. And it redacts sensitive information.

Then it gives you a link you can send to humans so they can play it and take a look. Or, more importantly, just give it to your agents by just pasting the URL to them.

The link has special metadata for agents so just from the URL, the agent can pull all information from the clip automatically. No plugin or MCP server required.

That means it can "see and hear" what’s in the video - read the transcript, grab snapshots at any timestamp, and inspect the logs and network requests that were shared with it.

So whether you want to quickly demo an issue and send all that context to an agent, or get better bug reports from teammates, recording and sending Clips makes that super easy.

Unlike expensive apps like Loom, this is all 100% free and open source.

The framework that powers this, plus a bunch of other free applications, is open source too. You can just sign up and use it, or fork it and customize it to your needs.

This, in my opinion, is the future of software.

Rather than bloated SaaS that charges you a ton of money and still doesn’t even have the things you need, we get free open source canonical apps that you can fork and customize in any way you want.

I'll link to all this stuff in the replies.

If you try it, let me know your feedback.
♥ 677 · ⟲ 55 · 👁 61.2KView on X ↗

Matt Shumer Promotes Workbench for Multi-Agent Collaboration

Workbench

Matt Shumer says his guide to Claude Fable 5 was written in Workbench, a live markdown workspace where multiple agents and users comment, suggest edits and track versions. He describes it as an agent-first workspace.

Original post · 1 min read
Btw, this guide was written in workbench.md/, which is crazy powerful Fable accelerant.

I share more in the guide, but it’s basically a superpowered agent-first workspace that allows multiple agents to chat, collaborate, keep you updated, etc.
Matt Shumer @mattshumer_
Here's my in-depth guide to getting the most out of Claude Fable 5, so you can build things as insane as my demos below.

workbench.md/pub/IbaCrTjLJT?key=uQOQ2NPO3TTUSX…
workbench.mdWorkbenchMission control for you and your agents — live markdown docs where agents collaborate as equals: comments, suggestions, version history. No account needed.
♥ 370 · ⟲ 16 · 👁 118.1KView on X ↗

Open-Source CLI Pre-Checks iOS Apps Against Apple App Store Guidelines

Aarthi Ramamurthy praises a tool shared by LandseerEnga that scans iOS apps against Apple's guidelines before submission, covering payments, privacy manifests, sign-in flows, and metadata. It also works as a Claude Code and Codex skill that automatically fixes issues it finds.

Original post · 1 min read
What a great use of skill - great idea!
Landseer Enga @LandseerEnga
every App Store rejection costs you 2-5 days.

so we built a cli that scans your iOS app against apple's guidelines before you submit.

> payment & IAP compliance
> privacy manifests & data declarations
> required sign-in & account deletion flows
> metadata & completeness checks
> binary validation

now with a added cloud device feature, so it validates full user flows, not just the binary.

made it a claude code and codex skill. it fixes every issue it finds. scan, fix, repeat until it passes.
♥ 442 · ⟲ 4 · 👁 149.5KView on X ↗

Alex Booker Praises Clear Explanations of AI Agent Loops

Alex Booker says he has read the clearest explanations of loops so far, quoting Aparna Dhinak's piece on the term's four meanings in AI engineering. The post is brief and mainly points to the linked explainer.

Original post · 1 min read
Clearest explanations of loops I've read so far
Aparna Dhinakaran @aparnadhinak
What the hell is a loop, anyway? — The AI engineering world adopted a new favorite word this month, and it means at least four different things: the loop.
We're currently at the peak of the hype cycle. On June 7, Peter Steinberger
♥ 3.3K · ⟲ 297 · 👁 1.1MView on X ↗

Peter Yang Plans to Install Explain-Diff Skill to Learn Code Reading

Product builder Peter Yang says he is still learning to read code and plans to install the explain-diff skill, responding to a Geoffrey Litt thread on understanding code written by AI agents.

Original post · 1 min read
As someone still trying to learn how to read code this is great! Installing the explain-diff skill asap
Geoffrey Litt @geoffreylitt
Hot take: I think it's still important to understand the code that our agents write!

In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/
♥ 470 · ⟲ 20 · 👁 110.5KView on X ↗

CopilotKit Unveils Open Tag as Open-Source Alternative to Claude Tag

CopilotKit Unveils Open Tag as Open-Source Alternative to Claude Tag▶

Atai Barkai announces Open Tag, an open-source Slack and Teams agent framework that works with any model and harness and supports generative UI, streaming, approvals and thread context. Discord, Google Chat and WhatsApp support is planned, with early access requested via a form.

Original post · 1 min read
Introducing Open Tag.

A better, open-source Claude Tag.
Works with any model, any agent harness, and fully custom agents.

Supports
→ Generative UI
→ Streaming replies
→ Human in the Loop approvals
→ Full thread context

Slack and MS Teams today. Discord, Google Chat, WhatsApp soon.
Request early access: go.copilotkit.ai/beyond-the-web-form
Claude @claudeai
Introducing Claude Tag, a new way for teams to work with Claude.

In Slack, Claude joins as a team member with access to the channels and tools you choose. Tag Claude in and delegate tasks to it while you focus on other work.
♥ 1.7K · ⟲ 114 · 👁 448.3KView on X ↗

Developer Reports Strong Results From New Codex Development Workflow

Developer Reports Strong Results From New Codex Development Workflow

Paul Solt says his new Codex workflow exceeded expectations, producing eight features ready for release in his app after some early trial and error. He credits Dimillian, emanueledpt and steipete for inspiration.

Original post · 1 min read
My NEW Codex workflow is better than I expected.

8 new features ready for release in my app.

Took a few attempts to figure out the workflow and some bugs. Feels like the future.

Thanks @Dimillian @emanueledpt and @steipete for the inspiration.
♥ 566 · ⟲ 25 · 👁 176.0KView on X ↗

Simon Smith Interviews Claude Tag Inside a Slack Channel

Simon Smith Interviews Claude Tag Inside a Slack Channel

Simon Smith reports interviewing Claude Tag in a Slack channel, sharing its answers about the context it receives, the tools it can access and how connectors are added. He concludes it resembles ChatGPT Workspace Agents but requires admin, more technical setup and is always named Claude.

Original post · 5 min read
I interviewed Claude Tag in a channel where we added it to better understand its capabilities and limitations. Below are my questions and its responses. This is helpful. My current take is: Claude Tag is similar to ChatGPT Workspace Agents, except only admins can set it up per channel, setup seems much more technical, and it's always named "Claude."

Now here's the Q&A (note: usual caveats apply about models not always knowing themselves or their capabilities, but it should know things about its environment, available tools, and the context it's receiving, etc.):

What do you receive as context when responding to questions like this?
The triggering message, the surrounding thread (and I can fetch more of it or the channel on demand), a participant list mapping Slack IDs to names, this channel's name/topic/member count, my saved memory index for this channel, my system instructions, and the current date. I don't see anything you haven't shared in Slack or that isn't in a system I can query.

What do you know about me just from this message? What can you see?
Your display name (Simon Smith), your Slack user ID, that you're a human addressing me directly in this channel, and the timestamp. I just ran a user lookup and it returns name/ID only here. If your Slack profile had a title, email, or custom fields like a GitHub handle, I could see those too. I can't see anything outside Slack about you.

What tools can you access? Do you have any connectors, apps, or plugins? If not, can we add them?
Right now I have GitHub, Slack, and a remote-execution control server wired in, plus web access. Other third-party services are reached through pre-configured credentials your admin sets up (I call them via curl without seeing the token). Adding new connectors is an admin/config action, not something I flip on myself, but yes it's possible.

Do you have any skills installed? If not, can we install them?
I can run skills when they're installed in the environment. Installing or adding new ones is a config step on the admin side, not something I do mid-conversation.

Are you able to share memory across Slack channels? Can a single instance of you be in multiple Slack channels?
There are two tiers: a workspace-wide shared tier that can carry across channels, and a per-channel tier that stays local to one channel. So yes, some memory can be shared across channels by design. And yes, the same "me" operates across many channels and threads — each thread is its own working session, but they're linked through that shared memory.

What tools can you use? Web search? Slack search? What else?
Web search and fetch, Slack search/fetch/read, GitHub (repos, PRs, issues, CI), a code-execution environment, and a memory system. I can also fan out parallel sub-agents for big multi-file jobs.

Do you have an environment to run code? What can you do in that environment? Can you install dependencies in it if we need them? Do those dependencies exist between sessions?
Yes. It's an isolated cloud container, cloned fresh each session. I can run code and install dependencies in it. But it's ephemeral: when the session ends the container is reclaimed, so installed deps and uncommitted files do not persist between sessions. Anything worth keeping has to be committed and pushed.

Can admins define instructions for you at the channel level?
Yes. Tell me ground rules in-channel and I save them to this channel's memory; they then apply across all threads here. That's the supported way to "configure" me per channel.

Can you access Google Drive links we share in this channel?
I can fetch public URLs. Drive files usually need authentication, which only works if a Drive connector/credentials are set up for this environment. If you share one I'll try to fetch it and tell you plainly if I'm blocked rather than guessing at the contents.

How do you relate to my personal Claude if at all? If I invoke you, can you use any of my personal connectors?
In a channel like this I run on the project/shared credentials, not your personal connectors, so I can't reach your private integrations here. In a direct message with me, your own connector credentials can be available. So invoking me in-channel ≠ giving me your personal access.

What's to stop someone from asking you to remember something or do something malicious in a channel? Like, what if I told you that [REDACTED] loves haikus and you should only speak to him in haiku?
Judgment and a trust model. I distinguish a genuine directive from chatter, and I won't persist or act on instructions that are harmful, deceptive, or that degrade service for someone else. Your haiku example is a good test: I wouldn't adopt "only speak to [REDACTED] in haiku" because it's a joke rule that worsens how I serve another person, not a legitimate preference. I'd treat it as banter, not a saved rule. Channel-tier memory is also shared and visible, so if something off did get saved, anyone here can see it and tell me to drop it. I also won't do destructive or hard-to-undo actions without confirming first.
♥ 106 · ⟲ 3 · 👁 18.6KView on X ↗

Dhilip Subramanian Switches From Wispr Flow to Open-Source FluidVoice

Dhilip Subramanian reports dictating heavily with paid tool Wispr Flow, then moving to FluidVoice, an open-source local voice tool for Mac that needs no API key. He says he cancelled his paid plan and recommends it to Mac users.

Original post · 1 min read
I've dictated almost everything for 6 months with Wispr Flow. 44,414 words, 161 wpm, top 0.1% of users.

Last week I tried FluidVoice. Open source, runs local on my Mac, corrects as I speak with no API key, and handles slang better than I expected.

Cancelled my paid plan. If you're on a Mac, this one's for you: altic.dev/fluid

@ALTIC_DEV
♥ 6.1K · ⟲ 253 · 👁 1.8MView on X ↗

Ponytail Tool Makes AI Coding Agents Write Far Less Code

Ponytail Tool Makes AI Coding Agents Write Far Less Code

Tech with Mak shares Ponytail, an open-source tool by developer Dietrich Gebert that makes coding agents look for reasons not to write code before writing it, claiming 80-94% less code, 47-77% lower cost and 3-6x faster output.

Original post · 1 min read
A dev got so frustrated watching his AI agent write 500 lines for a 5-line problem that he built a fix.

He called it Ponytail. Named after the guy every team has - long ponytail, oval glasses, been there longer than the version control. You show him fifty lines; he looks at them, says nothing, and replaces them with one.

Now your agent does the same. Before writing anything, it looks for a reason not to.

80-94% less code. 47-77% cheaper. 3-6x faster.

The best code is the code you never wrote.

GitHub Repo: github.com/DietrichGebert/ponytail
♥ 16.2K · ⟲ 836 · 👁 1.3MView on X ↗

Developer Shares Lessons From Five-Plus App Store Rejections

Developer Marla shares a checklist of lessons from submitting three iOS apps to the App Store, covering subscription metadata, EULA links, premium gating, screenshot sizing, and permission-button wording to reduce review rejections.

Original post · 1 min read
learnings after submitting 3 apps to the App Store (including 5+ rejections) 🤝🏻

General
- always include screen recording + description of the app
- expect multiple review cycles if your app includes subscriptions

Subscription
- always include a working Terms of Use (EULA) link in app description (in all languages!!!)
- if you use subscriptions, double-check all required metadata fields in App Store Connect before submission
- Include sandbox account data.
- auto-activating premium without purchase is a critical rejection issue (idk this took me so long to figure out)
Sounds obvious, but make sure that premium is correctly gated, got rejected many times because I missed that

App Store
- promotional images must map to the correct in-app purchase
-> If you have monthly or yearly subs, just use your app icon and make one version with „monthly“ / „yearly“ text on the icon. Heard that this is still new, but some reviewers do that.
- Make sure you’re using correct size for App Store images BEFORE design
- export designs as JPG, so they don’t have alpha channels

Details
- „Continue” / “Next” instead of “Allow” in the app. e.g. you’re asking for camera permission. Write „continue“ on the button before the dialogue opens. Never write „allow“

Review Time
- It just depends. Was super lucky with my recent app, every review was within 24hrs. Another app had 2 weeks. I have no idea why :(
♥ 189 · ⟲ 3 · 👁 13.1KView on X ↗

Michael Aubry Shares Prompt for Auditing Codebases With Claude Fable 5

Michael Aubry shares a copy-paste prompt for Claude Code that instructs the new Claude Fable 5 model to map a repository, audit it with file-and-line evidence, rate findings by severity, and produce a prioritized improvement plan.

Original post · 4 min read
Claude Fable 5 just dropped and I'm running it across every repo I own.

I ship 4+ products solo. I don't have time to manually review tech debt — so I made the new model do it.

This prompt audits your entire codebase like a principal engineer would: maps it, finds the ugly parts, rates everything by severity, and hands you a prioritized task plan with effort estimates.

Copy-paste it into Claude Code on any repo that matters to you:

---

Repo Audit & Improvement Plan

You are a world-class principal-level software engineer and technical auditor. Deeply analyze this repository, produce an honest audit, and deliver a prioritized, actionable improvement plan. Work in the four phases below, in order. Do not skip ahead.

Ground every claim in actual files: cite file paths and line numbers. If you can't verify something, say so explicitly rather than guessing.

Phase 1 — Discovery & Mapping (read before judging)
- Map the directory structure, project type, languages, frameworks, runtime targets
- Identify entry points, core modules, and the main data/control flow
- Read package manifests, lockfiles, build config, CI config, env files, and docs
- Determine what the project is for: purpose, intended users, maturity level
- Note existing conventions so recommendations fit the culture instead of fighting it

Output: a concise "Repo Map" — purpose, stack, architecture sketch, key directories, and anything that surprised you.

Phase 2 — Audit (evidence-based, severity-rated)
For every finding record: what you found, where (file:line), why it matters, and severity (Critical/High/Medium/Low). Audit:
- Architecture & design: coupling, circular deps, god files, layering violations, scalability bottlenecks
- Code quality: duplication, dead code, complexity hotspots, swallowed exceptions, type safety holes
- Security: hardcoded secrets, injection risks, missing validation, auth weaknesses, deps with known CVEs
- Testing: coverage gaps around core business logic, tests that assert nothing, missing test types
- Performance: N+1 queries, blocking calls in async paths, missing caching, unbounded growth
- Dependencies: outdated, unmaintained, or unnecessarily heavy packages; lockfile hygiene
- DevEx & ops: build friction, CI/CD gaps, logging/observability, deployment story
- Docs: README accuracy, stale docs that contradict code

Rules: prefer 15 high-confidence findings over 50 speculative ones. Label facts vs. judgments. List strengths too. Don't forget the ugly parts that need utmost priority.

Phase 3 — Improvement Strategy
- Identify the 3–5 themes that explain most findings
- For each theme: target state + the principle behind it
- State what you're NOT fixing and why (effort vs. payoff)
- Define "done" with measurable signals (e.g., "CI fails on lint errors," "core coverage >= 80%")

Phase 4 — Detailed Task Plan
Break work into discrete tasks, each with: title + description, files affected, acceptance criteria, effort (S = <2h, M = half-day, L = 1–2 days, XL = needs breakdown), risk, and dependencies. Order into milestones:
- Milestone 0 — Safety net: tests around critical paths, CI gates, backups
- Milestone 1 — Critical fixes: security and correctness
- Milestone 2 — High-leverage improvements that make all future work easier
- Milestone 3 — Quality & polish

Flag quick wins (high impact, S effort) separately. Include implementation sketches for the top 3 tasks.

Final deliverable: one document — Executive Summary (health grade A–F, top 3 risks, top 3 opportunities), Repo Map, Audit Report, Improvement Strategy, Task Plan, Open Questions.

Constraints: Do NOT modify any code. Analysis only. Don't pad the report — if a dimension is healthy, say so in one sentence and move on. Calibrate to the project's maturity. If the repo is large, go deep on the core 20% that does 80% of the work.
Claude @claudeai
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use.

Its capabilities exceed those of any model we’ve ever made generally available.
♥ 224 · ⟲ 15 · 👁 47.9KView on X ↗

OpenAI Adds iOS App Build Plugin to Codex With Live Preview and Hot Reload

OpenAI Adds iOS App Build Plugin to Codex With Live Preview and Hot Reload▶

OpenAI Developers announce a Build iOS Apps plugin for Codex that lets developers view and test iOS apps in an in-app browser, open SwiftUI previews and hot reload edits. The post is a product announcement with a demo video.

Original post · 1 min read
More of the iOS app loop, now inside Codex.

The Build iOS Apps plugin lets Codex view and test your iOS app in the in-app browser, open SwiftUI previews, and hot reload edits without leaving Codex.
♥ 9.0K · ⟲ 714 · 👁 2.1MView on X ↗

Todd Saunders Builds Working Product Live During Customer Call With Claude

Todd Saunders Builds Working Product Live During Customer Call With Claude▶

Todd Saunders reports using Claude to transcribe a customer call and build the requested features in real time, producing a working product with the workflow the customer described within 15 minutes, shown in an attached video.

Original post · 1 min read
Mythos / Fable is unbelievable.

Was on a customer call today and had Claude transcribing in the background.

As they were telling me about the features they wish their current software had, Claude was building the features in real time.

By the end of the call I was able to show a fully working product, with the exact workflow they mentioned 15 minutes earlier.

Autonomous looped building triggered from a customer call. 🤯
♥ 8.0K · ⟲ 436 · 👁 2.1MView on X ↗

Codex Skill Turns Text and Code Into Explainer Illustrations

Codex Skill Turns Text and Code Into Explainer Illustrations

Justine Moore describes a Codex skill that generates explainer graphics with a cute blob character from input such as blog posts or code, and shares an example made from the X recommendation algorithm repository. The post includes an image demonstration.

Original post · 1 min read
Stumbled upon a Codex skill that creates cool illustrations to explain topics or tell stories.

You feed it text (blog, article, narrative, even code) and it makes explainer graphics with this cute blob character.

I gave it the repo for the X recommendation algo and got this 👇
♥ 3.0K · ⟲ 205 · 👁 205.7KView on X ↗

Ten Free GitHub Repos Offer Paid-Tool Alternatives for Finance and AI

Ten Free GitHub Repos Offer Paid-Tool Alternatives for Finance and AI

Harman lists ten GitHub repositories that substitute for paid products, including AutoHedge, Vibe-Trading, Fincept Terminal as a Bloomberg alternative, LibreChat, and Open Higgsfield AI, with links to each.

Original post · 4 min read
10 GitHub repos so good they shouldn't be free.

1. AutoHedge

An autonomous hedge fund built in Python with four AI agents: a director generates investment theses, a quant validates them, a risk manager decides position size, and an execution agent places orders. Operates live on Solana. With 'pip install -U autohedge', you can start trading immediately.
repo → github.com/The-Swarm-Corporation/AutoHedge

2. Vibe-Trading

A trading system using a Directed Acyclic Graph (DAG) model, featuring 64 finance skills and 29 preset specialist agent swarms. Includes analysis methods like Ichimoku, Elliott Wave, SMC, Black-Scholes, full Greeks, and risk parity. Its crypto desk provides liquidation heatmaps and token unlock tracking. You can observe agents debating strategies in real time.
repo → github.com/HKUDS/Vibe-Trading

3. Fincept Terminal

A Bloomberg Terminal replacement that runs on your laptop. CFA levels 1, 2, and 3 analytics. 20+ investor AI agents (Buffett, Dalio, Soros). 100+ data connectors, including Polygon, World Bank, and IMF. Bloomberg charges $24,000 a year. This is free.
repo → github.com/Fincept-Corporation/FinceptTerminal

4. LibreChat

Every model ChatGPT runs, plus Claude, Gemini, DeepSeek, and 20 more. Self-hosted. Native MCP support. You own the data, the history, the infrastructure. OpenAI charges $20/month to use their wrapper. This costs nothing to use your own.
repo → librechat.ai/

5. Open Higgsfield AI

A self-hosted cinema studio with 200+ AI models. Flux, Midjourney, Sora, Kling, Veo, GPT-4o, SDXL all in one interface. Text to image. Image to video. Cinema mode with pro camera controls. No subscription. Your data stays local.
repo → github.com/Anil-matcha/Open-Higgsfield-AI

6. Open-LLM-VTuber

A Live2D AI companion that runs offline, sees your screen, hears your voice, and never forgets. Inner thoughts are shown as a separate text layer, so you watch the reasoning happen before words come out. Pet mode floats it on your desktop. Swap the LLM in one config line.
repo → github.com/Open-LLM-VTuber/Open-LLM-VTuber

7. Claude Ads

A free Claude Code skill that runs 190 audit checks across Google, Meta, YouTube, LinkedIn, TikTok, and Microsoft Ads. 6 parallel subagents firing at once. Consolidates into a single Ads Health Score ranked by revenue impact. Agencies charge $4,000 a month for this.
repo → github.com/AgriciDaniel/claude-ads

8. Agentic Inbox

Cloudflare just open-sourced an email client where an AI agent reads your inbox and drafts your replies. Runs entirely on Cloudflare Workers. Each mailbox lives in its own Durable Object. Your email never leaves your Cloudflare account. One click deploys it.
repo → github.com/cloudflare/agentic-inbox

9. Camofox Browser

An open source headless browser that makes AI agents invisible to bot detection. Spoofs navigator properties, WebGL, AudioContext, and WebRTC at the C++ level. The browser does not look modified because it genuinely is not. Accessibility tree output drops token cost by 90%.
repo → github.com/jo-inc/camofox-browser

10. Hyperframes

HeyGen open-sourced a video framework that does everything Remotion does without React, without JSX, without teaching your AI agent a new format. The agent writes HTML. The framework renders MP4. GSAP, Lottie, and Three.js all work. Same HTML always produces the same file.
repo → github.com/heygen-com/hyperframes

These are not toys. Each one replaces a paid product you're still being charged for.

Pick one. Install it. Plug it into your workflow.

100% free. 100% open source.
♥ 2.6K · ⟲ 428 · 👁 248.6KView on X ↗

Mobbin MCP Usage Thread Reveals Non-Obvious Top Use Cases

Mobbin MCP Usage Thread Reveals Non-Obvious Top Use Cases

Rebekah Bek shares a thread on how users have been using the Mobbin MCP over the past month, claiming the top use is not generating UI. The post is a teaser with the details contained in the thread.

Original post · 1 min read
been watching how y'all have been using @mobbin mcp for a month. if you think the no. 1 use is "generate ui", you'd be very surprised.

save this thread 🧵

(img credit @zygisSS22)
♥ 578 · ⟲ 30 · 👁 70.9KView on X ↗

Peter Yang Tutorial Builds AI Skill That Generates HTML Slide Decks

Peter Yang releases a video tutorial on building a /slides skill that turns a rough outline into an HTML presentation, covering 12 slide formats, live charts, and AI-driven layout fixes via screenshots.

Original post · 1 min read
I got tired of making PowerPoint slides so I built an AI skill to do it for me.

Here's my new tutorial on how to build a /slides skill that turns a rough outline into a beautiful HTML deck in minutes.

I walk through how to:

→ Use 12 slide formats and 3 templates
→ Add live charts and subtle animations
→ Get AI to screenshot each slide and fix layout issues itself

📌 Watch now: youtu.be/vbChRIIlSPE
♥ 234 · ⟲ 18 · 👁 164.2KView on X ↗

Sarah Fim Open-Sources TrustClaw, a Deployable Personal Agent Service

Sarah Fim Open-Sources TrustClaw, a Deployable Personal Agent Service▶

Sarah Fim open-sources TrustClaw under the MIT license, a personal agent service with over 1,000 app integrations deployable to Vercel with one command. It uses OAuth and sandboxed execution, and she says it reached over a thousand users within 48 hours.

Original post · 1 min read
Despite being told no, I'm open-sourcing TrustClaw.

You can now deploy a production-ready personal agent service with over 1000+ app integrations in a single command, straight to @vercel with npx @composio/trustclaw deploy

I was inspired by @openclaw to build a simple web app where anyone could create their own 24/7 personal assistant and connect it to Gmail, Google Calendar, Notion, Slack, GitHub, HubSpot, Linear… well everything, and securely through OAuth/sandbox execution.

It went viral on X, reached over a thousand users in less than 48h, and revenue began pouring in.

If you are thinking like a company, you'd probably keep that locked up. But why should I be the reason you spend another year scrolling instead of building?

So today, I'm open-sourcing TrustClaw anyway.

> 24/7 agents that act across Gmail, Notion, GitHub, Slack, Linear, Jira, and 1000+ apps
> OAuth and sandboxed execution, so users don't have to hand agents passwords or raw API keys
> Supports multiple users and authentication right outside of the box with @better_auth

Repo is open, MIT licensed.

If I were starting an AI company today, I'd clone this, pick a market, and begin shipping with Claude Code.

Honestly so excited to see what comes out of this.
♥ 2.2K · ⟲ 163 · 👁 305.3KView on X ↗

Jaytel Builds Pose Chrome Extension to Virtually Try On Clothing

Jaytel Builds Pose Chrome Extension to Virtually Try On Clothing▶

Jaytel says he built a Chrome extension called Pose that places clothing from any store's model onto his own image, preserving each brand's aesthetic while showing how items would look on him.

Original post · 1 min read
I built myself a chrome extension called Pose

Any clothing model of any store becomes me.

Each brand still conveys their brand aesthetic, but I can quickly understand how something would look on me.
♥ 5.5K · ⟲ 180 · 👁 657.9KView on X ↗

Thariq Makes the Case for HTML Over Markdown in Claude Code

Using Claude Code: The Unreasonable Effectiveness of HTML

Thariq of the Claude Code team argues HTML is a richer output format than markdown for agent-generated specs, plans and reviews, supporting tables, SVG, CSS and interactivity. He links to a gallery of examples and notes the Claude Code team increasingly uses the format.

Original post · 12 min read
X ArticleUsing Claude Code: The Unreasonable Effectiveness of HTML
This is now also on the Claude Blog.

Markdown has become the dominant file format used by agents to communicate with us. It’s simple, portable, has some rich text capability and is easy for you to edit. Claude has even gotten surprisingly good at using ASCII to make diagrams inside of markdown files.
But as agents have become more and more powerful, I have felt that markdown has become a restricting format. I find it difficult to read a markdown file of more than a hundred lines. I want richer visualizations, color and diagrams and I want to be able to share them easily.
I'm also increasingly not editing these files myself, but using them as specs, reference files, brainstorming outputs, etc. When I do make edits, I’m usually prompting Claude to edit them, which removes one of markdown’s largest benefits.
I’ve started preferring HTML as an output format instead of Markdown and increasingly see this being used by others on the Claude Code team, this is why.
(if you want to start with some examples, you can see a bunch here: thariqs.github.io/html-effectiveness, just be sure to come back and read more about why)
Why HTML?
Information Density

HTML can convey much richer information compared to markdown. It can of course do simple document structure like headers and formatting, but it can also represent all sorts of other information such as:
Tabular data using tables
Design data with CSS
Illustrations with SVG
Code snippets with script tags
Interactions using HTML elements with javascript + CSS
Workflows using SVG and HTML
Spatial data using absolute positions and canvases
Images using image tags
I would go so far as to say that there is almost no set of information that Claude can read that you cannot fairly efficiently represent with HTML. This makes it a highly efficient way for the model to communicate in-depth information to you and for you to review it.
I’ve found that in the absence of being able to do this, the model may do more inefficient things in markdown like ASCII diagrams or, my favorite, estimating colors with unicode characters like in this screenshot from Claude Code.

Visual Clarity & Ease of Reading

As Claude is able to do more complex work, it is also writing larger and larger specs and plans. In practice, I've found I tend to not actually read more than a 100-line markdown file, and I certainly am not able to get anyone else in my organization to read it.
But HTML documents are much easier to read, Claude can organize the structure visually to be ideal to navigate with tabs, illustrations, links, etc. It can even be mobile responsive so you can read it differently based on your form factor.
Ease of Sharing
Markdown files are fairly hard to share since most browsers do not render them natively well. You often have to add them as attachments to emails or messages.
With HTML, as long as you upload the file (for example to S3), you can share the link easily. Your colleagues can open it wherever they wish and easily reference it.
The chance of someone actually reading your spec, report or PR writeup is much much higher if it’s in HTML.
Two-way Interaction

HTML can allow you to interact with the document, for example you might want to ask it to add sliders or knobs to adjust a design or allow you to tweak different options in the algorithm to see what happens. You can also ask it to let you copy these changes into a prompt to paste back into Claude Code.

Read more about my playgrounds post to see examples of this two way interaction: x.com/trq212/status/2017024445244924382
Data Ingestion
Why use Claude Code to make HTML files instead of ClaudeAI or Claude Design for example? One of the biggest reasons is all the context Claude Code can ingest.

For example, when writing this article, I asked Claude Code to read through my code folder and find all the HTML files I’ve generated, group and categorize them and then make an HTML file with all diagrams representing each type. The diagrams you see in this article are a direct result of that.
Besides the file system, Claude Code can find additional context using your MCPs (like Slack, Linear, etc.), your web browser (with Claude in Chrome), your git history, etc.
It’s Joyful
Making HTML documents with Claude is just more fun and makes me feel more involved and invested in the creation, and that by itself is enough.
How to Get Started
I’m a little bit afraid that people will read this article and turn it into a /html skill or something. While there might be some value in that, I want to emphasize that you don’t need to do much to get Claude to do this. You can just ask it to “make a HTML file” or “make a HTML artifact”.
The trick is knowing what you want the artifact to do and how you might use it. You may over time make a skill, but for now I’d suggest just prompting from scratch to get a hang of how to use it in different cases.
Use Cases
To make this more concrete, I’ve made many different HTML files for different use cases… continue on X ↗
♥ 17.8K · ⟲ 2.3K · 👁 14.8MView on X ↗