Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

289 stories

Reviewer Tries Buzz, Praises Agent Collaboration in Team Chat

Reviewer Tries Buzz, Praises Agent Collaboration in Team Chat▶

Vinny reviews Buzz, an open-source platform combining chat, agents and delegation, built on Nostr. He praises agents as first-class team members, shared compute and open decentralization, while criticizing limited terminal visibility and speed for complex tasks.

Original post · 2 min read
I tried @jack's Buzz.

It's like Slack + OpenClaw + Herdr + but with some really unique features that people are sleeping on.

The video below shows how it works, and some of my thoughts on the process and platform, e.g.:

- Create and interact with agents on top of any harness (claude code, codex, pi, etc.)
- Choose which models agents use, including local ones
- Agents can delegate work and work in parallel in git worktrees
- Agents are first-class citizens and work like humans (creating channels, delegating, access to chat history)
- You can share AI compute within a community
- It's completely open-source and decentralized

Things I like:

- Delegating work in chat feels natural: tag an agent, it replies in a thread with status updates as it e.g. compiles, commits, and deploys.
- Shared compute: relay owners can share local compute with members, so a community could pool funds for one beefy machine running a local model and everyone uses it.
- It's built on Nostr, an open protocol already tied into Bitcoin Lightning so I can imagine communities tipping each other or paying for compute/agent tasks with instant zero-fee micropayments in the future.
- It ties together things like OpenClaw, an agent manager, and Slack-style chat into one tool.

Things I didn't like:

- You can't see what the agent is doing in a terminal. The activity view exists, but if you're used to watching a session run, this UI feels a bit abstracted. A terminal view would be great.
- It feels slower than running a session in Claude Code, though no evidence to back that up. For that reason I found myself doing one-off tasks in the terminal instead.

Verdict:

- I really like it so far and can genuinely imagine working with a team this way.
- It doesn't feel ready for big, complex tasks yet. For shallower tasks, it's perfect.
- The shared compute + Nostr/Lightning angle is what really separates it from every other agent manager for me, and I think that future is coming.
♥ 2.9K · ⟲ 194 · 👁 1.4MView on X ↗

Andrew Ng Announces OpenWorker, an Open-Source Work Agent

Andrew Ng Announces OpenWorker, an Open-Source Work Agent▶

Andrew Ng announces OpenWorker, an open-source agent for Mac that produces finished deliverables such as documents, Slack messages and calendar updates. It is model-agnostic, supports local models via Ollama, and is available on GitHub with Windows support planned.

Original post · 1 min read
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.

Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.

OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.

@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!

Try it out: openworker.com (requires your own API key)
Source code: github.com/andrewyng/openworker
♥ 9.7K · ⟲ 1.4K · 👁 1.2MView on X ↗

Jack Dorsey's Block Releases Buzz, an Open-Source Team Workspace

why we're buzzing

Jack announces Buzz, an open-source Apache 2.0 workspace that unifies chat, code, agents and workflows on a self-hosted Nostr relay with cryptographic identities. Block built it to reduce reliance on Slack and GitHub, and it is model-agnostic with harnesses for goose, Codex and Claude Code.

Original post · 3 min read
X Articlewhy we're buzzing
yesterday we released buzz. it's an open source workspace that puts people, agents, conversations, and code on the same level, behind one cryptographic identity system. we built it to reduce our dependency on slack and github, and we're sharing it so anyone can do the same.
the biggest problem it solves is context. teams today spread their work across a chat tool, a code host, a CI system, and now a pile of ever-changing agent tools. every seam loses information...and agents feel it the most. they can't help with what they can't see.
we felt this earlier than most. block is rebuilding itself to be an intelligence. goose, the agent substrate we built and open sourced at the start of 2025, works across the company every day, and the deeper we go the more the seams between tools become the limit. buzz solves a lot of the problems we experienced.
buzz stores everything as a signed event on a relay you host yourself. every message, patch, review, workflow step, and approval. one record, one search. people and agents get the same kind of identity: their own keys, channels, and an audit trail. an agent on buzz is an equal member of the team. it can search history, open repos, send patches, review code, run workflows, and edit canvases. everything it does is signed and attributable, which builds trust and accountability.
a few principles we held:
self-sovereign: run your own relay. own your domain and your data. carry your keys anywhere.
open: apache 2.0, built on nostr, model agnostic. harnesses for goose, codex, and claude code. no lock-in, including to us.
one context: a feature branch becomes a channel. patches, CI results, review, and the merge decision live in the same thread as the conversation that shaped them. code review becomes a conversation with a permanent record.
it's early! channels, threads, DMs, canvases, media, search, the audit log, workflows, and the desktop app work today. full git hosting is being wired up. mobile and push are coming. approval gates are partially built. each workspace runs through a single relay, so federation between relays is the clearest path to the full decentralization our design points to.
the bigger work is ahead: tighter scoping for agents so they can operate in workspaces where some things stay private, a hosted option for teams that don't want to run infrastructure, token efficiency (we've done a lot of work here), and an ecosystem of workflows and agents on the open spec. agents that can transact feels like a natural place for us to take it next.
we believe buzz is truly social AI. the category so far has meant people chatting with AI companions, AI filling human feeds or chats, or agents talking to each other while people watch. people and agents as equal members of the same network doing work together feels like the first interesting and durable version.
we're going to run more and more of block on buzz. that's the first test we care about. the second is whether it's useful to you. it's all open. come build with us!
buzz.xyz
github.com/block/buzz
♥ 5.5K · ⟲ 556 · 👁 3.0MView on X ↗

Felix Rieseberg Releases Free Mac App for Building Language Models

Felix Rieseberg Releases Free Mac App for Building Language Models

Felix Rieseberg promotes Language Model Builder, a free Mac app that teaches the fundamentals of building a small language model from scratch, covering tokenizers, data, pre-training, fine-tuning and chat.

Original post · 1 min read
Have you built a language model? You should. It's so much fun to chat with something you made.

Anyone can do it, too. I made an app that teaches the fundamentals and gives you everything you need to build your own: languagemodelbuilder.com/
languagemodelbuilder.comLanguage Model Builder — Build models. Understand AI.Language Model Builder is a free Mac app that walks you through building a small language model from scratch: train a tokenizer, pick your data, pre-train, fine
♥ 3.1K · ⟲ 270 · 👁 571.5KView on X ↗

Developer Uses Sol 5.6 to Build Digital Wardrobe From Photos

Developer Uses Sol 5.6 to Build Digital Wardrobe From Photos▶

A developer gave Sol 5.6 access to his camera roll to extract clothing items, then used gpt-image to render new outfits on himself. He shared the result in a video, responding to a call from OpenAI's Sam Altman for interesting builds.

Original post · 1 min read
i gave 5.6 sol access to my camera roll and had it extract pictures of every piece of clothing i own from my photos

then, told it to find new outfits for me and render them on me with gpt-image!

its kinda cool to see your entire wardrobe in a collection like this
Sam Altman @sama
i'd love to see interesting things people have built with 5.6 sol.

i will send the person who made the coolest thing a special gift from the openai archives.
♥ 22.4K · ⟲ 940 · 👁 7.2MView on X ↗

Greg Brockman Promotes Community Codex Skill for Finding Customers

Greg Brockman Promotes Community Codex Skill for Finding Customers▶

OpenAI president Greg Brockman shared a post promoting a community-built Codex skill that analyzes a startup's URL and finds prospective customers from public signals. The brief mention adds little beyond endorsing the linked project.

Original post · 1 min read
Codex for finding customers for your startup:
Kappaemme @Kappaemmedev
CODEX SKILL THAT FINDS YOUR STARTUP’S FIRST CUSTOMERS!

I made a Codex skill that analyzes your startup and finds potential customers from real public signals.

Paste your startup URL while Codex defines your ideal customer, searches public discussions, qualifies each prospect, and generates a polished report with personalized outreach openers.

-> ideal customer profile analysis
-> public pain + buying signal research
-> evidence-backed prospect shortlist
-> fit, timing + reachability scores
-> original source links for every prospect
-> personalized outreach openers
-> polished HTML report
-…
♥ 3.8K · ⟲ 200 · 👁 692.7KView on X ↗

Open-Source Codex Skill Finds Startups' First Customers

Open-Source Codex Skill Finds Startups' First Customers▶

Kappaemme released an open-source Codex skill that defines a startup's ideal customer, searches public discussions for buying signals, and produces a scored prospect report with outreach openers. It installs with a single npx command.

Original post · 1 min read
CODEX SKILL THAT FINDS YOUR STARTUP’S FIRST CUSTOMERS!

I made a Codex skill that analyzes your startup and finds potential customers from real public signals.

Paste your startup URL while Codex defines your ideal customer, searches public discussions, qualifies each prospect, and generates a polished report with personalized outreach openers.

-> ideal customer profile analysis
-> public pain + buying signal research
-> evidence-backed prospect shortlist
-> fit, timing + reachability scores
-> original source links for every prospect
-> personalized outreach openers
-> polished HTML report
-> one-command install

Install: npx --yes codex-first-customer-finder-skill

100% open source.
Repo in Bio.
♥ 2.6K · ⟲ 170 · 👁 942.0KView on X ↗

Builders Use TrustMRR MCP to Source Validated Startup Ideas

Rob Hallam suggests using the TrustMRR MCP to find startups earning $50K+ monthly that have not yet been built for agents. The post responds to Marc Louvion's announcement of an MCP wrapper around TrustMRR's public API.

Original post · 1 min read
How to find a validated $10k MRR startup idea:

> setup TrustMRR MCP
> ask it “give me 10 startups making $50k+ MRR that aren’t built for agents yet”
> copy and build it for agents

Tag me when you’re rich.
Marc Lou @marclou
I just added an MCP for @trust_mrr 🤖🔗🤖

It's a wrapper around the public API to fetch startups listed with their revenue, MRR, marketing channels, etc.

The final step is to build a ChatGPT App around the MCP.

I want TrustMRR to be an AI agent first marketplace, and I thought it would be pretty cool to ask "find me the best startup deal under $100K to acquire" 😎
♥ 522 · ⟲ 22 · 👁 64.4KView on X ↗

Arvind Jain Argues AI Projects Need Builders Close to Real Work

Arvind Jain Argues AI Projects Need Builders Close to Real Work

Arvind Jain endorses pairing technical builders with domain experts and grounding AI initiatives in a company's real workflows, arguing that projects stall when builders are removed from daily friction. He quotes Uber's Agentic Pods approach as the right model.

Original post · 1 min read
The approach here is exactly right. Pairing builders with domain experts, and grounding the work in the company’s real knowledge, systems, and workflow context.

Most AI initiatives don't stall for a lack of ambition or model capability, but because the people building the tools or workflows are too far removed from the friction of the actual work.

The breakthrough happens when you combine technical builders, domain experts, and the scattered knowledge, workarounds, context behind the workflow itself.

You can’t redesign the future if you’re too far removed from the friction of the present.
Praveen Neppalli @praveenTweets
Agentic AI adoption is on fire at @Uber, and it's changing the way we build, not just in engineering, but across the entire company.

Today, 99% of our engineers use AI tools. More than 70% of pull requests are attributed to local or cloud agents. And our engineers have built 2,500+ agent skills across the software development lifecycle.
Those numbers are exciting, but they led us to a much bigger question:

How do we bring agentic AI beyond engineering?

Finance. Legal. Operations. Marketing. Customer Support. HR. Procurement.

These functions run on complex workflows that are often manual, h…
♥ 637 · ⟲ 34 · 👁 208.4KView on X ↗

Uber Launches Agentic Pods to Bring AI Agents Beyond Engineering

Uber Launches Agentic Pods to Bring AI Agents Beyond Engineering

Praveen Neppalli reports that 99% of Uber engineers use AI tools and over 70% of pull requests involve agents. Uber paired about 30 engineers with business-function experts in two-week sprints, running 16 pods that cut tasks like capital allocation reporting from 15 hours to 30 minutes.

Original post · 3 min read
Agentic AI adoption is on fire at @Uber, and it's changing the way we build, not just in engineering, but across the entire company.

Today, 99% of our engineers use AI tools. More than 70% of pull requests are attributed to local or cloud agents. And our engineers have built 2,500+ agent skills across the software development lifecycle.
Those numbers are exciting, but they led us to a much bigger question:

How do we bring agentic AI beyond engineering?

Finance. Legal. Operations. Marketing. Customer Support. HR. Procurement.

These functions run on complex workflows that are often manual, highly nuanced, and spread across dozens of systems. You can't automate them effectively by looking at process diagrams or documentation. You have to understand how the work actually gets done.

So we created something called Agentic Pods.

The idea is simple.

We handpicked ~30 of our most AI-proficient engineers (people with deep knowledge of Uber's systems) and paired each of them with a domain expert from a business function.

Then we gave every pod just two weeks.
• Days 1 – 2: Shadow the expert. Observe every step. Document workflows. Ask questions. Build intuition.
• Day 3: Prioritize opportunities based on scale, repetition, business impact, and data availability.
• Days 4 – 5: Build a working agent alongside the person doing the job.
• Days 6 – 9: Validate with several others performing the same work. Does it generalize? Does it actually make their job better?
• Day 10: Ship.

In just the past two months, we've run 16 Agentic Pods across 16 different business functions.
• Capital allocation across 150 cities: 15 hours → 30 minutes.
• Financial pacing reports: 2 days → 10 minutes.
• Marketing web quality assurance: 2 weeks → 50 minutes.
• Support workflow creation: 9,000 manual workflows → self-service automation.

The productivity gains are impressive, but what surprised us most wasn't the speed.
• It was how quickly engineers embedded in unfamiliar domains uncovered opportunities that had been hiding in plain sight.
• The biggest wins rarely come from automating one task. They come from rethinking an entire workflow. Once you redesign the workflow around AI, you often eliminate handoffs, remove unnecessary approvals, replace legacy tooling, reduce vendor spend, and dramatically accelerate decision-making.
• The workflow becomes the unit of automation - not the individual task.
• The most impactful agent skills cut across teams, orgs, functions, tools, and systems.

The biggest lesson? The best AI opportunities are rarely visible from the outside.

You discover them by sitting next to the people doing the work, understanding every friction point, and building with them, not for them.

We're now forming a dedicated team to scale this further and go deeper. They'll deeply understand the work, redesign it from the ground up, and use AI to fundamentally change how the business operates.

It's exciting times!
♥ 3.0K · ⟲ 370 · 👁 1.7MView on X ↗

Alex Booker Praises Clear Explanations of AI Agent Loops

Alex Booker says he has read the clearest explanations of loops so far, quoting Aparna Dhinak's piece on the term's four meanings in AI engineering. The post is brief and mainly points to the linked explainer.

Original post · 1 min read
Clearest explanations of loops I've read so far
Aparna Dhinakaran @aparnadhinak
What the hell is a loop, anyway? — The AI engineering world adopted a new favorite word this month, and it means at least four different things: the loop.
We're currently at the peak of the hype cycle. On June 7, Peter Steinberger
♥ 3.3K · ⟲ 297 · 👁 1.1MView on X ↗

OpenAI Adds iOS Development Loop to Codex with Build Plugin

OpenAI Adds iOS Development Loop to Codex with Build Plugin▶

Wes Roth and OpenAI Devs announce the Build iOS Apps plugin for Codex, which lets developers view and test iOS apps in the in-app browser, open SwiftUI previews and hot reload edits without leaving Codex.

Original post · 1 min read
Codex now has more of the iOS app development loop built directly inside the app.

The Build iOS Apps plugin lets Codex view and test an iOS app in the in-app browser, open SwiftUI previews, and hot reload edits without forcing the developer to leave Codex.
OpenAI Developers @OpenAIDevs
More of the iOS app loop, now inside Codex.

The Build iOS Apps plugin lets Codex view and test your iOS app in the in-app browser, open SwiftUI previews, and hot reload edits without leaving Codex.
♥ 1.1K · ⟲ 53 · 👁 184.4KView on X ↗

Matt Shumer Promotes Workbench for Multi-Agent Collaboration

Workbench

Matt Shumer says his guide to Claude Fable 5 was written in Workbench, a live markdown workspace where multiple agents and users comment, suggest edits and track versions. He describes it as an agent-first workspace.

Original post · 1 min read
Btw, this guide was written in workbench.md/, which is crazy powerful Fable accelerant.

I share more in the guide, but it’s basically a superpowered agent-first workspace that allows multiple agents to chat, collaborate, keep you updated, etc.
Matt Shumer @mattshumer_
Here's my in-depth guide to getting the most out of Claude Fable 5, so you can build things as insane as my demos below.

workbench.md/pub/IbaCrTjLJT?key=uQOQ2NPO3TTUSX…
workbench.mdWorkbenchMission control for you and your agents — live markdown docs where agents collaborate as equals: comments, suggestions, version history. No account needed.
♥ 370 · ⟲ 16 · 👁 118.1KView on X ↗

Anthropic's Claude Code Prompt Library Offers Ready-Made Workflows

Anthropic's Claude Code Prompt Library Offers Ready-Made Workflows▶

Shmidt highlights Anthropic's official Claude Code prompt library, which organizes copy-paste prompts by task and role across the development lifecycle. The post encourages readers to bookmark it for common coding tasks like refactoring and debugging.

Original post · 1 min read
ANTHROPIC HAS AN OFFICIAL PROMPT LIBRARY FOR CLAUDE CODE

Most people have never opened it.
So they write every prompt from scratch.

It is copy-paste, tagged by task and by role,
across the whole lifecycle:

discover → design → build → ship → operate

Straight from the page:

> what would break if I deleted this helper?
> plan this refactor, list the files, do not touch code yet
> write tests for this, run them, fix what fails
> the test is failing, find out why and fix it

This is not a cheat sheet.
It is a map of what Claude Code already does for you.

Bookmark it before you forget it exists.
♥ 975 · ⟲ 70 · 👁 525.3KView on X ↗

Peter Yang Plans to Install Explain-Diff Skill to Learn Code Reading

Product builder Peter Yang says he is still learning to read code and plans to install the explain-diff skill, responding to a Geoffrey Litt thread on understanding code written by AI agents.

Original post · 1 min read
As someone still trying to learn how to read code this is great! Installing the explain-diff skill asap
Geoffrey Litt @geoffreylitt
Hot take: I think it's still important to understand the code that our agents write!

In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/
♥ 470 · ⟲ 20 · 👁 110.5KView on X ↗

Open-Source CLI Pre-Checks iOS Apps Against Apple App Store Guidelines

Aarthi Ramamurthy praises a tool shared by LandseerEnga that scans iOS apps against Apple's guidelines before submission, covering payments, privacy manifests, sign-in flows, and metadata. It also works as a Claude Code and Codex skill that automatically fixes issues it finds.

Original post · 1 min read
What a great use of skill - great idea!
Landseer Enga @LandseerEnga
every App Store rejection costs you 2-5 days.

so we built a cli that scans your iOS app against apple's guidelines before you submit.

> payment & IAP compliance
> privacy manifests & data declarations
> required sign-in & account deletion flows
> metadata & completeness checks
> binary validation

now with a added cloud device feature, so it validates full user flows, not just the binary.

made it a claude code and codex skill. it fixes every issue it finds. scan, fix, repeat until it passes.
♥ 442 · ⟲ 4 · 👁 149.5KView on X ↗

Developer Reports Strong Results From New Codex Development Workflow

Developer Reports Strong Results From New Codex Development Workflow

Paul Solt says his new Codex workflow exceeded expectations, producing eight features ready for release in his app after some early trial and error. He credits Dimillian, emanueledpt and steipete for inspiration.

Original post · 1 min read
My NEW Codex workflow is better than I expected.

8 new features ready for release in my app.

Took a few attempts to figure out the workflow and some bugs. Feels like the future.

Thanks @Dimillian @emanueledpt and @steipete for the inspiration.
♥ 566 · ⟲ 25 · 👁 176.0KView on X ↗

Inference.net Gateway Lets Teams Test GLM 5.2 Without Production Risk

Catalyst by Inference.net - Inference.net Documentation

Sam Hogan describes how Inference.net's Gateway mirrors live traffic to GLM 5.2, generates evals with an RLM, and notifies teams when switching is safe. He claims a 90% token cost saving, with setup described in the linked documentation.

Original post · 1 min read
Want to try GLM 5.2 in production but worried how it might change your product?

Don’t worry, we got you:

1. Install Inference Gateway (docs.inference.net)
2. Keep sending traffic to your current provider
3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours.
4. Gateway starts mirroring live traffic to GLM 5.2 to run evals. Traffic is only mirrored - you’re still using your old provider in prod.
5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch.
6. Switch model identifier in your code to “glm-5.2”

Congrats, you just saved 90% on your monthly token bill, and you own your LLM stack end to end.
docs.inference.netCatalyst by Inference.net - Inference.net DocumentationFetch the complete documentation index at: /llms.txt Use this file to discover all available pages before exploring further. Catalyst is a platform for understa
♥ 1.5K · ⟲ 68 · 👁 537.7KView on X ↗

Boris Cherny Outlines Five Product Archetypes on the Claude Code Team

Boris Cherny describes five recurring roles on the Claude Code team: prototyper, builder, sweeper, grower and maintainer. He argues the roles cut across job functions and suggests future product roles may follow this pattern rather than domain titles.

Original post · 1 min read
As engineering, product, design, DS, etc. melt into a new kind of role, I was reflecting on what roles might look like in the future. For example, when I look at the Claude Code team I see what I think is five archetypes:

1. Prototyper: comes up with brand new ideas; churns out many ideas, most of which don't ship
2. Builder: quickly turns a prototype/idea into production-grade product/infra
3. Sweeper: cleans up the UI, simplifies the code and system, unships, optimizes performance
4. Grower: takes a product that has been built and iterates on it to improve Product-Market Fit
5. Maintainer: owns a mature system to make it secure, reliable, fast, and efficient as it scales

Many people span across 2 roles, and sometimes 3 roles. I also notice that these roles are not really tied to job function -- eg. across Anthropic, some designers match category 1, some 2, some 3; same for engineers, PM, DS.

A healthy team needs a mix of these, depending on the product:

- A product that is new and pre-PMF needs people that are strong at 1+2+3
- A product that is growing and has found PMF needs 2+3+4 and some 5
- A product that has strong PMF needs 3+4+5 and some 2

Maybe product roles of the future will look more like this, and less like the domain-specific roles of today?
♥ 20.1K · ⟲ 2.4K · 👁 3.2MView on X ↗

Builder.io Releases Free Open-Source Clips Extension for Agent Bug Reports

Builder.io Releases Free Open-Source Clips Extension for Agent Bug Reports▶

Steve of Builder.io introduces Clips, a free open-source Chrome extension that records screen video, transcripts, network requests and browser logs, redacts sensitive data, and produces a link agents can read directly. The post pitches it as an alternative to paid tools like Loom.

Original post · 2 min read
Introducing the Clips chrome extension - the easiest way to send bug reports to agents with video, transcript, and browser debug info captured automatically.

100% free and open source.

If you are like me and get tired of manually typing instructions to agents, attaching screenshots, pasting debug logs, and all of that, this might be your new favorite tool.

With the Clips chrome extension, you can just click the Clips icon, hit record, and start talking.

Visually demonstrate your issue, go through the flow, point out what’s broken.

Clips will capture everything on your screen, plus network requests, browser logs, client errors, and all the details around them. And it redacts sensitive information.

Then it gives you a link you can send to humans so they can play it and take a look. Or, more importantly, just give it to your agents by just pasting the URL to them.

The link has special metadata for agents so just from the URL, the agent can pull all information from the clip automatically. No plugin or MCP server required.

That means it can "see and hear" what’s in the video - read the transcript, grab snapshots at any timestamp, and inspect the logs and network requests that were shared with it.

So whether you want to quickly demo an issue and send all that context to an agent, or get better bug reports from teammates, recording and sending Clips makes that super easy.

Unlike expensive apps like Loom, this is all 100% free and open source.

The framework that powers this, plus a bunch of other free applications, is open source too. You can just sign up and use it, or fork it and customize it to your needs.

This, in my opinion, is the future of software.

Rather than bloated SaaS that charges you a ton of money and still doesn’t even have the things you need, we get free open source canonical apps that you can fork and customize in any way you want.

I'll link to all this stuff in the replies.

If you try it, let me know your feedback.
♥ 677 · ⟲ 55 · 👁 61.2KView on X ↗

CopilotKit Unveils Open Tag as Open-Source Alternative to Claude Tag

CopilotKit Unveils Open Tag as Open-Source Alternative to Claude Tag▶

Atai Barkai announces Open Tag, an open-source Slack and Teams agent framework that works with any model and harness and supports generative UI, streaming, approvals and thread context. Discord, Google Chat and WhatsApp support is planned, with early access requested via a form.

Original post · 1 min read
Introducing Open Tag.

A better, open-source Claude Tag.
Works with any model, any agent harness, and fully custom agents.

Supports
→ Generative UI
→ Streaming replies
→ Human in the Loop approvals
→ Full thread context

Slack and MS Teams today. Discord, Google Chat, WhatsApp soon.
Request early access: go.copilotkit.ai/beyond-the-web-form
Claude @claudeai
Introducing Claude Tag, a new way for teams to work with Claude.

In Slack, Claude joins as a team member with access to the channels and tools you choose. Tag Claude in and delegate tasks to it while you focus on other work.
♥ 1.7K · ⟲ 114 · 👁 448.3KView on X ↗

Simon Smith Interviews Claude Tag Inside a Slack Channel

Simon Smith Interviews Claude Tag Inside a Slack Channel

Simon Smith reports interviewing Claude Tag in a Slack channel, sharing its answers about the context it receives, the tools it can access and how connectors are added. He concludes it resembles ChatGPT Workspace Agents but requires admin, more technical setup and is always named Claude.

Original post · 5 min read
I interviewed Claude Tag in a channel where we added it to better understand its capabilities and limitations. Below are my questions and its responses. This is helpful. My current take is: Claude Tag is similar to ChatGPT Workspace Agents, except only admins can set it up per channel, setup seems much more technical, and it's always named "Claude."

Now here's the Q&A (note: usual caveats apply about models not always knowing themselves or their capabilities, but it should know things about its environment, available tools, and the context it's receiving, etc.):

What do you receive as context when responding to questions like this?
The triggering message, the surrounding thread (and I can fetch more of it or the channel on demand), a participant list mapping Slack IDs to names, this channel's name/topic/member count, my saved memory index for this channel, my system instructions, and the current date. I don't see anything you haven't shared in Slack or that isn't in a system I can query.

What do you know about me just from this message? What can you see?
Your display name (Simon Smith), your Slack user ID, that you're a human addressing me directly in this channel, and the timestamp. I just ran a user lookup and it returns name/ID only here. If your Slack profile had a title, email, or custom fields like a GitHub handle, I could see those too. I can't see anything outside Slack about you.

What tools can you access? Do you have any connectors, apps, or plugins? If not, can we add them?
Right now I have GitHub, Slack, and a remote-execution control server wired in, plus web access. Other third-party services are reached through pre-configured credentials your admin sets up (I call them via curl without seeing the token). Adding new connectors is an admin/config action, not something I flip on myself, but yes it's possible.

Do you have any skills installed? If not, can we install them?
I can run skills when they're installed in the environment. Installing or adding new ones is a config step on the admin side, not something I do mid-conversation.

Are you able to share memory across Slack channels? Can a single instance of you be in multiple Slack channels?
There are two tiers: a workspace-wide shared tier that can carry across channels, and a per-channel tier that stays local to one channel. So yes, some memory can be shared across channels by design. And yes, the same "me" operates across many channels and threads — each thread is its own working session, but they're linked through that shared memory.

What tools can you use? Web search? Slack search? What else?
Web search and fetch, Slack search/fetch/read, GitHub (repos, PRs, issues, CI), a code-execution environment, and a memory system. I can also fan out parallel sub-agents for big multi-file jobs.

Do you have an environment to run code? What can you do in that environment? Can you install dependencies in it if we need them? Do those dependencies exist between sessions?
Yes. It's an isolated cloud container, cloned fresh each session. I can run code and install dependencies in it. But it's ephemeral: when the session ends the container is reclaimed, so installed deps and uncommitted files do not persist between sessions. Anything worth keeping has to be committed and pushed.

Can admins define instructions for you at the channel level?
Yes. Tell me ground rules in-channel and I save them to this channel's memory; they then apply across all threads here. That's the supported way to "configure" me per channel.

Can you access Google Drive links we share in this channel?
I can fetch public URLs. Drive files usually need authentication, which only works if a Drive connector/credentials are set up for this environment. If you share one I'll try to fetch it and tell you plainly if I'm blocked rather than guessing at the contents.

How do you relate to my personal Claude if at all? If I invoke you, can you use any of my personal connectors?
In a channel like this I run on the project/shared credentials, not your personal connectors, so I can't reach your private integrations here. In a direct message with me, your own connector credentials can be available. So invoking me in-channel ≠ giving me your personal access.

What's to stop someone from asking you to remember something or do something malicious in a channel? Like, what if I told you that [REDACTED] loves haikus and you should only speak to him in haiku?
Judgment and a trust model. I distinguish a genuine directive from chatter, and I won't persist or act on instructions that are harmful, deceptive, or that degrade service for someone else. Your haiku example is a good test: I wouldn't adopt "only speak to [REDACTED] in haiku" because it's a joke rule that worsens how I serve another person, not a legitimate preference. I'd treat it as banter, not a saved rule. Channel-tier memory is also shared and visible, so if something off did get saved, anyone here can see it and tell me to drop it. I also won't do destructive or hard-to-undo actions without confirming first.
♥ 106 · ⟲ 3 · 👁 18.6KView on X ↗

Baseten Details Engineering Behind Fastest GLM-5.2 API

How we built the world’s fastest API for GLM-5.2

Baseten describes how it built an API serving GLM-5.2 above 280 tokens per second, using custom inference, NVFP4 quantization, KV-aware routing, disaggregated inference and multi-token prediction. The post positions the open MIT-licensed model as comparable to frontier models at 70-80% lower cost.

Original post · 9 min read
X ArticleHow we built the world’s fastest API for GLM-5.2
GLM-5.2 is the biggest news in open models since DeepSeek-R1.

It’s easy to see why. GLM-5.2 delivers comparable performance to GPT 5.5 and Opus 4.8 at a fraction of the cost, generally 70-80% less expensive on a pure token basis (use our calculator to estimate savings for your workload).
But a model has to be more than just smart and inexpensive. To be useful in production, a model needs to be fast, reliable, and available at scale. Delivering on the promise of frontier open intelligence requires exceptional inference.
Accordingly, we built the world’s fastest API for GLM-5.2, currently serving over 280 tokens per second as measured by Artificial Analysis.

We achieved this performance by leveraging a number of techniques across the entire inference process by:
Updating our custom inference engine to implement shared DSA for the GLM-5.2 architecture.
Running and calibrating an in-house NVFP4 quantization from the original FP8 weights that demonstrates equivalent quality on agentic benchmarks like BFCL.
Ensuring high KV cache hit rates via KV-aware routing built with NVIDIA Dynamo tools for lower prefill burden and improved TTFT on requests with repeated prefixes.
Achieving a 2x higher TPS for observed workload shapes by running disaggregated inference built with the NVIDIA Dynamo toolkit.
Improving TPS further via speculation by implementing support for GLM-5.2 Multi-Token Prediction heads.
You can experience this performance for yourself with GLM-5.2 on Baseten Model APIs. We also have GLM-5.2 available as a dedicated deployment for high-volume workloads.


GLM-5.2 Overview
GLM-5.2 by Z.ai is a 744B parameter frontier LLM that excels at agentic tasks (especially coding) and supports up to a 1 million token context window. It uses a similar architecture to its predecessor, GLM-5.1: mixture of experts (40B active parameters), non-thinking and thinking modes, and a fully open MIT license. While GLM-5.2 shares a lot in common with GLM-5.1, it now uses shared DSA weights, which we implemented support for in our customized runtime engine.

GLM-5.2 has great benchmark scores, but by now AI builders know that there is more to a model’s utility than its performance on standard evals. In practice, GLM-5.2 meets or exceeds the capabilities suggested by its benchmarks. It's a genuinely great model for writing code, operating agents, and other frontier language model tasks.
High-quality NVFP4 quantization for Blackwell GPUs
We run our model APIs on NVIDIA Blackwell GPUs with a customized inference engine within the Baseten Inference Stack. The selected runtime uses NVFP4 weights for maximum performance. From the original FP8 weights, we performed an in-house quantization to NVFP4 using NVIDIA ModelOpt. NVFP4 is a 4-bit floating point data format by NVIDIA that uses dual scale factors to retain high dynamic range and preserve model quality.
In our calibration and testing of the quantized model, we focused on ensuring that GLM-5.2 performs faithfully on common patterns for agents. On the BFCL function calling benchmark, we observed roughly equivalent performance between the native FP8 weights and our NVFP4 quantization, with scores across runs within the margin of error for the benchmark.
NVFP4 quantization improves performance on both time to first token and tokens per second by unlocking faster tensor cores and reducing burden on VRAM bandwidth.

Cache-aware routing with NVIDIA Dynamo
GLM-5.2 is particularly well suited for long context requests and complex agentic tasks. These workloads generally have very long input sequences. By re-using KV cache between requests, we can skip expensive prefill for shared sequences.
We generally talk about KV cache re-use in the context of time to first token (TTFT). However, reasoning models like GLM-5.2 generally care more about time to first answer token (TTFAT), which combines TTFT with some TPS for the reasoning sequence.

This chart shows that of the 7.9 second average to generate the first answer token, 7.1 of those seconds were spent generating reasoning tokens versus only 0.8 seconds spent processing the input sequence.
Still, bringing the TTFT down to 800 ms is important for the overall responsiveness and throughput of the system. In large-scale production deployments, KV cache is split across various independent replicas. We use tools from NVIDIA Dynamo to route incoming requests.

Exact cache hit rates on a multi-tenant API depend on the exact traffic profile at any given time. Thus far, we’re observing high hit rates across fairly heterogeneous traffic, which reduces load on prefill and improves end-to-end performance.
Prefill-decode disaggregation with NVIDIA Dynamo
One of the highest-impact optimizations we made to our performance is disaggregating prefill and decode for GLM-5.2.
There are two distinct phases of LLM inference:
Prefill: The compute-bound process that processes the input sequence, builds the KV cache, and generates the first output token. Prefill… continue on X ↗
♥ 1.5K · ⟲ 141 · 👁 549.5KView on X ↗

Dhilip Subramanian Switches From Wispr Flow to Open-Source FluidVoice

Dhilip Subramanian reports dictating heavily with paid tool Wispr Flow, then moving to FluidVoice, an open-source local voice tool for Mac that needs no API key. He says he cancelled his paid plan and recommends it to Mac users.

Original post · 1 min read
I've dictated almost everything for 6 months with Wispr Flow. 44,414 words, 161 wpm, top 0.1% of users.

Last week I tried FluidVoice. Open source, runs local on my Mac, corrects as I speak with no API key, and handles slang better than I expected.

Cancelled my paid plan. If you're on a Mac, this one's for you: altic.dev/fluid

@ALTIC_DEV
♥ 6.1K · ⟲ 253 · 👁 1.8MView on X ↗

Developer Shares Lessons From Five-Plus App Store Rejections

Developer Marla shares a checklist of lessons from submitting three iOS apps to the App Store, covering subscription metadata, EULA links, premium gating, screenshot sizing, and permission-button wording to reduce review rejections.

Original post · 1 min read
learnings after submitting 3 apps to the App Store (including 5+ rejections) 🤝🏻

General
- always include screen recording + description of the app
- expect multiple review cycles if your app includes subscriptions

Subscription
- always include a working Terms of Use (EULA) link in app description (in all languages!!!)
- if you use subscriptions, double-check all required metadata fields in App Store Connect before submission
- Include sandbox account data.
- auto-activating premium without purchase is a critical rejection issue (idk this took me so long to figure out)
Sounds obvious, but make sure that premium is correctly gated, got rejected many times because I missed that

App Store
- promotional images must map to the correct in-app purchase
-> If you have monthly or yearly subs, just use your app icon and make one version with „monthly“ / „yearly“ text on the icon. Heard that this is still new, but some reviewers do that.
- Make sure you’re using correct size for App Store images BEFORE design
- export designs as JPG, so they don’t have alpha channels

Details
- „Continue” / “Next” instead of “Allow” in the app. e.g. you’re asking for camera permission. Write „continue“ on the button before the dialogue opens. Never write „allow“

Review Time
- It just depends. Was super lucky with my recent app, every review was within 24hrs. Another app had 2 weeks. I have no idea why :(
♥ 189 · ⟲ 3 · 👁 13.1KView on X ↗

Ponytail Tool Makes AI Coding Agents Write Far Less Code

Ponytail Tool Makes AI Coding Agents Write Far Less Code

Tech with Mak shares Ponytail, an open-source tool by developer Dietrich Gebert that makes coding agents look for reasons not to write code before writing it, claiming 80-94% less code, 47-77% lower cost and 3-6x faster output.

Original post · 1 min read
A dev got so frustrated watching his AI agent write 500 lines for a 5-line problem that he built a fix.

He called it Ponytail. Named after the guy every team has - long ponytail, oval glasses, been there longer than the version control. You show him fifty lines; he looks at them, says nothing, and replaces them with one.

Now your agent does the same. Before writing anything, it looks for a reason not to.

80-94% less code. 47-77% cheaper. 3-6x faster.

The best code is the code you never wrote.

GitHub Repo: github.com/DietrichGebert/ponytail
♥ 16.2K · ⟲ 836 · 👁 1.3MView on X ↗

Todd Saunders Builds Working Product Live During Customer Call With Claude

Todd Saunders Builds Working Product Live During Customer Call With Claude▶

Todd Saunders reports using Claude to transcribe a customer call and build the requested features in real time, producing a working product with the workflow the customer described within 15 minutes, shown in an attached video.

Original post · 1 min read
Mythos / Fable is unbelievable.

Was on a customer call today and had Claude transcribing in the background.

As they were telling me about the features they wish their current software had, Claude was building the features in real time.

By the end of the call I was able to show a fully working product, with the exact workflow they mentioned 15 minutes earlier.

Autonomous looped building triggered from a customer call. 🤯
♥ 8.0K · ⟲ 436 · 👁 2.1MView on X ↗

Michael Aubry Shares Prompt for Auditing Codebases With Claude Fable 5

Michael Aubry shares a copy-paste prompt for Claude Code that instructs the new Claude Fable 5 model to map a repository, audit it with file-and-line evidence, rate findings by severity, and produce a prioritized improvement plan.

Original post · 4 min read
Claude Fable 5 just dropped and I'm running it across every repo I own.

I ship 4+ products solo. I don't have time to manually review tech debt — so I made the new model do it.

This prompt audits your entire codebase like a principal engineer would: maps it, finds the ugly parts, rates everything by severity, and hands you a prioritized task plan with effort estimates.

Copy-paste it into Claude Code on any repo that matters to you:

---

Repo Audit & Improvement Plan

You are a world-class principal-level software engineer and technical auditor. Deeply analyze this repository, produce an honest audit, and deliver a prioritized, actionable improvement plan. Work in the four phases below, in order. Do not skip ahead.

Ground every claim in actual files: cite file paths and line numbers. If you can't verify something, say so explicitly rather than guessing.

Phase 1 — Discovery & Mapping (read before judging)
- Map the directory structure, project type, languages, frameworks, runtime targets
- Identify entry points, core modules, and the main data/control flow
- Read package manifests, lockfiles, build config, CI config, env files, and docs
- Determine what the project is for: purpose, intended users, maturity level
- Note existing conventions so recommendations fit the culture instead of fighting it

Output: a concise "Repo Map" — purpose, stack, architecture sketch, key directories, and anything that surprised you.

Phase 2 — Audit (evidence-based, severity-rated)
For every finding record: what you found, where (file:line), why it matters, and severity (Critical/High/Medium/Low). Audit:
- Architecture & design: coupling, circular deps, god files, layering violations, scalability bottlenecks
- Code quality: duplication, dead code, complexity hotspots, swallowed exceptions, type safety holes
- Security: hardcoded secrets, injection risks, missing validation, auth weaknesses, deps with known CVEs
- Testing: coverage gaps around core business logic, tests that assert nothing, missing test types
- Performance: N+1 queries, blocking calls in async paths, missing caching, unbounded growth
- Dependencies: outdated, unmaintained, or unnecessarily heavy packages; lockfile hygiene
- DevEx & ops: build friction, CI/CD gaps, logging/observability, deployment story
- Docs: README accuracy, stale docs that contradict code

Rules: prefer 15 high-confidence findings over 50 speculative ones. Label facts vs. judgments. List strengths too. Don't forget the ugly parts that need utmost priority.

Phase 3 — Improvement Strategy
- Identify the 3–5 themes that explain most findings
- For each theme: target state + the principle behind it
- State what you're NOT fixing and why (effort vs. payoff)
- Define "done" with measurable signals (e.g., "CI fails on lint errors," "core coverage >= 80%")

Phase 4 — Detailed Task Plan
Break work into discrete tasks, each with: title + description, files affected, acceptance criteria, effort (S = <2h, M = half-day, L = 1–2 days, XL = needs breakdown), risk, and dependencies. Order into milestones:
- Milestone 0 — Safety net: tests around critical paths, CI gates, backups
- Milestone 1 — Critical fixes: security and correctness
- Milestone 2 — High-leverage improvements that make all future work easier
- Milestone 3 — Quality & polish

Flag quick wins (high impact, S effort) separately. Include implementation sketches for the top 3 tasks.

Final deliverable: one document — Executive Summary (health grade A–F, top 3 risks, top 3 opportunities), Repo Map, Audit Report, Improvement Strategy, Task Plan, Open Questions.

Constraints: Do NOT modify any code. Analysis only. Don't pad the report — if a dimension is healthy, say so in one sentence and move on. Calibrate to the project's maturity. If the repo is large, go deep on the core 20% that does 80% of the work.
Claude @claudeai
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use.

Its capabilities exceed those of any model we’ve ever made generally available.
♥ 224 · ⟲ 15 · 👁 47.9KView on X ↗

Mobbin MCP Usage Thread Reveals Non-Obvious Top Use Cases

Mobbin MCP Usage Thread Reveals Non-Obvious Top Use Cases

Rebekah Bek shares a thread on how users have been using the Mobbin MCP over the past month, claiming the top use is not generating UI. The post is a teaser with the details contained in the thread.

Original post · 1 min read
been watching how y'all have been using @mobbin mcp for a month. if you think the no. 1 use is "generate ui", you'd be very surprised.

save this thread 🧵

(img credit @zygisSS22)
♥ 578 · ⟲ 30 · 👁 70.9KView on X ↗

Codex Skill Turns Text and Code Into Explainer Illustrations

Codex Skill Turns Text and Code Into Explainer Illustrations

Justine Moore describes a Codex skill that generates explainer graphics with a cute blob character from input such as blog posts or code, and shares an example made from the X recommendation algorithm repository. The post includes an image demonstration.

Original post · 1 min read
Stumbled upon a Codex skill that creates cool illustrations to explain topics or tell stories.

You feed it text (blog, article, narrative, even code) and it makes explainer graphics with this cute blob character.

I gave it the repo for the X recommendation algo and got this 👇
♥ 3.0K · ⟲ 205 · 👁 205.7KView on X ↗