Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Agents & Dev Tools

Coding agents, developer tools, workflows, open source

Jared Palmer Releases Kev Open Source Decision Model Family Based on Qwen3

Jared Palmer Releases Kev Open Source Decision Model Family Based on Qwen3

Jared Palmer announces Kev-0.6B, 4B and 8B, open source Apache 2.0 decision models built on Qwen3 with LoRA and a pointer head. He reports Kev-8B scores 79.6% out of domain versus 85.7% for Jev, and says the models are drop-in compatible with the TypeSafe API.

Original post · 1 min read
UPDATE: Kev-0.6B, 4B, and 8B are now available. Kev is a family of small open source Jev-like decision models you can train and run yourself.

This new family is based on Qwen3 using the same LoRA + small pointer head technique as before, but scaled up.

Out of domain, on data Kev never trained on: Kev-8B 79.6%, Jev 85.7%.

• Drop-in TypeSafe System One API; their SDK works with one `base_url` change
• Kev-4B serves on a 32 GB Mac in bf16: ~300 ms for five questions, ~40 ms on an H100
• Repeated documents hit a KV cache: 2-2.5x faster
• Apache 2.0 License. Kev-4B trains in 40 minutes on one H100. Kev-8B in 83 minutes.

Code, weights, evals: github.com/jaredpalmer/kev
Jared Palmer @jaredpalmer
Kev-0.5B: A tiny open source Jev-like decision model with a TypeSafe-compatible API based on Qwen2.5-0.5B that you can train and run on a MacBook Pro.

Model card and weights are available on GitHub

github.com/jaredpalmer/kev
♥ 2.7K · ⟲ 227 · 👁 291.9KView on X ↗

Astra AI Agent Edits Five-Minute Yosemite Vlog From 20 GB of Footage

Astra AI Agent Edits Five-Minute Yosemite Vlog From 20 GB of Footage▶

Shubhankar reports that an AI agent named Astra processed over 20 GB of 360-degree video, pulled soundtracks, and edited a five-minute vlog in Davinci Resolve in under an hour from a single prompt. He estimates the job would have taken him six or more hours manually.

Original post · 1 min read
My mind is blown!

Astra diligently went thru 20+ gigs of 360-video, cropped it from insta360-studio, downloaded and mixed a bunch of soundtraks, edited the whole vid in davinci resolve, and exported a 5-minute vlog, in under an hour.

from a simple prompt: "make a casey-neistat style vlog of my trip to Yosemite with my friends"

This would've taken me 6+ hours to edit. Holy shit, what a world.
Shubhankar @_shubhankar
agi test: astra is editing my vlog using davinci resolve and insta360 studio, report back soon
♥ 1.3K · ⟲ 40 · 👁 136.1KView on X ↗

Tesla Owner Replaces Subscription With Open-Source sunnypilot on Comma Four

Tesla Owner Replaces Subscription With Open-Source sunnypilot on Comma Four▶

Mike LaBarbera says he canceled his Tesla FSD subscription and installed a Comma Four device running sunnypilot, an open-source fork of comma.ai's openpilot, to get Level 2 driver assistance in his Model Y. A video accompanies the post.

Original post · 1 min read
My FSD subscription ended last January, and this isn’t Tesla Autopilot.

I installed a comma four in February, eight months of Level 2 ADAS I own outright in my Model Y. This is sunnypilot, an open source fork of @comma_ai's openpilot. No longer renting a closed system.
♥ 1.9K · ⟲ 63 · 👁 1.8MView on X ↗

Danielle Morrill Credits Claude Code With $10,000 in Savings

Danielle Morrill says Claude Code helped her cancel $4,480 in annual recurring charges and find $5,500 in one-time savings over about three hours. She describes doing it casually while eating and watching TV.

Original post · 1 min read
Claude Code helped me cancel $4,480/year in recurring charges and another $5,500 in one-time savings tonight over the course of 3 hours while eating a burrito and watching TV with my dogs

And I thought I was dialed in personal finance but whoa
♥ 152 · ⟲ 0 · 👁 30.1KView on X ↗

Opus 5.5 Builds Interactive Browser Game Assets With Blender and Three.js

Levelsio says Claude still struggles with sound, producing tinny synth effects, while praising a third iteration of a browser game built with Opus 5.5 using Blender, image generation and three.js. The quoted creator calls it a glimpse of the future of game development.

Original post · 1 min read
This is getting really good

The last thing Claude is bad at now is sound, it always does these "ticky" synth sounds that are way worse in quality than the visuals it now produces
Shikhar @xikhar
Third iteration with Opus 5.5 medium.

I am in awe. It turned blender, image-gen, and three.js into this beauty, which runs on your browser.

An understatement to say that Anthropic cooked. You are looking at the future of game dev.
♥ 1.6K · ⟲ 31 · 👁 336.2KView on X ↗

Field Notes From SpaceX AI Team Offer Multi-Bot Agent Playbook

Field Notes From SpaceX AI Team Offer Multi-Bot Agent Playbook▶

cat.png shares a free MIT-licensed repository of notes from three SpaceX engineers who shipped a product in 72 hours using Grok Bot agents. The notes cover a chief-of-staff bot, scoped specialist bots, proof-before-merge rules, cost figures and 40 documented antipatterns. The quoted post describes a 14-page guide on a growth lead building bots to replace his own work.

Original post · 2 min read
I still can't f**king get why nobody runs Grok Bot the way the people who built it do.

Three SpaceX engineers shipped a product from an empty repo in 72 hours, live on stream. 433 PRs.

I took notes all three days and turned them into a repo. Free, MIT.

The setup:

● One chief of staff. The only bot you talk to. You build every specialist through it, so it knows who does what.

● One bot, one job. Scope it like a job description. Split when the scope grows, not before.

● Drafts before sends. Write access is earned, not granted.

● No proof, no merge. The bot reproduces the bug, fixes it, attaches a video to the PR.

● Four human gates only: migrations, deploys, money, permissions. Everything else merges on its own.

The numbers nobody puts in the demo:

→ Support ticket: $1-2 with a reasoning pass. $0.20 once low-complexity tickets go to a script.

→ Full case-study deck: $20-30. By hand: 4-5 hours.

→ One team, one month, this stack: 2,500 PRs merged to prod.

→ Launch day: 4,000 games, 2,000 users, $0 revenue.

Cost is a property of your setup, not the tool. And autonomy is not a business model.

40 things broke on air. Every one is written down with the rule that came out of it.

Inside the repo:

> AGENTS.md for your root
> 9 playbooks
> 69 bot roles
> 40 antipatterns
> 14-page guide

The tools are not the moat. Everyone has them. The moat is a loop that decides what not to ship.

Save it before you wire your next agent 👇
github.com/unicodef1wn/grokbot-field-notes
unicode @unicodef1wn
SpaceX AI's growth lead built a bot to replace himself. I turned his session into a 14-page PDF.

He started with six bots. One job each.

→ researcher, product marketer, website ops, performance, analyst
→ each one takes the handoff from the one before it
→ he leaves comments in the doc, the bot reads them and redrafts

Then he built the seventh bot.

It read the other six. It read every place he had to step in and fix them.

Then it ran three campaigns and told him nothing until they were done.

His words: "I don't want to talk to any of the other bots."

Also in the PDF:

→ 8 prompts, copie…
♥ 1.0K · ⟲ 56 · 👁 179.4KView on X ↗

David Launches Jevgrep, a Context Collection CLI for Coding Agents

David Launches Jevgrep, a Context Collection CLI for Coding Agents▶

David announces jevgrep, a command-line research agent built on TypeSafe's Jev that finds relevant code by describing what it does. He claims it cuts coding agent costs by 40% on SWE-bench and links to the GitHub repository.

Original post · 1 min read
Introducing jevgrep - a research agent CLI powered by jev from @typesafeai that reduces your coding agent cost by 40% (verified on SWE-bench)

Make sure to use the built in skill so your coding agent knows to use jg for context collection github.com/dzhng/jevgrep
github.comGitHub - dzhng/jevgrep: Find code by asking what it does. A CLI for coding agents that uses Jev to discover relevant files and source context.Find code by asking what it does. A CLI for coding agents that uses Jev to discover relevant files and source context. - dzhng/jevgrep
♥ 4.9K · ⟲ 326 · 👁 447.7KView on X ↗

Startups Blacksmith, Namespace and Depot Tackle CI Build Bottlenecks

Deedy Das lists three startups, Blacksmith, Namespace and Depot, that address slow continuous integration builds, which he says have become a bottleneck as coding costs fall. He cites a Lindy post describing stratospheric CI spend.

Original post · 1 min read
As cost of coding has plummeted, long builds (CI) has become a bottleneck for many software cos. Three of the biggest startups solving this:

1. Blacksmith (@useblacksmith)
2. Namespace (@namespacelabs)
3. Depot (@depotdev)
Flo Crivello @Altimor
CI has become the top bottleneck of every engineering team I talk to (including Lindy). Our CI spend has become stratospheric. twitter.com/DavidOndrej1/status/21040010168869…
♥ 1.0K · ⟲ 50 · 👁 141.1KView on X ↗

CopilotKit Open-Sources OpenMuse Personal Assistant on GitHub

CopilotKit Open-Sources OpenMuse Personal Assistant on GitHub▶

Atai Barkai announces OpenMuse, an open source self-hostable personal assistant compatible with any agent harness, offering computer use, app connectors, goal tracking and mobile and web support. It is built with CopilotKit and AG-UI and is linked to a Muse personal AI agent post from Meta.

Original post · 1 min read
🎉 Introducing 𝙾𝚙𝚎𝚗𝙼𝚞𝚜𝚎

An open source, self-hostable personal assistant that works with any agent harness.

Includes:
- Computer use: browser, terminal & files
- Connectors for your personal apps
- Ideas, goals & progress tracking
- Built for Mobile and Web

Repo → github.com/CopilotKit/OpenMuse

Powered by @CopilotKit and AG-UI.

Clone this template and customize it however you want.
Muse @Muse
Introducing Muse, your personal AI agent from Meta that gets things done across every part of life.

Download the Muse app and get started: Muse.ai
♥ 5.9K · ⟲ 523 · 👁 1.2MView on X ↗

Grok Engineers Share Their Fourteen Work Bots in Podcast With Peter Yang

Lauren, an engineering lead on Grok @bot, says she enjoyed a conversation with Peter Yang about the fourteen bots the team uses for work and life. The linked episode covers a design bot, an engineering lead bot and tips on trusting bots with more work.

Original post · 1 min read
had a lot of fun chatting with @petergyang about our grok @bot setups!
Peter Yang @petergyang
"Everything I touch with my keyboard and mouse, I try to delegate to my bots."

Here's my new episode with @poteto and @pengzheng_, the eng and design leads for Grok @bot, where they showed me the 14 bots they use for work and life, including:

→ A design bot that turns one keyframe into a full user flow
→ An eng lead bot that manages a team of eng bots
→ How to trust your bots with more of your work

Some quotes from both:

"I like to call it the Michelin kitchen…when you say software factory, it has this connotation of mass manufactured slop."

"Sometimes I actually don't even look at the PR…
♥ 709 · ⟲ 32 · 👁 84.3KView on X ↗

TradingView Releases Official MCP Server for Claude Integration

TradingView Releases Official MCP Server for Claude Integration

Miles Deutscher announces TradingView's official MCP server for Claude, explaining how users can connect their TradingView account to agents and set it up in under three minutes with a guide.

Original post · 1 min read
TradingView recently released the official MCP server for Claude.

You can turn your agents into quant desks by giving them access to your TV account.

If you haven't set it up yet, don't worry, I got you.

Connect Claude x TradingView in <3 minutes:
♥ 1.4K · ⟲ 222 · 👁 127.7KView on X ↗

Opus 5.5 and Open Edit Produce Video in 15 Minutes, Open Source

Opus 5.5 and Open Edit Produce Video in 15 Minutes, Open Source▶

Sabba Keynejad says Opus 5.5 combined with Open Edit created a video in 15 minutes and that the whole workflow is open source. Resource links are promised in the post but are not included in the text.

Original post · 1 min read
Opus 5.5 + Open Edit made this in 15 minutes.

Wild how fast video editing is changing.

And the whole thing is open source.

Resources below ↓
♥ 24 · ⟲ 4 · 👁 4.3KView on X ↗

Grok Bot Marketplace Launches With Shelf of Prebuilt Team Agents

Grok Bot Marketplace Launches With Shelf of Prebuilt Team Agents

Denis Labelle shares a list of 16 Grok Bot articles covering topics such as work, solutions engineering, PMs, GTM, designing, engineering, and support. The list is linked to a quoted post from Eric Zakariasson announcing the Grok Bot Marketplace, where users can import prebuilt teammates.

Original post · 1 min read
Grok Bot Articles by Bot team (16)

1. Work: x.com/joshkim/status/2095633918611636627
2. Solutions Engineers: x.com/kiaraplds/status/2099296631698944298
3. PMs: x.com/n2parko/status/2088664030789681260
4. Get started with X in Grok Bot: x.com/ericzakariasson/status/2097340203790913733
5. Marketplace:
6. Run multiple teams of Grok Bots: x.com/ericzakariasson/status/2092982710465970425
7. Intro to Grok Bot: x.com/mattyp/status/2087252657589412119
8. Chat is all you need: x.com/mattyp/status/2089758434921160877
9. 8 templates to get inspired: x.com/mattyp/status/2094833468400447618
10. Guide to pstack Pt. 1: x.com/poteto/status/2094457600259842065
11. Guide to pstack Pt. 2: x.com/poteto/status/2097732320606507506
12. GTM: x.com/kristaletz/status/2089103618121314689
13. Designing: x.com/johnbai/status/2092019324797989023
14. Engineering: x.com/lingxi/status/2094493172516966781
15. Support: x.com/davidgan/status/2093397573277229288
16. GTM: x.com/BrianJ671/status/2089481016507285929
eric zakariasson @ericzakariasson
Grok Bot Marketplace is live! — We launched the Grok Bot Marketplace last week on Grok Bot · Bot Marketplace! It's a shelf of teammates other people already built for real jobs. Pick one, import it into your sidebar, and you're
♥ 393 · ⟲ 46 · 👁 46.8KView on X ↗

Veed Open-Sources OpenEdit for Agent-Driven Video Editing

Veed Open-Sources OpenEdit for Agent-Driven Video Editing▶

Sabba Keynejad introduces OpenEdit, an open-source agent-driven pipeline for creating subtitles, motion graphics, slides and rendered videos. He argues the challenge is native editing with fonts, branding and repeatable templates, not generating video from code. A linked post by Deedy Das describes producing a launch video for about $2.

Original post · 1 min read
Generating a good video from code is not the problem.

The problem is how you edit it natively.

How you use use fonts, branding and build repeatable templates.

And that’s why we built OpenEdit.

github.com/veedstudio/open-edit
Deedy @deedydas
Opus 5.5 is incredible at instructional video generation.

I made this launch video for a inference startup in 1min for ~$2. Videos like these used to take weeks if not months and a lot of coordination with agencies and 1000x the costs.

Humans broadly prefer video to text. This changes the substrate of communication. These videos actually help communicate technical ideas in seconds (photorealistic video gen like Seedance is not very useful here).
- changes how often marketing should be talking about products and launches
- change how sales people can talk about technical products to their cus…
github.comGitHub - veedstudio/open-edit: Open-source, agent-driven editing pipeline: create subtitles, motion graphics, slides, edit and render videos.Open-source, agent-driven editing pipeline: create subtitles, motion graphics, slides, edit and render videos. - veedstudio/open-edit
♥ 297 · ⟲ 5 · 👁 50.6KView on X ↗

DHH Ports Omarchy Screensaver Engine From Rust to Assembly

DHH Ports Omarchy Screensaver Engine From Rust to Assembly▶

David Heinemeier Hansson reports porting the ttfx screensaver engine from Rust to x86-64 assembly, using Opus 5.5 for a one-shot translation, with the pull request claiming up to 17x speedup. The linked PR describes 9.8x faster than Rust and 322x faster than Python.

Original post · 1 min read
I ported the Omarchy screensaver engine (ttfx) from Rust to x86-64 assembler, and it's up to 17x faster!! One-shot translation by Opus 5.5. We keep drilling until the agentic drill bit hits bedrock! github.com/omacom/ttfx/pull/35/
github.comAdd an x86-64 assembly engine: 9.8x faster than Rust, 322x faster than Python by dhh · Pull Request #35 · omacom/ttfxAdds an x86-64 assembly engine for all 37 effects, linked into the Rust binary and picked automatically at runtime. It produces byte-identical output to the Rus
♥ 5.0K · ⟲ 209 · 👁 1.4MView on X ↗

Hamel Husain and Shreya Shankar Release Evals Skill for AI Coding Agents

Hamel Husain and Shreya Shankar Release Evals Skill for AI Coding Agents

Lenny Rachitsky recommends installing a new evals skill from Hamel Husain and Shreya Shankar that guides AI coding agents in building product-specific AI evals. The linked GitHub repo collects these skills, and his post cites examples of evals improving results at Ramp, Shopify, Harvey and Cursor.

Original post · 1 min read
Pro tip: Install this new evals skill from @HamelHusain and @sh_reya, it'll save you many hours and a lot of mistakes

github.com/ai-evals-course/evals-skills
Lenny Rachitsky @lennysan
Evals have been coming up more and more in my conversations with podcast guests and PM friends.

Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them:
— @tryramp took its automatic receipt collection from 35% to 83% accuracy.
— @Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced.
— @harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score.
— @cursor_ai tuned its Auto Balance routing, …
github.comGitHub - ai-evals-course/evals-skills: Skills that guide AI coding agents to help you build product-specific AI evals.Skills that guide AI coding agents to help you build product-specific AI evals. - ai-evals-course/evals-skills
♥ 1.4K · ⟲ 93 · 👁 230.1KView on X ↗

DJ Sampath Praises Meta's Muse Agent Sandbox Design

DJ Sampath says Meta's Muse agent gets its sandboxing largely right, citing a per-user virtual machine, credential isolation, a second approval agent named Sentinel, and a full audit trail. He notes Cisco has provisioned microVMs for over 80,000 employees and plans to scale the approach to users.

Original post · 1 min read
Spent the weekend with Meta's @Muse, and they got a lot right.

- Every user gets their own VM.
- The agent never sees a password.
- A 2nd agent (Sentinel) has to approve anything that leaves the box.
- Every action is in the audit trail.

That's the most careful consumer agent sandbox I've used.

I don't know if folks realize but a VM for every user, at Meta scale, is kinda huge! It's never been done before where we simply provision a linux microVM for every single consumer.

We recently did this within @Cisco for over 80,000 of our employees. And now within our products as well as we start to scale it out to every single one of our users.

1/n
♥ 957 · ⟲ 47 · 👁 197.6KView on X ↗

Jason Fried Endorses Graphical, a Tool for Designing Visual Languages

Jason Fried praises Graphical, a tool by Josh Puckett for designing visual languages and styling components with coding agents. The post quotes Puckett's launch announcement and links to graphicalui.com.

Original post · 1 min read
Josh always brings the good stuff.
joshpuckett @joshpuckett
Introducing Graphical.

It's a simple, powerful, and fun tool to design visual languages, style components, and work with coding agents to create interfaces that look memorable and feel unique.

I hope you'll check it out at graphicalui.com/!
♥ 578 · ⟲ 18 · 👁 114.7KView on X ↗

DHH Says Future Role of Programming Frameworks Remains Uncertain

DHH writes that nobody knows the future role of programming languages and frameworks, possibly including a path from prompt to microcode, and advises developers to make the most of current tools.

Original post · 1 min read
You're going to have to deal with the fact that nobody knows exactly what the future role of programming languages and frameworks are. Maybe it all does go away and it's a straight shot from prompt to microcode! But best you can do today is get the most out of what's here now.
♥ 5.1K · ⟲ 293 · 👁 165.4KView on X ↗

AJ Asver Details Harness That Cuts Repetitive Agent Costs 90%

AJ Asver Details Harness That Cuts Repetitive Agent Costs 90%

AJ Asver describes a new harness built with Jev that learns repetitive tasks and moves steps from LLM calls to code. He says processing 100,000 compliance alerts fell from over $290K on Opus 5 to under $26K using agentrun().

Original post · 1 min read
We built a new harness using @typesafeai's Jev that cuts the cost of repetitive work by 90%. The harness learns the job as it runs, moving steps from LLM calls to code.

Running 100,000 compliance alerts costs >$290K on Opus 5.

With agentrun() we got it down to <$26K.
Miguel Ríos Berríos @MiguelriosEN
♥ 2.2K · ⟲ 124 · 👁 286.6KView on X ↗

HarnessRouter Offers Unified Interface for Agent Harnesses

HarnessRouter Offers Unified Interface for Agent Harnesses▶

Akshay Pachaar describes HarnessRouter, an open-source layer that runs multiple agent harnesses, including Codex, Claude Code, Hermes and Jev's System One, under one interface via the Unified Harness Protocol. He links to the repository and a related article.

Original post · 1 min read
Finally, an OpenRouter for agent harnesses!

(including System One by Jev)

Devs just open-sourced a plug-and-play infrastructure layer that lets you run any harness under a single interface, like:

- Codex
- Hermes
- Claude code
- DeepSeek Harness
- System One, powered by Jev
- And 9 more agent harnesses

This means you can bring Jev into the same product that already uses Codex, Claude Code, or another supported harness, without writing another implementation for sessions, streaming, files, cancellation, and failure handling.

Here's the repo: github.com/HarnessRouter/harnessrouter

(don't forget to star it ⭐ )

The harnesses run locally, and the Unified Harness Protocol (UHP) defines the common task interface with an OpenAI Responses-compatible API.

If you want to dive deeper, my recent article explains why model routing is not the same as harness routing, and what it takes to support multiple harnesses.

It also covers UHP, the full local setup, a working API call, and how sessions and files work.

Read it below.
Akshay 🚀 @akshay_pachaar
Run Any Agent Harness Under One Interface — How UHP and HarnessRouter standardize agent execution across Codex, Claude Code, Hermes, and other runtimes.

When an agent product integrates one harness directly, its backend starts depending on
♥ 2.2K · ⟲ 306 · 👁 358.0KView on X ↗

Researchers' Jev Method Claims 63x Cheaper LLM Output Checking

Researchers' Jev Method Claims 63x Cheaper LLM Output Checking

Codila reports on a Chinese research PDF testing the Jev approach across 44 benchmarks, where asking a single question achieved a 0.886 median AUROC. The post claims checking cost $0.30 versus $18.96 with LLM judges, roughly 63 times cheaper.

Original post · 1 min read
Chinese students just found the best way to use JEV for any LLM or AI agent - released a PDF research

the shift: I pasted it into Claude and GPT - and cut my costs by~63х

here’s what they found across 44 benchmarks:

1 → 7,193 responses, 10 types of failure. Jev was tested on hallucinations, prompt injections, data leaks, and other AI failures

2 → One simple question worked: 0.886 median AUROC, beating trained baselines on 25 of 31 benchmarks without task-specific training

3 → Context beat clever prompting - give Jev the source or rule it needs to check the answer against

4 → Keep the probability, not just "yes" or "no" - Fitting a threshold on 10 labeled examples raised median F1 from 0.706 to 0.793

5 → Among the 50% most confident decisions, median accuracy reached 0.933 - send uncertain cases for another review

6 → Jev even helped uncover labeling errors in three benchmarks. Sometimes the test’s "correct answer" was the problem

7 → 11.4 questions per call, with 0.31-second median latency - on 19 benchmarks, checking cost $0.30 vs $18.96 with LLM judges - roughly 63× cheaper

the result: It will made your setup CHEAPER and FASTER than what 95% of people are running

Copy the Jev setup researchers tested across 44 benchmarks - then read the full Jev architecture ↓
codila @0xCodila
Jev is the "Internet" moment for the AI industry

It tells your agents and LLMs what to do next, in milliseconds and at almost zero cost

If you set it up correctly, you will have the AI engineer’s stack for 2028

In this article, I show you how
♥ 822 · ⟲ 138 · 👁 97.9KView on X ↗

Anthropic Offers Free Claude Code Credits Through October 7

Anthropic Offers Free Claude Code Credits Through October 7

Claude's official developer account announced free credits, claimable via a link or the /claim-credit command in the CLI with a connected GitHub account. Claims are due by October 7, and terms apply.

Original post · 1 min read
Claude is handing out free credits 🫡
ClaudeDevs @ClaudeDevs
Follow the link below to claim the credit or run /claim-credit in the CLI. You’ll need GitHub connected to start a session. Claim by Oct 7. Terms apply.

claude.ai/code/claim-credit/10
♥ 8.8K · ⟲ 247 · 👁 2.8MView on X ↗

Developer Sets Strict Testing Rules for AI Coding Agents

Ansh Nanda shares AGENTS.md rules banning unit tests written after code and favoring end-to-end tests with verifiable artifacts. He responds to a quoted post about an AI agent generating trivial unit tests.

Original post · 1 min read
At the top of my AGENTS.md:

- NEVER write unit tests after you write code.
- Highly prefer E2E tests as the sole testing mechanism. Use them to verify complex features work. At the end of E2E tests, produce a verifiable and repeatable artifact.
- If you must test a system in isolation, FIRST write all the ways it could fail, THEN write the code.
dex @dexhorthy
leave it to your boy opus to add 10 unit tests to ensure a constant string contains various substrings
♥ 5.6K · ⟲ 246 · 👁 1.3MView on X ↗