Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

289 stories

Hamel Husain and Shreya Shankar Release Evals Skill for AI Coding Agents

Hamel Husain and Shreya Shankar Release Evals Skill for AI Coding Agents

Lenny Rachitsky recommends installing a new evals skill from Hamel Husain and Shreya Shankar that guides AI coding agents in building product-specific AI evals. The linked GitHub repo collects these skills, and his post cites examples of evals improving results at Ramp, Shopify, Harvey and Cursor.

Original post · 1 min read
Pro tip: Install this new evals skill from @HamelHusain and @sh_reya, it'll save you many hours and a lot of mistakes

github.com/ai-evals-course/evals-skills
Lenny Rachitsky @lennysan
Evals have been coming up more and more in my conversations with podcast guests and PM friends.

Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them:
— @tryramp took its automatic receipt collection from 35% to 83% accuracy.
— @Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced.
— @harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score.
— @cursor_ai tuned its Auto Balance routing, …
github.comGitHub - ai-evals-course/evals-skills: Skills that guide AI coding agents to help you build product-specific AI evals.Skills that guide AI coding agents to help you build product-specific AI evals. - ai-evals-course/evals-skills
♥ 1.4K · ⟲ 93 · 👁 230.1KView on X ↗

Alex Finn Demos Grok Bot Voice Agents in Tesla Cybertruck

Alex Finn Demos Grok Bot Voice Agents in Tesla Cybertruck▶

Alex Finn says he had early access to the new Grok Bot for Tesla and praises having the car drive while he talks to a set of agents. He shares a video ride in his Cybertruck demonstrating the release.

Original post · 1 min read
Grok Bot just released for Tesla and I'm blown away

I was lucky enough to have early access. Having your car drive you around while you talk to an army of agents is incredible

In this video I take you for a ride in my Cybertruck and show you just how awesome this new release is
♥ 3.3K · ⟲ 314 · 👁 891.6KView on X ↗

Matthew Berman Pitches Jev for Automating Instagram Short-Form Research

Matthew Berman Pitches Jev for Automating Instagram Short-Form Research▶

Matthew Berman says Jev can scan a library of 12 million shorts to analyze what makes top creators work, replacing manual research for a real estate client seeking Instagram growth. The post is part of a wave of praise for Jev's marketing use.

Original post · 1 min read
jev has made the impossible POSSIBLE.

so this morning my real estate buddy asked me how to break out on instagram.

I could:
- ask her to scroll.
- find her favorite creators
- map out each hook
- figure out what makes them work.
- rework it in her voice
- rinse & repeat every single week.

Or i could have jev go through a library of 12 million shorts and do it for her.
jaffa @dsqjaffa
Jev is INSANE for Marketing — You can't escape it.
Literally EVERYWHERE you look on your feed is filled with something about "jEv is iNsAnE"... and it's generated tens of millions of views on X in less than a week.

I'm guilty of
♥ 598 · ⟲ 44 · 👁 94.0KView on X ↗

CopilotKit Open-Sources OpenMuse Personal Assistant on GitHub

CopilotKit Open-Sources OpenMuse Personal Assistant on GitHub▶

Atai Barkai announces OpenMuse, an open source self-hostable personal assistant compatible with any agent harness, offering computer use, app connectors, goal tracking and mobile and web support. It is built with CopilotKit and AG-UI and is linked to a Muse personal AI agent post from Meta.

Original post · 1 min read
🎉 Introducing 𝙾𝚙𝚎𝚗𝙼𝚞𝚜𝚎

An open source, self-hostable personal assistant that works with any agent harness.

Includes:
- Computer use: browser, terminal & files
- Connectors for your personal apps
- Ideas, goals & progress tracking
- Built for Mobile and Web

Repo → github.com/CopilotKit/OpenMuse

Powered by @CopilotKit and AG-UI.

Clone this template and customize it however you want.
Muse @Muse
Introducing Muse, your personal AI agent from Meta that gets things done across every part of life.

Download the Muse app and get started: Muse.ai
♥ 5.9K · ⟲ 523 · 👁 1.2MView on X ↗

Prajwal Tomar Pairs Jev With GPT-6 Astra and Higgsfield Workflow

Prajwal Tomar Pairs Jev With GPT-6 Astra and Higgsfield Workflow▶

Prajwal Tomar describes a motion graphics workflow where GPT-6 Astra plans concepts, Higgsfield renders variants, and Jev selects the best render automatically. He says a client adopted the setup and promotes an accompanying article course.

Original post · 1 min read
I gave Jev access to an AI motion graphics studio and honestly this is kind of terrifying.

So last week my workflow was basically GPT-6 Astra comes up with the concept, Higgsfield renders 5 versions, and then I just sit there watching all 5 trying to figure out which one actually works.

That last part was killing me. Every single round I had to stop and watch videos for like 10 minutes.

Now I plugged Jev in and it's honestly wild.

It looks at every asset and picks what's worth rendering before Higgsfield even touches it. No explanation, just a decision in like 100 milliseconds.

Astra plans it, Higgsfield renders it, Jev decides it.

So basically I went from rendering five versions and hoping one of them works to just rendering the one Jev already picked.

I set this up for a client last week and they absolutely loved it. They're not even working with their AI creative agency anymore.

If you want to turn GPT-6 Astra into a motion graphics studio, read the article below.
Prajwal Tomar @PrajwalTomar_
How to Turn GPT-6 Astra Into a Motion Graphics Studio (Full Course) — GPT-6 Astra is two weeks old and it is still the only thing my feed talks about. Two days ago I wrote about what it does to web design once you show it real websites. Today is the thing I actually
♥ 26 · ⟲ 1 · 👁 6.4KView on X ↗

Steve Yegge Proposes Agentic Technical Program Managers for Enterprises

Steve Yegge argues that AI agents acting as technical program managers, who drive projects without direct authority, could be the most direct route for coding agents into enterprises. He describes his own agent-based TPM seats in his Wheelhouse factory that use email, Slack and nagging to drive projects, and links to his site.

Original post · 4 min read
Hear me out: Agentic TPMs (Technical Program Managers.) I had this idea in Sydney while chatting with Martha McKeen at CBA. I think this winds up being the most immediate and direct way that coding agents can make their way into the enterprise, and it will set the stage for true AI employees rolling in next year.

So. Build-side agents are great but they don't escape the SDLC. Only devs are using them. There are a handful of business people vibe coding SaaS, but for the most part, non-engineers aren't using coding agents to help with their jobs. Right? Not yet.

Autonomous 24x7 unmanned queue-based "operator" agents, like the ones that handle internal or external customer issues, are great. But they are narrowly scoped, and generally require devs involved to set them up and maintain them.

Neither builder nor operator agents are automatically going viral internally and helping run the company. They stay in their lanes. But what if their lane was to help run projects?

I was a TPM at Amazon in 1999. Bezos brought in high-powered engineers with people skills to run difficult cross-functional projects and programs. TPMs are used at Google, Uber, Netflix, and other companies, and they are always in high demand and short supply.

I have a class of agents in my Wheelhouse factory that act just like TPMs. They have external email and Slack, and talk to my accountant, lawyers, players. Each one has a project lane and drives it. They use Progress By Nagging, which... works.

A TPM owns delivery, but has no authority, and no resources. They can only ask, observe, document, and report. This is just like my TPM Wheelhouse seats, who have been helping me drive dozens of projects to completion, large and small, for months.

Agents, particularly smarter models, will go to great lengths to document the hell out of everything in the domain where they're operating. They'll capture all the tribal knowledge and unwritten rules. They can create topological maps of your project, org dependencies, and workflows. They'll bulldoze through silos and knowledge-hoarders and figure out how the company actually works, and document it all. And nag people along the way.

This kind of agent sits well in constraint-space. They're cheap: You don't need to use the fanciest models; anyone with Opus or Sol access could have a TPM agent. And TPM agents have low risk and blast radius, because they cannot act. Unlike builder agents, which create new problems (like merge-queue and code-review bottlenecks), TPM agents simply shine a light on the org, and nudge things along.

It doesn't matter what format they're recording their findings in. It could be Sanskrit and hieroglyphics. When it comes time to merge their findings with those of other TPM agents, it will all translate trivially into your company brain.

Anyone in the company can stand up a TPM agent. It's like a personal chief of staff. There's no dependency on engineers. Everyone can do it; it doesn't even have to have a paced rollout. And there's no product to buy, no tech to install, maybe just a Skill you give people. Maybe you put a company wrapper on it. But it's just an agent that's playing the TPM role.

TPM agents will wind up training human orgs on human-agent interactions. Humans start getting emails or DMs from agents, work-related, and will have to get comfortable replying and interacting. Companies can push the social side along without waiting for engineers to finish messing with the SDLC, which honestly will never finish.

Other kinds of agents struggle at enterprises because they lack context. TPM agents will build that missing context as their exhaust, no joke; they've done it for my game without me even asking. TPM agents are the jungle explorers that will map out your organization, and you'll discover all sorts of fun stuff, like that you had 3 teams doing the same thing. TPM agents are a low-risk, high-impact way to start figuring out how AI can help you run your project, or organization.

I'll write a blog post about this, but feel free to start now. Go! Just give me credit when you win big with this idea. And if you want my help, ping me on yegge.ai.
♥ 466 · ⟲ 33 · 👁 42.9KView on X ↗

Jev Model Router Mod Routes Claude Code Tasks Across Models

Jev Model Router Mod Routes Claude Code Tasks Across Models▶

Alvaro Cintas describes jev-model-router, an early-access Claude Code mod built on function hooks that uses Jev to pick a model per turn based on task complexity and risk. He provides install steps, a link, and configuration notes for API keys.

Original post · 1 min read
You can now use Jev right inside Claude Code 🤯

It's called jev-model-router, an early access mod built on Claude Code's new function hooks. Before every turn, it checks in with Jev and asks:

> how mechanical the task is
> how much reasoning it needs
> whether it's risky

Then it routes:
→ it'll move up to a stronger model on weak evidence, and only drops to a cheaper one when it's confident the task is simple
→ every decision gets logged in your transcript
→ if the call fails, your request runs untouched

Setup:
1. copy the install command: npx claude-code-templates@latest --mod productivity/jev-model-router
2. paste it inside Claude Code
3. run claude with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 set
4. accept the trust prompt on first launch

Link: aitmpl.com/component/mod/productivity/jev-mode…

Also works with no api key, it just falls back to a built-in classifier with no confidence score.

For real jev routing, add your typesafeApiKey or gatewayApiKey to ~/.claude/settings.json.

Follow me for more AI workflows and tutorials.
♥ 447 · ⟲ 51 · 👁 43.4KView on X ↗

Developer Builds Fast Search Over 6,000 Y Combinator Startups

Developer Builds Fast Search Over 6,000 Y Combinator Startups▶

Aayan says he built a search tool on Jev that indexes more than 6,000 Y Combinator startups and returns results in under a second. He reports about 90 million tokens and $2.70 in total testing costs, and shares a demo video.

Original post · 1 min read
Jev (@typesafeai) is so insane & cheap for search!!

> 6000+ @ycombinator Startups indexed.
> Sub 1 second search results.
> 90M tokens & $2.7 in total testing costs.

Search any startup in a second, in any way!
- Color - Niche - Your Competitor - Age - Image - etc...

> watch the entire video, it's so freaking cool omg!

> this is the coolest thing i have ever built for fun! (worked on it for 2 days straight!)
♥ 1.7K · ⟲ 56 · 👁 129.7KView on X ↗

Lauren Tan's Grok Bot Rules Shared as Installable Skill

Lauren Tan's Grok Bot Rules Shared as Installable Skill▶

cat.png recommends a set of Grok Bot team rules based on Lauren Tan's approach, including one bot per job, a chief-of-staff router, and human approval for money and deploys. He links a GitHub repo containing the rules as an installable skill.

Original post · 1 min read
I still can't f**king get why people keep adding more Grok Bots instead of copying the rules the best teams run them on.

A Japanese Grok Bot team I’ve been studying is built around Lauren “poteto” Tan’s rules.

The useful part:

> One bot, one job.
> One chief of staff routes everything.
> Draft before send.
> No proof, no "done".
> Money, deploys, permissions stay behind human approval.

If a bot keeps making the same mistake, turn the fix into a rule or skill.

That’s basically the whole game.

Less prompting. Better roles. Better guardrails.

All the rules are listed here, and you can set them up as a skill in your Grok Bot:

github.com/unicodef1wn/lauren-poteto-rules-

Save it before you build your next Grok Bot team 👇
unicode @unicodef1wn
One of the SpaceXAI engineers building Grok Bot explained how she gets coding agents she can actually trust.

I turned Lauren Tan’s talks and extended Q&A into one installable skill.

Her main point:

If an agent cannot verify its own work, you are still the verification system.

Agent writes the code → you open the app → find what it broke → send screenshots back → repeat

Her workflow changes that loop.

The skill teaches your agent to:

1. Read the affected code before guessing the cause
2. Reproduce the bug before changing anything
3. Run the real user flow, not just build and typecheck
4.…
♥ 133 · ⟲ 13 · 👁 22.1KView on X ↗

Peter Yang Asks How Amazon Can Detect Human Versus AI Access

Peter Yang Asks How Amazon Can Detect Human Versus AI Access

Peter Yang asks how Amazon could tell whether a request comes from a human or an AI agent when both appear to originate from the same IP address, responding to a post about Amazon cutting off Muse.

Original post · 1 min read
Dumb question - how is Amazon going to be able to tell if it's a human or AI accessing it assuming it's coming from the human's IP address?
BuccoCapital Bloke @buccocapital
Amazon cuts off Muse.

While I am bullish Meta and Muse, I think many people are overlooking the digital knife fight that’s about to occur

Nobody wants to get commoditized or layered here. Let the games begin
♥ 283 · ⟲ 4 · 👁 159.6KView on X ↗

AJ Asver Details Harness That Cuts Repetitive Agent Costs 90%

AJ Asver Details Harness That Cuts Repetitive Agent Costs 90%

AJ Asver describes a new harness built with Jev that learns repetitive tasks and moves steps from LLM calls to code. He says processing 100,000 compliance alerts fell from over $290K on Opus 5 to under $26K using agentrun().

Original post · 1 min read
We built a new harness using @typesafeai's Jev that cuts the cost of repetitive work by 90%. The harness learns the job as it runs, moving steps from LLM calls to code.

Running 100,000 compliance alerts costs >$290K on Opus 5.

With agentrun() we got it down to <$26K.
Miguel Ríos Berríos @MiguelriosEN
♥ 2.2K · ⟲ 124 · 👁 286.6KView on X ↗

Colin McDermott Releases Grok Bot Template For Jev Classification

Jev router by Colin

Colin McDermott shares a Grok Bot template built on Jev that provides fast, calibrated classification for routing, urgency, labels and rubrics, and links to the bot on x.ai.

Original post · 1 min read
I built a Grok Bot x Jev template. Give all your agents a fast, calibrated classifier.

Try it here: x.ai/bot/lS9XaHCr9QTTHhNtb0VQX
x.aiJev router by ColinA Jev-powered Grok Bot for fast, calibrated classification. Choice, Score, and yes/no judgements for routing, urgency, labels, and rubrics — before your...
♥ 457 · ⟲ 30 · 👁 89.1KView on X ↗

Developer Lauren Shares Method for Shipping 2,500 PRs Monthly

Developer Lauren Shares Method for Shipping 2,500 PRs Monthly▶

Developer lauren (@poteto) posts a free video walkthrough of how she shipped 2,500 pull requests to production last month. The talk was originally planned for Cursor Compile in London and is viewable on X at 2x speed.

Original post · 1 min read
here's how i shipped 2,500 PRs last month to production

this was originally supposed to be for Cursor Compile in London. i couldn't make it since i was livestreaming for Grok @Bot Galaxy so i'm making it available for free here on X! watch it on 2x speed, i talk slowly
♥ 16.0K · ⟲ 1.4K · 👁 3.5MView on X ↗

Beacon Open-Sources Self-Improving Memory Layer for Coding Agents

Beacon Open-Sources Self-Improving Memory Layer for Coding Agents▶

Avi Chawla describes Beacon, an open-source tool from Asymptote Labs that captures coding-agent sessions across Claude Code, Codex, Cursor and OpenCode. It uses the Jev model to score runs and turn approved corrections into reusable skills, linking to the GitHub repository.

Original post · 2 min read
Another insane Jev use case!

Jev is making it dramatically cheaper to evaluate what actually happened inside an agent run.

And finally, someone open-sourced a self-improving memory layer that can put that signal to work across agent harnesses:

- Claude Code
- Codex
- Cursor
- OpenCode, and 20+ more

Beacon by @asymptotelabs continuously captures your agent history across harnesses and uses Jev to identify which runs are actually worth learning from.

It then turns the highest-signal workflows, corrections, and debugging patterns into reusable skills.

GitHub repo: github.com/Asymptote-Labs/agent-beacon

(don’t forget to star it ⭐ )

Beacon preserves the complete session history. But preserving a run and learning from it are two different things.

Most coding-agent sessions contain routine exploration, failed commands, and fixes that only apply to one task. The trace can remain available for inspection without turning every detail into guidance for future agents.

Jev scores each run for evidence, reuse potential, and human correction signals. An application policy then decides whether to promote, review, or discard it.

The recording shows this in action.

Claude receives a coding task, modifies the implementation, and runs the tests. I then provide an edge-case correction, so Claude updates the code and adds regression coverage.

Beacon automatically captures the complete session. Jev evaluates whether the correction contains a reusable engineering lesson.

Once approved, that lesson becomes available to other coding agents working on the project.

Since it works across harnesses:
- Claude Code sessions can teach Codex.
- Cursor debugging can improve OpenCode.

So a problem solved by one agent should not need to be learned from scratch by another.

If you want to dive deeper into Jev, I also wrote a hands-on guide to building this Jev-style decision path with open models, entirely locally.

Read it below.
Avi Chawla @_avichawla
Build your own Jev (100% local) — Everything you need to turn an open-source LLM into a fast, local decision engine without retraining it. It covers next-token scoring, fixed choices with probability distributions, SGLang, and a
♥ 2.0K · ⟲ 239 · 👁 291.8KView on X ↗

HarnessRouter Offers Unified Interface for Agent Harnesses

HarnessRouter Offers Unified Interface for Agent Harnesses▶

Akshay Pachaar describes HarnessRouter, an open-source layer that runs multiple agent harnesses, including Codex, Claude Code, Hermes and Jev's System One, under one interface via the Unified Harness Protocol. He links to the repository and a related article.

Original post · 1 min read
Finally, an OpenRouter for agent harnesses!

(including System One by Jev)

Devs just open-sourced a plug-and-play infrastructure layer that lets you run any harness under a single interface, like:

- Codex
- Hermes
- Claude code
- DeepSeek Harness
- System One, powered by Jev
- And 9 more agent harnesses

This means you can bring Jev into the same product that already uses Codex, Claude Code, or another supported harness, without writing another implementation for sessions, streaming, files, cancellation, and failure handling.

Here's the repo: github.com/HarnessRouter/harnessrouter

(don't forget to star it ⭐ )

The harnesses run locally, and the Unified Harness Protocol (UHP) defines the common task interface with an OpenAI Responses-compatible API.

If you want to dive deeper, my recent article explains why model routing is not the same as harness routing, and what it takes to support multiple harnesses.

It also covers UHP, the full local setup, a working API call, and how sessions and files work.

Read it below.
Akshay 🚀 @akshay_pachaar
Run Any Agent Harness Under One Interface — How UHP and HarnessRouter standardize agent execution across Codex, Claude Code, Hermes, and other runtimes.

When an agent product integrates one harness directly, its backend starts depending on
♥ 2.2K · ⟲ 306 · 👁 358.0KView on X ↗

Muse Reveals Specs of Its Ubuntu Virtual Machines for Users

Tanay Jaipuria lists the resources given to Muse users' virtual machines, including 2 vCPUs from a 126-core AMD EPYC, 8GB of RAM, 100GB of persistent disk and Ubuntu 24.04.

Original post · 1 min read
The specs of Muse VM users get:

- 2 vCPUs (a slice of a 126-core AMD EPYC)
- 8GB RAM
- 100GB persistent disk
- Ubuntu 24.04
♥ 1.2K · ⟲ 36 · 👁 147.2KView on X ↗

Tobi Lütke Argues Code Should Be Judged on Merit, Not Origin

Tobi Lütke responds to a report that KDE's draft AI policy discourages disclosing LLM use, saying code should be accepted on merit with a human accountable for it. He says good code is good and slop is slop regardless of how it was made.

Original post · 1 min read
This is the way. Accept code on merit and ensure that a person takes accountability for it. Doesn’t matter if it was typed, chiseled, generated, or bit-flipped via magnetized needle on a chip.

If it’s good, it’s good. If it’s slop, it’s slop.
The Lunduke Journal @LundukeJournal
KDE is working on an official AI / LLM policy, and it reads like the rules of Fight Club.

In short, KDE’s AI policy:

1) Encourages using AI, as long as a human is kept “in the loop”.

2) But you can’t tell anyone that you used AI.

“Don’t disclose LLM usage”.

“Don’t add ‘Assisted-by: [some LLM]” (as is done in the Linux kernel).

“Nobody in KDE should know if you use an LLM”.

In other words: “Welcome to developing KDE with AI. The first rule of developing KDE with AI is: you do not talk about developing KDE with AI.”

invent.kde.org/plasma/plasma-workspace/-/work_…
♥ 4.0K · ⟲ 228 · 👁 233.7KView on X ↗

Fastbrowse Launches as Open-Source, Low-Cost Browser Agent

Fastbrowse Launches as Open-Source, Low-Cost Browser Agent

Furqan Rydhan introduces fastbrowse, an experimental open-source browser agent where an LLM plans actions and each claim cites an exact quote from the page. He says it is significantly cheaper and faster for agents to browse the web.

Original post · 1 min read
Introducing fastbrowse.

An open-source fast browser agent.

Jev picks each action directly from the page, an LLM plans and every claim cites an exact quote.

It's significantly cheaper and faster for agents to browse the web now.

Early and experimental, but very promising.

fastbrowse.ai
♥ 578 · ⟲ 37 · 👁 71.9KView on X ↗

SpaceXAI Engineer Describes Running Twenty Agents Under One Chief of Staff

SpaceXAI Engineer Describes Running Twenty Agents Under One Chief of Staff▶

Sheema Moto relays a post from SpaceXAI engineer Peng Zheng, who says he moved from 250 unsuccessful applications to a $850,000 offer by running about 20 agents managed by a Chief of Staff agent. The post promotes a 40-minute workshop on GrokBot and agent automation.

Original post · 1 min read
SpaceXAI engineer Peng Zheng:

"250 applications, two years, zero offers. Then he stopped applying as one engineer and showed up as one engineer running 20 agents. We offered him $850,000.

I don't use GrokBot like Google anymore. I built a 24/7 system once - now it automates 95% of my life and work every day. Only 1% of people run a single Chief of Staff agent that manages the other ~20 agents and knows everything about them.

That's not a skill gap. That's a stack gap, and it takes one evening to close."

GrokBot → Chief of Staff → 20 Agents → Auto-Delegation → 24/7 System

In a 40-minute workshop, SpaceXAI engineers show how to stop managing every agent by hand - and how they actually use GrokBot.

Research → Build → Launch → Improve

This will save you 20 hours of useless agent tutorials.
♥ 349 · ⟲ 33 · 👁 28.6KView on X ↗

Jared Palmer Releases Kev Open Source Decision Model Family Based on Qwen3

Jared Palmer Releases Kev Open Source Decision Model Family Based on Qwen3

Jared Palmer announces Kev-0.6B, 4B and 8B, open source Apache 2.0 decision models built on Qwen3 with LoRA and a pointer head. He reports Kev-8B scores 79.6% out of domain versus 85.7% for Jev, and says the models are drop-in compatible with the TypeSafe API.

Original post · 1 min read
UPDATE: Kev-0.6B, 4B, and 8B are now available. Kev is a family of small open source Jev-like decision models you can train and run yourself.

This new family is based on Qwen3 using the same LoRA + small pointer head technique as before, but scaled up.

Out of domain, on data Kev never trained on: Kev-8B 79.6%, Jev 85.7%.

• Drop-in TypeSafe System One API; their SDK works with one `base_url` change
• Kev-4B serves on a 32 GB Mac in bf16: ~300 ms for five questions, ~40 ms on an H100
• Repeated documents hit a KV cache: 2-2.5x faster
• Apache 2.0 License. Kev-4B trains in 40 minutes on one H100. Kev-8B in 83 minutes.

Code, weights, evals: github.com/jaredpalmer/kev
Jared Palmer @jaredpalmer
Kev-0.5B: A tiny open source Jev-like decision model with a TypeSafe-compatible API based on Qwen2.5-0.5B that you can train and run on a MacBook Pro.

Model card and weights are available on GitHub

github.com/jaredpalmer/kev
♥ 2.7K · ⟲ 227 · 👁 291.9KView on X ↗

Deedy Lists Practical Use Cases for Muse and Instinct AI Agents

Deedy shares five uses for Muse and Instinct, including filing FOIA requests, issuing spend-limited cards, completing visa forms and buying reservations at opening time. He argues dark patterns on the web are being broken and that creativity now limits what is possible.

Original post · 1 min read
My favorite use cases for Muse / Instinct so far:

1. Submit FOIA requests to request data from the US government
2. Creating spend-limited Privacy cards to spend on subscriptions without having them recur
3. End to end filed an entire visa form for a country
4. Responded to a coordination mail for a wedding by finding the flight and hotel details
5. Look for reservations for restaurants or concerts when they open and purchase them immediately

A lot of the web was designed with dark patterns: increase friction to prevent enough humans from doing something, and now those walls are completely broken.

At this point, I feel like I’m squarely limited by creativity and understanding what is possible.
♥ 1.1K · ⟲ 38 · 👁 173.7KView on X ↗

Kevin Rose Releases Grok Bot That Indexes Saved Instagram Videos Offline

Min Choi reports that Kevin Rose shared a Grok Bot that transcribes and analyzes a user's saved Instagram videos, building a searchable local markdown wiki in the style of Karpathy. The quoted post says the tool runs offline and uses Grok Voice Transcribe and Grok Vision, with X and TikTok support coming later.

Original post · 1 min read
Kevin Rose just shared his Grok Bot.

It takes years of your saved IG video, runs Transcribe 2.0 + Vision on them, and builds a local Karpathy-style markdown wiki you can search and ask.
Kevin Rose @kevinrose
Built this Grok Bot, fully offline index of all your saved instagram videos via @karpathy-wiki-style .md. Uses @bot, @grok Voice Transcribe 2.0, Grok Vision + more.

(X + TikTok coming soon)
♥ 377 · ⟲ 21 · 👁 89.3KView on X ↗

LangChain Tests Jev as Fast Semantic Judge for Agent Evaluation

Harrison Chase highlights LangChain's testing of Jev against LLM judges on accuracy, repeatability, latency and cost. He argues Jev's cheap, fast verifiers suit online evaluation of large numbers of agent traces.

Original post · 1 min read
"jev as a judge"

cheap and fast semantic verifiers! great for evals - especially online evals where you want to grade LOTS of traces
LangChain @LangChain
We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation.
♥ 512 · ⟲ 46 · 👁 79.5KView on X ↗

Gergely Orosz Argues Scrum Only Suited Slow Release Cycles

Gergely Orosz argues Scrum made sense for teams shipping every two to three months but held back teams doing daily or continuous deployment, which startups and Big Tech abandoned a decade ago.

Original post · 1 min read
Scrum made sense at companies/teams that shipped a new version of their product every 2-3 months (or less frequent)

As soon as a team got to shipping daily or had CD: they only held teams back

Most startups + Big Tech moved on a decade+ ago, legacy companies remained only
Klaas @forgebitz
it's funny how scrum masters completely disappeared

and no one really cared
♥ 2.3K · ⟲ 112 · 👁 269.9KView on X ↗

Field Notes From SpaceX AI Team Offer Multi-Bot Agent Playbook

Field Notes From SpaceX AI Team Offer Multi-Bot Agent Playbook▶

cat.png shares a free MIT-licensed repository of notes from three SpaceX engineers who shipped a product in 72 hours using Grok Bot agents. The notes cover a chief-of-staff bot, scoped specialist bots, proof-before-merge rules, cost figures and 40 documented antipatterns. The quoted post describes a 14-page guide on a growth lead building bots to replace his own work.

Original post · 2 min read
I still can't f**king get why nobody runs Grok Bot the way the people who built it do.

Three SpaceX engineers shipped a product from an empty repo in 72 hours, live on stream. 433 PRs.

I took notes all three days and turned them into a repo. Free, MIT.

The setup:

● One chief of staff. The only bot you talk to. You build every specialist through it, so it knows who does what.

● One bot, one job. Scope it like a job description. Split when the scope grows, not before.

● Drafts before sends. Write access is earned, not granted.

● No proof, no merge. The bot reproduces the bug, fixes it, attaches a video to the PR.

● Four human gates only: migrations, deploys, money, permissions. Everything else merges on its own.

The numbers nobody puts in the demo:

→ Support ticket: $1-2 with a reasoning pass. $0.20 once low-complexity tickets go to a script.

→ Full case-study deck: $20-30. By hand: 4-5 hours.

→ One team, one month, this stack: 2,500 PRs merged to prod.

→ Launch day: 4,000 games, 2,000 users, $0 revenue.

Cost is a property of your setup, not the tool. And autonomy is not a business model.

40 things broke on air. Every one is written down with the rule that came out of it.

Inside the repo:

> AGENTS.md for your root
> 9 playbooks
> 69 bot roles
> 40 antipatterns
> 14-page guide

The tools are not the moat. Everyone has them. The moat is a loop that decides what not to ship.

Save it before you wire your next agent 👇
github.com/unicodef1wn/grokbot-field-notes
unicode @unicodef1wn
SpaceX AI's growth lead built a bot to replace himself. I turned his session into a 14-page PDF.

He started with six bots. One job each.

→ researcher, product marketer, website ops, performance, analyst
→ each one takes the handoff from the one before it
→ he leaves comments in the doc, the bot reads them and redrafts

Then he built the seventh bot.

It read the other six. It read every place he had to step in and fix them.

Then it ran three campaigns and told him nothing until they were done.

His words: "I don't want to talk to any of the other bots."

Also in the PDF:

→ 8 prompts, copie…
♥ 1.0K · ⟲ 56 · 👁 179.4KView on X ↗

SimSlim Passes 2,000 GitHub Stars for iOS Simulator Tool

GitHub - MobAI-App/simslim: Run more iOS simulators on one Mac by disabling background daemons a simulator doesn't need

Interlap announces that SimSlim, an open-source project that disables unneeded background daemons so more iOS simulators can run on one Mac, has passed 2,000 GitHub stars. The developer thanks contributors and testers.

Original post · 1 min read
SimSlim just crossed 2,000 GitHub stars!

Huge thanks to everyone who contributed, reported issues, tested changes, and helped make it much better than the thing I originally released.

github.com/MobAI-App/simslim
github.comGitHub - MobAI-App/simslim: Run more iOS simulators on one Mac by disabling background daemons a simulator doesn't needRun more iOS simulators on one Mac by disabling background daemons a simulator doesn't need - MobAI-App/simslim
♥ 233 · ⟲ 19 · 👁 25.8KView on X ↗

Developer Sets Strict Testing Rules for AI Coding Agents

Ansh Nanda shares AGENTS.md rules banning unit tests written after code and favoring end-to-end tests with verifiable artifacts. He responds to a quoted post about an AI agent generating trivial unit tests.

Original post · 1 min read
At the top of my AGENTS.md:

- NEVER write unit tests after you write code.
- Highly prefer E2E tests as the sole testing mechanism. Use them to verify complex features work. At the end of E2E tests, produce a verifiable and repeatable artifact.
- If you must test a system in isolation, FIRST write all the ways it could fail, THEN write the code.
dex @dexhorthy
leave it to your boy opus to add 10 unit tests to ensure a constant string contains various substrings
♥ 5.6K · ⟲ 246 · 👁 1.3MView on X ↗

Open-Source Laya-MLX Runs Fast Decision Models on Apple Silicon

Open-Source Laya-MLX Runs Fast Decision Models on Apple Silicon▶

A developer introduces laya-mlx, an open-source MLX port of the Laya typed decision model, claiming 7-14 ms decisions on an M3 Max and a demo playing Snake at 60 decisions per second. A GitHub repository is linked.

Original post · 1 min read
介绍比Jev快50倍,在你设备上跑的laya-mlx!

只在你的设备上占用最高1G内存

Laya是一个开源的类似于Jev的,基于文本输出概率的分类系统

我将其移植到MLX,并且做了一些性能优化!

视频中就是这个模型在我的本地M3Max上玩贪吃蛇

这个模型能够以每秒决策60次的速度玩贪吃蛇!

github.com/mizorewww/laya-mlx
github.comGitHub - mizorewww/laya-mlx: Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API. - mizorewww/laya-mlx
♥ 14.1K · ⟲ 1.2K · 👁 4.3MView on X ↗