Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Agents & Dev Tools

Coding agents, developer tools, workflows, open source

Microsoft Unveils Biggest Copilot Update, Adding Autopilot and Code

Microsoft Unveils Biggest Copilot Update, Adding Autopilot and Code▶

Satya Nadella announces a major Copilot overhaul framed as a new operating system for work, including a proactive enterprise agent, in-tenant app building, Chat and Cowork merged into Home, Office embedded in Copilot, and a proactive Today feed in Microsoft 365.

Original post · 1 min read
We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:

· Autopilot: proactive and long-running agent built for the enterprise
· Code: build apps with Copilot, hosted inside your company’s tenant
· Home: Chat + Cowork together
· Office: now fully embedded in Copilot (and Copilot embedded in Office, of course!)

Plus, you can invoke Copilot in Teams, and we’re introducing Today, a proactive experience that surfaces the most important information from across M365 without needing to ask for it.

The way we work is changing and so are our workflows. This update brings AI into that flow, from answering a question, to building an app, to getting work done on your behalf.
♥ 12.3K · ⟲ 1.6K · 👁 3.8MView on X ↗

Developer Says Opus 5.5 Recreated Adobe Apps in Rust and Open-Sourced Them

Developer Says Opus 5.5 Recreated Adobe Apps in Rust and Open-Sourced Them

Kevin Picchi claims one developer used Opus 5.5 to re-create all Adobe applications, port them to Rust and release them as open source. The post is brief and shares a media attachment.

Original post · 1 min read
This guy re-created all the adobe apps with opus-5.5, ported them to rust, and opensourced them

something really cool is happening
♥ 51.4K · ⟲ 2.7K · 👁 1.8MView on X ↗

LuxAlgo Releases Free Open-Source Edge Stats Trading Engine

Mr. Quant reports that LuxAlgo released Edge Stats, a free open-source statistics engine for technical traders that tests setups such as opening range breakouts and fair value gaps with confidence intervals. The quoted LuxAlgo announcement links to a GitHub repository and says it runs locally.

Original post · 1 min read
Edgeful just got Nuked☢️

Today @LuxAlgo released Edge Stats.

FREE. Open-source. No subscription.

Test ORBs, FVGs, Initial Balance, Gap Fills + 38 more setups.

You can also pull up every session on a chart for free using @velacharts

The industry is changing for good🔥
LuxAlgo @LuxAlgo
Today we're introducing Edge Stats, our free, open-source statistics engine for technical traders.

See how often your setup actually played out, with the sample size and confidence interval on every number, and view the actual sessions on @velacharts.

Built-in by default:

> Opening Range Breakouts (ORB)
> Fair Value Gaps (FVGs)
> Initial Balance Breakouts
> Gap Fills & Reversals
+ 38 more, each via the LuxAlgo Library.

Runs on your own computer with your data, or start instantly with free crypto data.

Free forever, no subscription.

github.com/LuxAlgo/edge-stats
♥ 594 · ⟲ 41 · 👁 60.6KView on X ↗

Grok Bot Guides Offer Playbooks for Building AI Teammates

Grok Bot Guides Offer Playbooks for Building AI Teammates

Ben Lang recommends asking a main Grok bot to review xAI's new Bot Guides and apply the best practices to one's own bots. The linked guides cover practical playbooks for AI teammates.

Original post · 1 min read
Pro tip: Tell your main @bot to go through the new guides and recommend best practices for your Bots

x.ai/bot/guides
x.aiGrok Bot · Grok Bot GuidesPractical playbooks for AI teammates.
♥ 5.0K · ⟲ 290 · 👁 523.1KView on X ↗

Olivia Moore Details Her Seven-Tool Personal AI Agent Stack

Olivia Moore lists the AI agents she uses, including Instinct for purchases, Muse for mini-apps, Tomo for goal tracking, Bot for X monitoring, OpenAI's Dots for long-running projects, Town for work email and docs, and Tab for outbound calls. She says she is using Dots to start a business.

Original post · 1 min read
My agent stack:

- @instinct - on-the-go purchases, form fill, local recs
- @Muse - mini-apps that require reliable connectors (ex. "track my sleep and movement, send me a survey daily")
- @tomo - motivation + goal tracking, also want to try their Interest Groups
- @bot - X feed monitoring + alerts
- Dots (@OpenAI) - deep, long-running projects...I'm currently making it start a business 👀
- @townai - everything work-related in email, Slack, docs
- @Tabdotbot - outbound phone calls / appts
♥ 479 · ⟲ 26 · 👁 49.1KView on X ↗

Levie Argues Evals Will Gate Enterprise Adoption of AI Agents

Box CEO Aaron Levie argues enterprises cannot automate work they cannot measure, so evals are essential for adopting AI agents. He says domain-specific evals will become a major opportunity, and a quoted post notes data labeling firms expect Fortune 1000 customers to drive revenue and company evals may become proprietary IP.

Original post · 1 min read
You can’t automate what you can’t measure. This means that evals are one of the gates to diffusion of AI in the enterprise.

We can test our deterministic processes through software, but most enterprises have no useful way of understanding how their non-deterministic processes are working today. Specifically the work that agents are doing for them.

Evals are mission critical for enterprises adopting AI because you have no other way of knowing what’s working, what’s broken, what changed, what improved, what you can do more of, etc. if you don’t have a good sense of how agents work in your environment today. All changes, upgrades, and deployments are downstream from good evals.

Not only are we going to get vastly more domain specific evals over time for the labs and across the industry, but every enterprise will also need a clear sense of how agents are performing in their environment as well. Huge opportunity.
Alex Lieberman @businessbarista
Just spoke to one of the big data labeling businesses.

Few interesting insights:
- They predict the majority of their revenue will come from Fortune 1000 enterprises, not labs in a few years
- They believe every company will want to own their intelligence, but owning intelligence does not necessarily mean using open source models
- A company’s evals will become their main proprietary IP given the improvement in agent performance after properly setting up & running internal eval environments
- Most enterprises haven’t graduated from coding agents and it’s largely due to not having the proper e…
♥ 450 · ⟲ 51 · 👁 134.0KView on X ↗

Anthropic Details How It Made claude.ai Three Times Faster

Claude

Boris Cherny highlights an Anthropic blog post describing how the team made claude.ai three times faster in two weeks using Claude to measure, debug and improve performance. The post includes prompts and methods for engineers optimizing their own apps.

Original post · 1 min read
If you've noticed how fast claude.ai/login and the Desktop app have become in the last few weeks, here's how we did it.

Lots of juicy learnings & techniques in the blog post for engineers working on speeding up your own apps.
ClaudeDevs @ClaudeDevs
We made claude​.ai 3x faster in two weeks.

Here’s how we use Claude to measure, debug and improve performance. Prompts and methods included.

claude.dev/blog/how-we-made-claude-ai-faster/
claude.aiClaudeClaude is Anthropic's AI, built for problem solvers. Tackle complex challenges, analyze data, write code, and think through your hardest work.
♥ 5.3K · ⟲ 168 · 👁 847.5KView on X ↗

Claude Design Head Explains Working Process With Opus 5.5

Claude Design Head Explains Working Process With Opus 5.5▶

Codez shares remarks from Claude Head of Design Jenny Wen, who says design is now mainly reference selection, CLAUDE.md and spec writing, and prompt design with Opus 5.5. The post links to a video course on motion design prompting.

Original post · 1 min read
Claude Head of Design, Jenny Wen:

"after Opus 5.5 the design process is actually dead. it's already 10x faster & cheaper than 99% of designers.

designing now is giving the right reference, building CLAUDE.md and spec, designing the right prompt - that's the new stack of a designer"

in a 1-hour speech, Head of Claude design explained how to use Claude's new models at 100% of their power

watch this today, then explore the full Opus 5.5 guide with prompt techniques and demos below
Movez @0xMovez
How to build motion design studio with Opus 5.5 ( Full-course ) — Most people who try motion design with Opus 5.5 end up with the same video: centered text on a gradient, everything fading in, a logo at the end.
They don't give it a reference, don't give it a
♥ 71 · ⟲ 7 · 👁 9.1KView on X ↗

Alex Finn Praises Hark, a Proactive Agent With Adaptive Interface

Alex Finn Praises Hark, a Proactive Agent With Adaptive Interface▶

Alex Finn says the proactive AI agent Hark impressed him after a week of early access, and shares a video explaining its features. The post describes an interface that adapts to each user.

Original post · 1 min read
This new AI agent blew my mind

Hark is a proactive agent with a custom interface that completely transforms based on who you are

They gave me early access a week ago and I've really enjoyed using it

Here's everything you need to know about Hark:
♥ 528 · ⟲ 24 · 👁 72.1KView on X ↗

Vercel Launches Drives for Sandbox in Public Beta

Guillermo Rauch argues successful AI agents need separated brain, hands and files components, and announces Vercel Drives, persistent storage for Vercel Sandbox now in public beta. Drives allow up to four mounts per sandbox and 16 TiB per Drive.

Original post · 1 min read
Muse, Instinct, OpenClaw, Claude Code…
All successful agents have 3 key components:

🧠 Brain → model, harness (logic)
👐 Hands → tools, computer, browser
🗃️ Files → memories, skills, repos

The 'easy' way is to throw all these in 1 stateful computer (a Mac Mini)

Like, you run 𝚌𝚕𝚊𝚞𝚍𝚎 or 𝚏𝚡 in your mac, you keep it running all day with 𝚌𝚊𝚏𝚏𝚎𝚒𝚗𝚊𝚝𝚎, it has storage, and CLIs and apps installed.

But if you want to cost-efficiently run agents in the cloud, you actually start breaking down these parts.

🧠 The harness can run in Fluid compute. To make it reliable across restarts, rollouts, crashes, you make its event log durable using Workflow.

👐 The hands can be a dedicated browser fleet like Browserbase/Kernel, a computer like Sandbox, and even more efficient lightweight tools like just-bash.

🗃️ 🆕 What was missing was a way to also decouple storage. Imagine you want to run a memory consolidation cron job every night ("dreaming"). You can read/write to the files directly without 'booting up' the agent's full computer.

Today we're introducing the perfect companion to Sandbox: Drives. We shipped the computer for agents, now we're giving you the 'external disk' you can attach at will. It's early, and we'll be expanding capabilities here quickly.

Btw, breaking apart the agent into these independent parts not only optimizes costs in a big way, it also *massively* improves security and auditability. I'd argue you can't even run a secure agent otherwise!
Vercel Developers @vercel_dev
Vercel Sandbox now has persistent storage with Drives, in public beta on every plan.

▪︎ Store agent workspaces, data, models, deps
▪︎ Read snapshots across parallel sandboxes
▪︎ Mount up to four Drives per sandbox
▪︎ Up to 16 TiB per Drive

vercel.com/changelog/drives-for-vercel-sandbox…
♥ 2.3K · ⟲ 138 · 👁 262.7KView on X ↗

TypeSafe Releases Claude Code Skill for Building on Jev

TypeSafe Releases Claude Code Skill for Building on Jev

Mnimiy describes an official TypeSafe skill that teaches Claude Code to structure calls to Jev, batching questions into one request to cut costs by 12.2 times in TypeSafe's cookbook. The post also outlines fallback strategies for when Jev's 70-500 ms checks fail or go silent.

Original post · 1 min read
it's f*cking gold

TypeSafe just dropped the official skill that teaches Claude Code to build on Jev

left alone, coding agents ask Jev one question per call and guess the field names.

this skill rewires the order before a single line gets written:

docs -> behaviour -> judgments -> one request -> code decides

every rule in it, broken down on two pages:

> see the failure behind each of its 11 instructions
> pick Choice, Noul or Score in one glance
> paste the prompt skeleton straight into your agent

12.2x

the bill drop when 13 questions share one call, measured in TypeSafe's cookbook.

a short file that does a lot of the thinking for your agent.
Mnimiy @Mnilax
it's f*cking insane

Jev sat between GPT and me, killing every draft that broke my rules before i saw it. good setup.

then Jev did not answer. the agent decided silence was safer and stopped sending me anything at all. took a second agent to unstick it.

your checker needs a branch for the moment it goes quiet:

> let it through tagged unverified and keep moving
> hand the risky span to the big model
> park it in a queue and retry in a minute
> wake a human when the action is irreversible

70-500 ms

one Jev call, by TypeSafe's own number. a half-second checker still takes down the whole pipe…
♥ 2.2K · ⟲ 174 · 👁 291.7KView on X ↗

Lauren Describes Using Grok Bot With Cursor Cloud Agents

Lauren Describes Using Grok Bot With Cursor Cloud Agents

Lauren describes a workflow in which a Grok bot in Slack hands coding tasks to Cursor cloud agents, organized into Projects, so the bot acts as a manager while each agent writes code on its own machine.

Original post · 1 min read
grok @bot combined with cloud agents is a magical experience! here's my favorite use case:

1. create a new team engineer bot and add it to slack
2. @ your bot whenever you want it to do some coding work
3. you can even tell it to create new Projects, which is a way to group together many related agents into one conversation

you can use any model available on cursor with these cloud agents, and since they each have their own computer, it frees up your bot to be more of a manager rather than write the code itself!
Grok Bot @bot
Grok Bot is now more powerful for building software.

Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.
♥ 1.5K · ⟲ 107 · 👁 3.1MView on X ↗

Viral Post Offers Claude Code Prompt for Regime-Detection Trading Bot

Viral Post Offers Claude Code Prompt for Regime-Detection Trading Bot

The post shares a prompt for Claude Code that builds a trading bot using a Hidden Markov Model to detect market regimes, with per-regime strategies, walk-forward testing and a Sharpe 1.5 threshold. It attributes the prompt to a leaked Jane Street quant document, and the post's claims about its origin are unverified.

Original post · 1 min read
A Jane Street quant got fired for leaking this internally, then posted it anyway

paste into Claude Code tonight. This one prompt builds a regime-detection trading bot:

fits a Hidden Markov Model to read calm, choppy, stressed and crash states

writes a separate playbook per state, never mixes strategies

switches only after a new state holds the lead for several candles, so it never flip-flops on noise

walks forward through real data with real costs, and refuses to ship unless it clears Sharpe 1.5 and beats buy-and-hold

ends every run with one line in caps: what could blow up this account

bookmark it
Roan @RohOnChain
Jev is the FASTEST AI model ever built for trading

It makes calibrated buy/sell decisions in under 100 ms

That is one real decision on every single block, 24/7

In this article I've shown EXACTLY how to build HFT trading system with Jev (from scratch)
♥ 1.7K · ⟲ 152 · 👁 295.3KView on X ↗

Call4.me Launches AI Agent That Makes Phone Calls for Users

Call4.me Launches AI Agent That Makes Phone Calls for Users▶

Nick Khami introduces call4.me, an AI agent that handles phone calls for tasks like canceling subscriptions and booking appointments. It works inside Claude Code and Codex, and the post links to callbay, which offers prepaid credits from $10.

Original post · 1 min read
introducing call4.me, an ai agent that handles phone calls for you.

you can use it to cancel subscriptions, book medical appointments, dinners, change flights, or anything else possible with a call.

most importantly, it works in claude code & codex. i spend most of my time in the terminal, so i want my productivity tools to live there as well. switching to imessage or another app annoys me.

i've now handled dozens of tasks over the phone using it & am in love. if you're like me & hate breaking concentration to dial, you'll likely love it as well.

call4.me/
call4.mecallbay: your AI agent makes phone calls for youyour AI agent makes phone calls for you. book dinners, doctor appointments, call dealerships. one prompt to install, prepaid credits from $10.
♥ 1.0K · ⟲ 49 · 👁 195.4KView on X ↗

AI Hedge Fund Backtester Hides Ticker and Dates to Prevent LLM Leakage

AI Hedge Fund Backtester Hides Ticker and Dates to Prevent LLM Leakage▶

Virat Singh explains that backtesting LLM-driven trading is difficult because returns are already encoded in model weights. He says the AI Hedge Fund backtester now hides the ticker, dates, and position size to limit this leakage, shown in an attached video.

Original post · 1 min read
Backtesting an LLM is hard.

The returns are already in the weights.

AI Hedge Fund's backtester now hides the ticker, dates, and size to limit leakage.
♥ 90 · ⟲ 9 · 👁 15.1KView on X ↗

Claude Code Creator Boris Cherny Explains Prompting Opus 5.5

Claude Code Creator Boris Cherny Explains Prompting Opus 5.5▶

A post shares a 12-minute video in which Boris Cherny, creator of Claude Code, says Opus 5.5 needs less prompting than earlier models, and links to an article on prompting the model, which Anthropic released days earlier.

Original post · 1 min read
Boris Cherny, creator of Claude Code:

"Opus 5.5 does in a day what used to take your team a month. Most people will keep using it wrong."

In 12 minutes he explains why Opus 5.5 needs less prompting than any model before, and why your old detailed prompts now work against you.

Watch it, then read the article below on how to prompt Opus 5.5 👇
Rahul @sairahul1
Claude Opus 5.5 dropped 3 days ago.

And it might be the strongest model Anthropic has shipped yet.

But most people are still prompting it like older Claude models.

I turned Anthropic's latest guidance into a complete Opus 5.5 prompting masterclass:
♥ 710 · ⟲ 63 · 👁 151.0KView on X ↗

Deedy Demonstrates Fully Automated Video Summaries of Research Papers

Deedy Demonstrates Fully Automated Video Summaries of Research Papers▶

Deedy says Claude Opus 5.5 can generate an eight-minute, 3Blue1Brown-style explainer video from any research paper, using a summary of a paper on regularized recursive self-improvement of agent harnesses as an example.

Original post · 1 min read
You can now generate an entire 3blue1brown style video from any research paper with Opus 5.5.

Here’s a 8min video summary of “Regularized Recursive Self Improvement of Agent Harnesses”.

The 90%ile educational YouTuber is fully automated.
♥ 5.6K · ⟲ 398 · 👁 332.4KView on X ↗

Instinct Adds Group Chat Support for Coordinating Plans With Friends

Julie Zhuo highlights a feature letting early-access users add Instinct to group chats, where the agent joins the whole group to coordinate plans, times and logistics without requiring friends to install it.

Original post · 1 min read
This is cool. Fantastic way for people to learn about the power of agents through their friends.
Noah Shinn @noahrshinn
Instinct in group chats

Starting today, early access users can add Instinct to group chats. A new Instinct joins and works for the whole group.

Making plans with friends usually turns into a frustrating back-and-forth over times and places. With Instinct in the group, you can explore options together, agree on a plan and get it done, all in one thread. Your friends don't even need Instinct to join in.

A few things it’s good at:
- Planning a weekend trip with friends, including dates and arrival times
- Coordinating logistics with a roommate
- Getting tickets as soon as they go on sale, then…
♥ 119 · ⟲ 4 · 👁 35.5KView on X ↗

Hiten Shah Open-Sources 16 Competitive Intelligence Skills for AI

Hiten Shah has released 16 open-source competitive intelligence skills covering market briefings, positioning, pricing, launches, battlecards, deal prep and win/loss analysis. He says a product built around them will be shown on Friday.

Original post · 1 min read
I just open-sourced 16 competitive intelligence skills.

They cover market briefings, positioning, pricing, launches, battlecards, deal prep, win/loss, and more.

Each one teaches AI a different method for the work.

On Friday, I’ll show you the product I’ve been building around them and what changes when those methods can use the market knowledge behind them. hiten.com/make-your-best-ai-work-reusable
♥ 153 · ⟲ 13 · 👁 21.7KView on X ↗

DHH Ports Campfire to Rust, Reporting 19 to 44 Times Speedup

DHH Ports Campfire to Rust, Reporting 19 to 44 Times Speedup

DHH says a beta Rust rewrite of Basecamp's Campfire chat app was produced with little manual effort, though the code is more verbose. The linked GitHub project describes a single binary using the Rails app's data that runs 19 to 44 times faster.

Original post · 1 min read
Been running a beta version of Campfire rewritten in Rust today! The code itself is still ugly as sin, and 6x as verbose, but the conversion was free*, and I never had to look at the Rust directly, so this is awesome 🤩 github.com/basecamp/once-campfire-rust
github.comGitHub - basecamp/once-campfire-rust: ONCE Campfire in Rust: one binary, the Rails app's data, 19–44× fasterONCE Campfire in Rust: one binary, the Rails app's data, 19–44× faster - basecamp/once-campfire-rust
♥ 2.3K · ⟲ 90 · 👁 1.1MView on X ↗

DHH Says Hand-Writing Code Is No Longer Economically Viable

Rails World 2026 Opening Keynote - DHH

David Heinemeier Hansson, creator of Ruby on Rails, argues that writing code by hand is no longer economically viable for most programmers at most companies. He frames this as a bright moment for software creation, in a Rails World 2026 keynote.

Original post · 1 min read
It's pencils down, people. Writing code by hand is no longer an economically viable skill for most programmers at most companies. But the future of making software has never been brighter. Don't you dare black pill this beautiful moment! youtube.com/watch?si=6FsfJXS22mk52gyu&v=vDjW_d…
youtube.comRails World 2026 Opening Keynote - DHHDHH opens Rails World 2026 in Austin with a keynote on the age of A...
♥ 8.9K · ⟲ 950 · 👁 2.9MView on X ↗

Steve Yegge Proposes Agentic Technical Program Managers for Enterprises

Steve Yegge argues that AI agents acting as technical program managers, who drive projects without direct authority, could be the most direct route for coding agents into enterprises. He describes his own agent-based TPM seats in his Wheelhouse factory that use email, Slack and nagging to drive projects, and links to his site.

Original post · 4 min read
Hear me out: Agentic TPMs (Technical Program Managers.) I had this idea in Sydney while chatting with Martha McKeen at CBA. I think this winds up being the most immediate and direct way that coding agents can make their way into the enterprise, and it will set the stage for true AI employees rolling in next year.

So. Build-side agents are great but they don't escape the SDLC. Only devs are using them. There are a handful of business people vibe coding SaaS, but for the most part, non-engineers aren't using coding agents to help with their jobs. Right? Not yet.

Autonomous 24x7 unmanned queue-based "operator" agents, like the ones that handle internal or external customer issues, are great. But they are narrowly scoped, and generally require devs involved to set them up and maintain them.

Neither builder nor operator agents are automatically going viral internally and helping run the company. They stay in their lanes. But what if their lane was to help run projects?

I was a TPM at Amazon in 1999. Bezos brought in high-powered engineers with people skills to run difficult cross-functional projects and programs. TPMs are used at Google, Uber, Netflix, and other companies, and they are always in high demand and short supply.

I have a class of agents in my Wheelhouse factory that act just like TPMs. They have external email and Slack, and talk to my accountant, lawyers, players. Each one has a project lane and drives it. They use Progress By Nagging, which... works.

A TPM owns delivery, but has no authority, and no resources. They can only ask, observe, document, and report. This is just like my TPM Wheelhouse seats, who have been helping me drive dozens of projects to completion, large and small, for months.

Agents, particularly smarter models, will go to great lengths to document the hell out of everything in the domain where they're operating. They'll capture all the tribal knowledge and unwritten rules. They can create topological maps of your project, org dependencies, and workflows. They'll bulldoze through silos and knowledge-hoarders and figure out how the company actually works, and document it all. And nag people along the way.

This kind of agent sits well in constraint-space. They're cheap: You don't need to use the fanciest models; anyone with Opus or Sol access could have a TPM agent. And TPM agents have low risk and blast radius, because they cannot act. Unlike builder agents, which create new problems (like merge-queue and code-review bottlenecks), TPM agents simply shine a light on the org, and nudge things along.

It doesn't matter what format they're recording their findings in. It could be Sanskrit and hieroglyphics. When it comes time to merge their findings with those of other TPM agents, it will all translate trivially into your company brain.

Anyone in the company can stand up a TPM agent. It's like a personal chief of staff. There's no dependency on engineers. Everyone can do it; it doesn't even have to have a paced rollout. And there's no product to buy, no tech to install, maybe just a Skill you give people. Maybe you put a company wrapper on it. But it's just an agent that's playing the TPM role.

TPM agents will wind up training human orgs on human-agent interactions. Humans start getting emails or DMs from agents, work-related, and will have to get comfortable replying and interacting. Companies can push the social side along without waiting for engineers to finish messing with the SDLC, which honestly will never finish.

Other kinds of agents struggle at enterprises because they lack context. TPM agents will build that missing context as their exhaust, no joke; they've done it for my game without me even asking. TPM agents are the jungle explorers that will map out your organization, and you'll discover all sorts of fun stuff, like that you had 3 teams doing the same thing. TPM agents are a low-risk, high-impact way to start figuring out how AI can help you run your project, or organization.

I'll write a blog post about this, but feel free to start now. Go! Just give me credit when you win big with this idea. And if you want my help, ping me on yegge.ai.
♥ 466 · ⟲ 33 · 👁 42.9KView on X ↗

Beacon Open-Sources Self-Improving Memory Layer for Coding Agents

Beacon Open-Sources Self-Improving Memory Layer for Coding Agents▶

Avi Chawla describes Beacon, an open-source tool from Asymptote Labs that captures coding-agent sessions across Claude Code, Codex, Cursor and OpenCode. It uses the Jev model to score runs and turn approved corrections into reusable skills, linking to the GitHub repository.

Original post · 2 min read
Another insane Jev use case!

Jev is making it dramatically cheaper to evaluate what actually happened inside an agent run.

And finally, someone open-sourced a self-improving memory layer that can put that signal to work across agent harnesses:

- Claude Code
- Codex
- Cursor
- OpenCode, and 20+ more

Beacon by @asymptotelabs continuously captures your agent history across harnesses and uses Jev to identify which runs are actually worth learning from.

It then turns the highest-signal workflows, corrections, and debugging patterns into reusable skills.

GitHub repo: github.com/Asymptote-Labs/agent-beacon

(don’t forget to star it ⭐ )

Beacon preserves the complete session history. But preserving a run and learning from it are two different things.

Most coding-agent sessions contain routine exploration, failed commands, and fixes that only apply to one task. The trace can remain available for inspection without turning every detail into guidance for future agents.

Jev scores each run for evidence, reuse potential, and human correction signals. An application policy then decides whether to promote, review, or discard it.

The recording shows this in action.

Claude receives a coding task, modifies the implementation, and runs the tests. I then provide an edge-case correction, so Claude updates the code and adds regression coverage.

Beacon automatically captures the complete session. Jev evaluates whether the correction contains a reusable engineering lesson.

Once approved, that lesson becomes available to other coding agents working on the project.

Since it works across harnesses:
- Claude Code sessions can teach Codex.
- Cursor debugging can improve OpenCode.

So a problem solved by one agent should not need to be learned from scratch by another.

If you want to dive deeper into Jev, I also wrote a hands-on guide to building this Jev-style decision path with open models, entirely locally.

Read it below.
Avi Chawla @_avichawla
Build your own Jev (100% local) — Everything you need to turn an open-source LLM into a fast, local decision engine without retraining it. It covers next-token scoring, fixed choices with probability distributions, SGLang, and a
♥ 2.0K · ⟲ 239 · 👁 291.8KView on X ↗

Wes Bos Details Unrestricted Features of Meta's Muse Assistant

Wes Bos Details Unrestricted Features of Meta's Muse Assistant

Developer Wes Bos shares findings from inspecting Meta's Muse, noting it can zip directories, ships with about 70 preloaded skills, deploys sites to Cloudflare, installs software freely, and has an optional email inbox feature.

Original post · 1 min read
I cracked open Meta's Muse - here are some interesting bits 🔽

1. You can ask it for the entire contents of /opt/ and it will zip it up for you

2. There a ~70 preloaded skills and CLI for everything from spotify to instragram-cli

3. "Spaces" is their website builder that uses Tanstack, Bun and Tailwind. Sharing a site deploys it to Cloudflare

4. It will install anything - like a torrent client

5. It has it's own "Muse DB" with a skill to query its own memories/convos with strict guard rails

6. There are a few things not enabled on my account, including "muse mail" which will give muse an inbox and ability to send email

Generally impressed that you can just do anything - it's not limited or watered down like I would have imagined
♥ 2.0K · ⟲ 69 · 👁 199.9KView on X ↗