Thursday, October 8, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Search

Latest stories — use the filters to narrow by keyword, section or date.

Farza Open-Sources Clicky on GitHub

GitHub - farzaa/clicky

Farza announces that his project Clicky is now open source and links to its GitHub repository. The post gives no further detail on what the project does.

Original post · 1 min read
Now open-source.

Go build.

github.com/farzaa/clicky
github.comGitHub - farzaa/clickyContribute to farzaa/clicky development by creating an account on GitHub.
♥ 238 · ⟲ 7 · 👁 30.2KView on X ↗

Ray Dalio Warns the US-Israel-Iran War Is Part of a Longer World War

The Big Thing: We Are In A World War That Isn’t Going To End Anytime Soon.

Investor Ray Dalio argues the US-Israel-Iran conflict is one piece of a broader world war that will not end soon, and that markets underestimate its duration. He covers the Strait of Hormuz, the risk of Iranian missiles and nuclear threats, and the US midterm elections.

Original post · 12 min read
X ArticleThe Big Thing: We Are In A World War That Isn’t Going To End Anytime Soon.
I will start off by wishing you well in these challenging times and by saying that the picture I paint in the following observations is not the picture I wish to be true; it is the picture that I believe to be true based on what I have learned and what the indicators that I use to objectively see things now suggest is true.
As a global macro investor for over 50 years who has needed to study all things that affected markets over the last 500 years to know how to deal with what’s coming at me, it appears to me that most people tend to focus on and react to the attention-grabbing things that are going on at the time—like what is going on with Iran now—and miss the much bigger, more important, and longer-term-evolving things that are driving what is going on and what is likely to happen. For today, that is most importantly that the US-Israel-Iran war is just part of a world war that we are in and that isn’t going to end anytime soon.
Certainly, what will happen with the Strait of Hormuz (most significantly, whether control of passage through it will be taken away from Iran and which countries are willing to spend how much blood and treasure to make that happen) will have many enormous repercussions all around the world. There are also the issues of whether Iran will still have a capacity to inflict harm on its neighbors with missiles and the threat of nuclear weapons, of how many troops the US is sending and what they will do, of the cost of gasoline, and of the upcoming US midterm elections.
All these near-term issues are important, but they lead people to miss the really big, even more important things. More specifically, because most people tend to have this short-term perspective, they now expect, and the markets are pricing in, that this war won’t last long and that when it ends we will get back to “normal.” Virtually nobody is talking about the fact that we are in the early stages of a world war that isn’t going to end anytime soon. Because I have this different perspective, I will now explain it.
Here are the really big things going on that I think we need to pay attention to:
1. We are now in a world war that isn’t going to end anytime soon.
While this sounds like hyperbole, it is indisputable that we are now in an interconnected world that has a number of shooting wars going on (e.g., the Russia-Ukraine-Europe-US war; the Israel-Gaza-Lebanon-Syria war; the Yemen-Sudan-Saudi Arabia-UAE war that also involves Kuwait, Egypt, Jordan, and other related countries; and the US-Israel-GCC-Iran war). Most of these wars involve major nuclear powers, and there are also significant non-shooting wars (i.e., trade, economic, capital, technology, and geopolitical influence wars) that most countries are in.
Together, these conflicts make up a very classic world war that is analogous to past “world wars.” For example, past “world wars” consisted of interrelated wars that were generally slipped into without any clear start dates or declarations of war. Those past wars combined into a classic world war dynamic that affected them all, as is happening with the current wars. I described that war dynamic in detail in Chapter 6, “The Big Cycle of External Order and Disorder,” of my book Principles for Dealing with the Changing World Order, which I published about five years ago, so it’s there if you want a more comprehensive description. That chapter covers the arc of what we are seeing happen and what is likely to happen.
2. Understanding how the sides are lining up and what their relationships are is very important.
It is quite easy to see objectively how the sides are lining up via indicators such as their treaties and formal alliances, their votes at the United Nations, their leaders’ statements, and their actions. For example, one can see how China is aligned with Russia and Russia is aligned with Iran, North Korea, and Cuba, and how that group is largely opposed to the United States, Ukraine (which is aligned with most European countries), Israel, the GCC states, Japan, and Australia.
These alliances matter a lot in imagining how things will go for the relevant players, so they need to be considered when observing what’s going on and what’s likely to happen. For example, we see this reflected in China’s and Russia’s actions at the UN on Iran needing to open the Strait of Hormuz. Similarly, as another example, while it’s said that China is particularly harmed by the closing of the Strait of Hormuz, that is wrong because China’s mutually supportive relationship with Iran will probably allow oil going to China to get through, and China’s relationship with Russia will ensure that China will get oil from Russia. China also has a lot of other energy (coal and solar) and a huge inventory of oil (about 90-120 days’ usage). Also noteworthy is that China consumes 80-90% of Iran’s oil output, which adds to the power of its relationship with Iran. All things considered, it appears that China and Russia are the relative economic… continue on X ↗
♥ 6.1K · ⟲ 1.3K · 👁 1.8MView on X ↗
AI7/10

NVIDIA Releases PersonaPlex 7B, an Open Full-Duplex Voice Model

NVIDIA Releases PersonaPlex 7B, an Open Full-Duplex Voice Model▶

Linus Ekenstam highlights NVIDIA's PersonaPlex 7B, an open-source real-time speech model that listens and speaks simultaneously and can interrupt mid-sentence. He claims it beats Gemini Live on dialog naturalness and runs locally without API costs.

Original post · 1 min read
NVIDIA just killed the awkward pause in voice AI 😱

PersonaPlex 7B is a real-time conversational model that listens AND speaks simultaneously. Like actually interrupts you mid-sentence like a human.

Beat Gemini Live on dialog naturalness. 18x faster interruptions.

100% open source. Run it locally. No API bill. No latency.
♥ 3.9K · ⟲ 433 · 👁 443.3KView on X ↗
AI8/10

Anthropic Growth Head Says Claude Now Automates Growth Experiments

Lenny Rachitsky summarizes an interview with Anthropic Head of Growth Amol Avasare, covering how Claude handles growth work, the shrinking PM-to-engineer ratio, and the continued need for PMs to align stakeholders. Anthropic's ARR reportedly grew from $1B to $19B in a year.

Original post · 4 min read
My biggest takeaways from @AnthropicAI's Head of Growth Amol Avasare:

1. Engineering is getting the most AI leverage—and it’s squeezing PMs and designers. With Claude Code, a five-engineer team now produces the output of 15 to 20 engineers. But PM and design productivity haven’t scaled proportionally. The result is a compressed ratio where one PM is effectively managing the output of a much larger engineering team. Anthropic's growth team is responding in two ways: hiring even more PMs (!), and formally deputizing product-minded engineers to act as mini-PMs for any project with less than two weeks of engineering time.

2. Anthropic is using Claude to automate its own growth. The internal initiative is called CASH (Claude Accelerates Sustainable Hypergrowth). It works across four stages: identifying opportunities, building features, testing quality, and analyzing results. Right now it handles copy changes and minor UI tweaks. The win rate is comparable to a junior PM with two to three years of experience, and improving rapidly.

3. The one part of PM work that AI can’t automate yet: getting six people in a room to agree. Amol and his head of design joke that even with AGI, it’ll still be impossible to align six stakeholders. Cross-functional coordination—managing opinions, navigating politics, mediating tradeoffs—remains the bottleneck that AI doesn’t touch for larger projects. This is why Amol believes PM roles aren’t going away, and may actually grow.

4. 60-80% of Anthropic’s growth team's projects have no PRD. For smaller work, kickoffs happen on Slack—messages back and forth with product-minded engineers who can push back and ask the right questions. For larger projects, Amol believes in a proper 30-minute cross-functional kickoff (legal, safeguards, stakeholders) to surface concerns early.

5. Adding friction to onboarding drives growth—if the friction helps users understand why the product is for them. His work Mercury, MasterClass, Calm, and now Anthropic, adding steps to onboarding flows consistently improved conversion. The key: cut annoying friction that doesn’t add value, but add friction that helps users understand why the product is for them.

6. AI companies need to focus on bigger bets, not better A/B tests. Amol’s argument: if your core product value is driven by AI, then the future value is orders of magnitude higher than today’s value, because model capabilities grow exponentially. In that world, micro-optimizations capture a shrinking share of a growing pie. Traditional growth teams do 60% to 70% small optimizations and 20% to 30% big swings. At Anthropic, they flip this ratio.

7. Amol built a weekly AI agent that scans Slack for cross-functional misalignment. Using Cowork with the Slack MCP, he has a scheduled task that looks across his projects and conversations and surfaces areas where teams are about to do overlapping work or pull in different directions. A colleague on the enterprise team already caught major misalignment that would have caused weeks of wasted effort.

8. A traumatic brain injury taught Amol the principle that now drives his work: freedom through constraints. In early 2022, a kick to the head during a Muay Thai sparring session caused a traumatic brain injury. Amol spent nine months off work and months relearning to walk, unable to look at screens or listen to music for more than 20 seconds. He was re-injured a month after joining Mercury and had to take two more months off. He’s still not fully healed. But the constraints—no alcohol, no caffeine, mandatory breaks, daily meditation—have become the habits that let him operate at the intensity Anthropic demands. “The true freedom in life is learning how to be content when you don’t get what you want.”
Lenny Rachitsky @lennysan
Anthropic is on an unprecedented growth run.

Just in the past year they grew from $1B to $19B ARR. They added $6B in ARR just in *February*. Companies like Palantir and Atlassian took 15-20 years to reach ~$5B ARR. Anthropic is adding that every month.

Amol Avasare is head of growth at Anthropic, and one of the most impressive people I've had on the podcast.

In his first ever public interview, Amol shares:
🔸 How Anthropic is automating growth experiments with Claude (their internal tool called “CASH”)
🔸 Why activation is the single highest-leverage growth problem in AI
🔸 Why Amol is hiri…
♥ 1.6K · ⟲ 172 · 👁 351.9KView on X ↗
AI9/10

New Yorker Investigation Details Board Concerns Over Sam Altman

New Yorker Investigation Details Board Concerns Over Sam Altman▶

A thread summarizes a New Yorker investigation into Sam Altman based on interviews, memos compiled by Ilya Sutskever and private notes from Dario Amodei. It describes alleged patterns of dishonesty that preceded Altman's firing and reinstatement at OpenAI.

Original post · 3 min read
The New Yorker just dropped a massive investigation into Sam Altman, based on over 100 interviews, the previously undisclosed "Ilya Memos," and Dario Amodei's 200+ pages of private notes. It's the most detailed account yet of the pattern of behavior that led to Sam's firing and rapid reinstatement at OpenAI. Here's the breakdown:

> Ilya compiled ~70 pages of Slack messages, HR documents, and photos taken on personal phones to avoid detection on company devices. He sent them to board members as disappearing messages. The first memo begins with a list headed "Sam exhibits a consistent pattern of . . ." The first item is "Lying."

> Dario kept detailed private notes for years under the heading "My Experience with OpenAI" (subheading: "Private: Do Not Share"), totaling 200+ pages. His conclusion: "The problem with OpenAI is Sam himself."

> Sam reportedly told Mira his allies were "going all out" and "finding bad things" to damage her reputation after the firing. Thrive put its planned $86B investment on hold and implied it would only close if Sam returned, giving employees financial incentive to back him.

> Sam texted Satya Nadella directly to propose the new board composition: "bret, larry summers, adam as the board and me as ceo and then bret handles the investigation." The two new members selected to oversee an independent inquiry into Sam were chosen after close conversations with Sam himself.

> Before OpenAI, senior employees at Loopt asked the board to fire Sam as CEO on two separate occasions over concerns about leadership and transparency. At Y Combinator, partners complained to Paul Graham about Sam's behavior, and Graham privately told colleagues "Sam had been lying to us all the time."

> OpenAI's superalignment team was promised 20% of the company's compute. Four people who worked on or with the team said actual resources were 1-2%, mostly on the oldest cluster with the worst chips. The team was dissolved without completing its mission.

> Sam told the board that safety features in GPT-4 had been approved by a safety panel. Helen Toner requested documentation and found the most controversial features had not been approved. Sam also never mentioned to the board that Microsoft released an early ChatGPT version in India without completing a required safety review.

> Sam made a secret pact with Greg and Ilya where he agreed to resign if they both deemed it necessary, essentially appointing his own shadow board. The actual board was alarmed when they learned about it.

> Sam struck a deal with Greg to become CEO while simultaneously telling researchers that Greg's authority would be diminished, and telling Greg something different.

> A board member described Sam as having "two traits almost never seen in the same person: a strong desire to please people in any given interaction, and almost a sociopathic lack of concern for the consequences of deceiving someone." Multiple sources independently used the word "sociopathic."

> OpenAI is reportedly preparing for an IPO at a potential $1 trillion valuation while securing government contracts spanning immigration enforcement, domestic surveillance, and autonomous weaponry in war zones.
♥ 14.1K · ⟲ 2.1K · 👁 3.3MView on X ↗

VC Ryan Sarver Details Building an AI Chief of Staff on OpenClaw

How I built a chief of staff on OpenClaw that's better than any human I've hired

Venture investor Ryan Sarver writes up how he built an AI chief of staff on OpenClaw with a markdown-based memory layer and a continuous improvement loop. He offers to open source the system if there is enough interest.

Original post · 12 min read
X ArticleHow I built a chief of staff on OpenClaw that's better than any human I've hired
I'm a VC in the middle of a fundraise, sitting on boards, helping portfolio companies, and angel investing on the side. I've worked with great human EAs and chiefs of staff over the years, so I know what high-leverage support actually looks like. When the first AI APIs came out, I tried to build an AI version of that as a product and couldn't make it work.
When OpenClaw launched I went deep immediately and haven't stopped. I have helped a number of friends set it up and each of them have asked what I have done to configure it and super power it. @ryancarson's post (link in the comments) about how he built his OpenClaw assistant was also great to see, and the response to it convinced me to finally write up what I've been building.
What I have now is more capable than any human chief of staff I've ever worked with. It never forgets a commitment, it handles the small stuff without being asked, flags the important stuff without being told, and it gets better every week. Plus it never sleeps and it never tires. There are still some bumps, but less and less each week.
If any of this is interesting, let me know. If there's enough interest I'll package the whole system up and open source it.
What makes a great chief of staff?
Before I walk through what I've built, it's worth thinking about what a great chief of staff actually does. Not the job description, the real leverage. The best ones I've worked with filtered the noise so only the right things reached me, made sure I walked into every meeting prepared and that nothing fell through after, kept the full picture of what was in flight and flagged what was slipping, tracked relationships and knew where things stood with every important person, and created the daily and weekly rhythm that kept everything moving.
Her name is Stella. She handles all of these, and I'll walk through each one below. But the two things that make my setup genuinely different from other OpenClaw builds are the memory layer underneath it all and the continuous improvement loop that makes the system get better every week. I want to start there because they're what make everything else compound.
Memory: the foundation
Session memory is a lie. Any assistant that treats conversation history as its working context will fail you at the most frustrating moments.
I built two layers. The first is daily notes: one markdown file per day (memory/YYYY-MM-DD.md) serving as a raw log of everything that happened. Meetings attended, decisions made, tasks added and completed, context that came up in conversation. A script called pulls from my sessions throughout the day and writes these automatically.
The second is long-term memory in MEMORY.md, curated by Stella herself. Key people, active projects, lessons learned, decisions made. She periodically synthesizes this from the daily notes, and it's what she reads on startup to orient herself on what matters right now.
Every meeting processed, every email triaged, and every task tracked feeds back into this picture continuously. Without this layer you have a capable assistant with amnesia. With it you have something closer to a person who's been working alongside you for months and never forgets anything.
I've also come to really value that all of this lives in flat markdown files rather than a database. I can open any memory file, read it, edit it if something's wrong, and understand exactly what the assistant knows. I can back the whole thing up to git and restore anything instantly. There's no abstraction layer between me and the assistant's understanding of my world, which means I trust it more and fix things faster when they're off.
Here's where the layers really come together. I'm managing a fundraise involving 100+ LP contacts across multiple countries. Stella tracks the full pipeline, keeps context on each LP and contact, and knows where every relationship stands. For first meetings, I've created a rule that she researches the fund and any recent content they or their partners have published, then preps me with what she found, how it maps to our thesis, and tailored talking points as part of the pre-meeting brief. For ongoing relationships, she knows exactly where they are in the pipeline, what was discussed and committed in our last meeting, and what the key issues are. You can't automate something as critical as a fundraise, but having this kind of structure underneath it means I'm spending my time on the conversations themselves rather than managing the process around them.
Kaizen: the system improves itself
This might be my favorite part, and the thing that makes it feel genuinely different from any assistant I've worked with, human or AI.
Every Friday, a cron job runs research. Stella scans the OpenClaw community, checks for new patterns, looks at what other builders are doing, and saves findings to `memory/kaizen-research-YYYY-MM-DD.md`. On Sunday morning we review it together. She summarizes the week's research, surfaces the top ideas worth tryi… continue on X ↗
♥ 2.2K · ⟲ 215 · 👁 1.2MView on X ↗

X Releases MCP Server, Jon Oringer Shares Setup Guide for OpenClaw

GitHub - xdevplatform/xmcp: MCP server for the X API

Jon Oringer shares steps to connect X to the OpenClaw agent using X's newly released XMCP server, which is hosted on GitHub. The guide covers OAuth setup, a tool allowlist for safety, and test prompts.

Original post · 2 min read
This is huge : @X released an MCP server today..

How to Connect X to your 🦞 :

**Step 1: Run the XMCP Server**

git clone github.com/xdevplatform/xmcp.git
cd xmcp
cp env.example .env

Edit the .env file with your X OAuth consumer key and secret. Set the callback URL to 127.0.0.1:8976/oauth/callback in your X Developer app.

For safety, add an allowlist such as:
X_API_TOOL_ALLOWLIST=searchPostsRecent,createPosts,getUsersMe,getPostsById,likePost,repostPost

Then run:
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python server.py

The server will be available at 127.0.0.1:8000/mcp. Complete the OAuth flow on first run and keep this process active.

**Step 2: Add XMCP in @OpenClaw**

Use the following command:

openclaw mcp set x '{
"url": "127.0.0.1:8000/mcp"
}'

Verify with:
openclaw mcp list
openclaw mcp show x

**Step 3: Test the Integration**

Restart the OpenClaw agent or reload MCP configuration if required.

Test by sending these prompts to OpenClaw in your chat app:
- Search recent posts about MCP on X and summarize the top trends
- Draft and post this thread on X
- Get my X profile information
- Like the latest post from @xdevplatform

OpenClaw will use the XMCP tools automatically when relevant.

**Key Benefits**

- OpenClaw provides persistent memory and works across multiple messaging platforms.
- XMCP delivers standardized access to X API functionality.
- Combined, they enable an agent that can research trends, post content, engage with posts, and report results within your existing chat workflows.

**Safety and Configuration Notes**

Start with a minimal tool allowlist in the XMCP .env file. Expand gradually after testing.
The allowlist can be updated and requires restarting the XMCP server.
Monitor logs in both the XMCP server and OpenClaw for troubleshooting.
X actions performed by the agent are public.

XMCP repository: github.com/xdevplatform/xmcp
OpenClaw MCP documentation: docs.openclaw.ai/cli/mcp
github.comGitHub - xdevplatform/xmcp: MCP server for the X APIMCP server for the X API. Contribute to xdevplatform/xmcp development by creating an account on GitHub.
♥ 2.0K · ⟲ 208 · 👁 312.2KView on X ↗

Citrini Research Publishes Field Report on Strait of Hormuz

Strait of Hormuz: A Citrini Field Trip

Citrini announces that the Field Report from its Analyst #3 on the Strait of Hormuz is live. The post links to the full report on the Citrini Research site.

Original post · 1 min read
Strait of Hormuz: A CitriniResearch Field Trip

The Field Report from Analyst #3 is live.

citriniresearch.com/p/strait-of-hormuz-a-citri…
citriniresearch.comStrait of Hormuz: A Citrini Field TripAnalyst #3 on Assignment
♥ 12.6K · ⟲ 1.3K · 👁 10.6MView on X ↗

Trainer Shares Fitness Tips, Leads With Advice to Quit Alcohol

A fitness coach claiming nine years of experience and 900 clients shares tips for weight loss, beginning with advice to stop drinking alcohol. The post is a short list with no further detail visible.

Original post · 1 min read
After 9 years in the gym and helping over 900 people lose between 9–36 kg, here are the best fitness tips I’ve learned:

1. Stop drinking alcohol.
♥ 6.0K · ⟲ 486 · 👁 5.5MView on X ↗
AI8/10

Google DeepMind Study Measures Manipulation Attacks on AI Agents

Google DeepMind Study Measures Manipulation Attacks on AI Agents

A post describes a Google DeepMind study with 502 participants across eight countries that catalogs 23 attack types against agents, including hidden HTML instructions and steganographic image commands. It reports that frontier models including GPT-4o, Claude and Gemini are vulnerable.

Original post · 5 min read
🚨 BREAKING: Google DeepMind just mapped the attack surface that nobody in AI is talking about.

Websites can already detect when an AI agent visits and serve it completely different content than humans see.

> Hidden instructions in HTML.
> Malicious commands in image pixels.
> Jailbreaks embedded in PDFs.

Your AI agent is being manipulated right now and you can't see it happening.

The study is the largest empirical measurement of AI manipulation ever conducted. 502 real participants across 8 countries.

23 different attack types. Frontier models including GPT-4o, Claude, and Gemini.

The core finding is not that manipulation is theoretically possible it is that manipulation is already happening at scale and the defenses that exist today fail in ways that are both predictable and invisible to the humans who deployed the agents.

Google DeepMind built a taxonomy of every known attack vector, tested them systematically, and measured exactly how often they work.

The results should alarm everyone building agentic systems.

The attack surface is larger than anyone has publicly acknowledged. Prompt injection where malicious instructions hidden in web content hijack an agent's behavior works through at least a dozen distinct channels.

Text hidden in HTML comments that humans never see but agents read and follow. Instructions embedded in image metadata.

Commands encoded in the pixels of images using steganography, invisible to human eyes but readable by vision-capable models.

Malicious content in PDFs that appears as normal document text to the agent but contains override instructions.

QR codes that redirect agents to attacker-controlled content.

Indirect injection through search results, calendar invites, email bodies, and API responses any data source the agent consumes becomes a potential attack vector.

The detection asymmetry is the finding that closes the escape hatch. Websites can already fingerprint AI agents with high reliability using timing analysis, behavioral patterns, and user-agent strings.

This means the attack can be conditional: serve normal content to humans, serve manipulated content to agents.

A user who asks their AI agent to book a flight, research a product, or summarize a document has no way to verify that the content the agent received matches what a human would see.

The agent cannot tell the user it was served different content.

It does not know. It processes whatever it receives and acts accordingly.

The attack categories and what they enable:
→ Direct prompt injection: malicious instructions in any text the agent reads overrides goals, exfiltrates data, triggers unintended actions
→ Indirect injection via web content: hidden HTML, CSS visibility tricks, white text on white backgrounds invisible to humans, consumed by agents
→ Multimodal injection: commands in image pixels via steganography, instructions in image alt-text and metadata
→ Document injection: PDF content, spreadsheet cells, presentation speaker notes every file format is a potential vector
→ Environment manipulation: fake UI elements rendered only for agent vision models, misleading CAPTCHA-style challenges
→ Jailbreak embedding: safety bypass instructions hidden inside otherwise legitimate-looking content
→ Memory poisoning: injecting false information into agent memory systems that persists across sessions
→ Goal hijacking: gradual instruction drift across multiple interactions that redirects agent objectives without triggering safety filters
→ Exfiltration attacks: agents tricked into sending user data to attacker-controlled endpoints via legitimate-looking API calls
→ Cross-agent injection: compromised agents injecting malicious instructions into other agents in multi-agent pipelines

The defense landscape is the most sobering part of the report.

Input sanitization cleaning content before the agent processes it fails because the attack surface is too large and too varied.

You cannot sanitize image pixels. You cannot reliably detect steganographic content at inference time.

Prompt-level defenses that tell agents to ignore suspicious instructions fail because the injected content is designed to look legitimate.

Sandboxing reduces the blast radius but does not prevent the injection itself. Human oversight the most commonly cited mitigation fails at the scale and speed at which agentic systems operate.

A user who deploys an agent to browse 50 websites and summarize findings cannot review every page the agent visited for hidden instructions.

The multi-agent cascade risk is where this becomes a systemic problem.

In a pipeline where Agent A retrieves web content, Agent B processes it, and Agent C executes actions, a successful injection into Agent A's data feed propagates through the entire system.

Agent B has no reason to distrust content that came from Agent A. Agent C has no reason to distrust instructions that came from Agent B.

The injected command travels through the pipeline with the same trust level as legitimate instructions. Google DeepMind documents this explicitly: the attack does not need to compromise the model.

It needs to compromise the data the model consumes. Every agentic system that reads external content is one carefully crafted webpage away from executing attacker instructions.

The agents are already deployed. The attack infrastructure is already being built. The defenses are not ready.
♥ 6.9K · ⟲ 1.6K · 👁 2.0MView on X ↗

Solo App Builder Launches Studio, Cites Rork Marketing Academy

Prajwal Tomar announces IgnytStudio, which aims to ship two AI-built mobile apps per month, and says distribution is the main challenge. He promotes the Rork Max UGC Marketing Academy, a course built on growth tactics behind viral apps.

Original post · 1 min read
You don't realize how BIG this is for solo app builders.

I'm launching my app studio this month (IgnytStudio). The goal is simple: ship 2 mobile apps per month using AI and scale them to actual revenue.

Building apps is the easy part now. A full native iOS app takes me 2-3 days max.

But getting users? That's the part I've been stuck on for weeks.

I can build. I just don't know how to get people to actually download and use what I ship.

Rork just dropped an entire marketing academy built from the growth system behind 2,000+ apps that went viral on TikTok.

This is exactly what was missing from my vibe coding stack.

At this point I'm pretty sure distribution is the ONLY moat left for solo builders.
Rork @rork
We just launched Rork Max UGC Marketing Academy.

The growth system behind 2K+ apps that went viral on TikTok.

Viral hooks, DM scripts, creator hiring guides, and contract templates. Everything to get you to $10K+ MRR

We want you to win.

Built with @wesocialgrowth.
♥ 391 · ⟲ 20 · 👁 69.2KView on X ↗

Andrew Farah Releases Fieldtheory CLI to Sync X Bookmarks Locally

Andrew Farah Releases Fieldtheory CLI to Sync X Bookmarks Locally▶

Andrew Farah shares his first open source project, a free CLI called fieldtheory that downloads and syncs X bookmarks locally so an agent can access them. The post shows install and sync commands plus viz and classify features.

Original post · 1 min read
sharing my first open source project

a CLI for downloading and syncing your X bookmarks locally so your agent can access them. it's free

› npm install -g fieldtheory
› login to your X account in a chrome tab
› ft sync (done!)

bonus:
› ft viz
› ft classify
♥ 4.3K · ⟲ 271 · 👁 556.4KView on X ↗
AI7/10

Google Releases Open Source App to Run Gemma 4 Offline on Phones

Google Releases Open Source App to Run Gemma 4 Offline on Phones▶

Paul Couvert highlights an official Google app that runs Gemma 4 models fully offline on iOS and Android. The app supports text, audio and image input through the E4B and E2B variants.

Original post · 1 min read
Friendly reminder that Google has an official app to run Gemma 4 on your phone.

- 100% open source
- Fully offline and private
- Multimodal with text/audio/image
- Works with Gemma E4B and E2B

And the app is available on both iOS and Android.

Steps and download below
♥ 5.3K · ⟲ 572 · 👁 727.1KView on X ↗

Andrej Karpathy Shares Idea File for Building LLM Knowledge Bases

Andrej Karpathy publishes a gist describing an 'idea file' approach, where an agent builds a personal LLM wiki from shared concepts rather than shared code. The post builds on his earlier thread on using LLMs to compile markdown knowledge bases from raw sources.

Original post · 1 min read
Wow, this tweet went very viral!

I wanted share a possibly slightly improved version of the tweet in an "idea file". The idea of the idea file is that in this era of LLM agents, there is less of a point/need of sharing the specific code/app, you just share the idea, then the other person's agent customizes & builds it for your specific needs.

So here's the idea in a gist format: gist.github.com/karpathy/442a6bf555914893e9891…

You can give this to your agent and it can build you your own LLM wiki and guide you on how to use it etc. It's intentionally kept a little bit abstract/vague because there are so many directions to take this in. And ofc, people can adjust the idea or contribute their own in the Discussion which is cool.
Andrej Karpathy @karpathy
LLM Knowledge Bases

Something I'm finding very useful recently: using LLMs to build personal knowledge bases for various topics of research interest. In this way, a large fraction of my recent token throughput is going less into manipulating code, and more into manipulating knowledge (stored as markdown and images). The latest LLMs are quite good at it. So:

Data ingest:
I index source documents (articles, papers, repos, datasets, images, etc.) into a raw/ directory, then I use an LLM to incrementally "compile" a wiki, which is just a collection of .md files in a directory structure. The wiki…
♥ 26.8K · ⟲ 2.8K · 👁 7.3MView on X ↗

Ruben Recommends Superpowers Brainstorming Skill for Creative Work

brainstorming — obra/superpowers

Ruben replies to Paul Solt that he uses the brainstorming skill from the obra/superpowers repository on skills.sh. The skill is meant to be used before creative work such as building features or modifying behavior, to explore user intent.

Original post · 1 min read
@PaulSolt I use brainstorming all the time

skills.sh/obra/superpowers/brainstorming
skills.shbrainstorming — obra/superpowersYou MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent,…
♥ 119 · ⟲ 4 · 👁 34.1KView on X ↗
AI7/10

Google Publishes Visual Guide to Gemma 4 Architectures

Google Publishes Visual Guide to Gemma 4 Architectures

Google Gemma shares a visual guide explaining the new Gemma 4 architectures and how they process text, images, and audio in the smaller models.

Original post · 1 min read
Who wants to know how Gemma 4 works?

This visual guide breaks down the new architectures and how they process text, images, and (for the smaller models) audio.

👇
♥ 4.2K · ⟲ 480 · 👁 246.9KView on X ↗

Addy Osmani Releases Open-Source Agent Skills for Coding Agents

Addy Osmani Releases Open-Source Agent Skills for Coding Agents

Addy Osmani of Google released Agent Skills, 19 engineering skills and 7 slash commands that enforce specs, tests, and reviews for coding agents including Claude Code and Cursor. The project is free, open source, and installable via npx.

Original post · 1 min read
🚨 You need to see this.

@addyosmani from Google just dropped his new Agent Skills and it's incredible.

It brings 19 engineering skills + 7 commands to AI coding agents, all inspired by Google best practices 🤯

AI coding agents are powerful, but left alone, they take shortcuts.

They skip specs, tests, and security reviews, optimizing for "done" over "correct." Addy built this to fix that.

Each skill encodes the workflows and quality gates that senior engineers actually use: spec before code, test before merge, measure before optimize.

The full lifecycle is covered:

→ Define - refine ideas, write specs before a single line of code
→ Plan - decompose into small, verifiable tasks
→ Build - incremental implementation, context engineering, clean API design
→ Verify - TDD, browser testing with DevTools, systematic debugging
→ Review - code quality, security hardening, performance optimization
→ Ship - git workflow, CI/CD, ADRs, pre-launch checklists

Features 7 slash commands: (/spec, /plan, /build, /test, /review, /code-simplify, /ship) that map to this lifecycle.

It works with:
✦ Claude Code
✦ Cursor
✦ Antigravity
✦ ... and any agent accepting Markdown. Baking in Google-tier engineering culture (Shift Left, Chesterton's Fence, Hyrum's Law) directly into your agent's step-by-step workflow!

`npx skills add addyosmani/agent-skills`

Free and open-source.

Repo link in 🧵↓
♥ 2.5K · ⟲ 339 · 👁 423.0KView on X ↗

Analysis Traces Rohit Sharma's Late-Career Batting Evolution

The Man Who Rewrote His Own Obituary

A cricket analyst examines how Rohit Sharma reworked his batting between 2021-23 and 2024-26 after bowlers began exploiting his weaknesses, citing strike rates, Runs Above Average and pitch maps.

Original post · 6 min read
X ArticleThe Man Who Rewrote His Own Obituary
We've seen that most batters tend to hit a saturation or a dip around the age of 35. And honestly, that's not surprising as it happens with many cricketers. Some choose to retire because the game catches up with them, and some just settle into whatever's left.
As a Rohit Sharma fan for a long time, it was disappointing to see that there was always this expectation, created widely throughout the world, that the same thing would happen with him. And that expectation wasn't wrong either, because the problems were there. He was struggling in the powerplay, with left arm seamers, against slow left arm bowlers. After playing 200 T20s, there is also enough data, enough evidence in your technique and your methodology, for bowlers to finally start figuring you out really well.
And most of the time, players don't really evolve much at that point in their careers. Bowlers know what strategies work against them, the weaknesses become well established.
But what Rohit did was genuinely different.
He looked at that reality and found a way to flip the script in a way that very few batters have ever managed. What happened to Rohit between 2021-23 and 2024-26 is, for me, one of the most remarkable technical and psychological evolutions you'll ever see from a batter at the back end of a legendary career. And the numbers make it very clear.
The Obituary That Was Being Written
Think back to the years 2021-23, Rohit was still a formidable force on the field but behind the scenes, a strategy was developing in dressing rooms across the globe.
Using your slow left-arm orthodox bowlers attack him from around the wicket, keep it on the stumps, and deliver those good-length balls. Guess what? It actually worked.
During that period, Rohit faced left-arm fast bowlers in the powerplay with a strike rate of 103.02, but his Runs Above Average (RAA) was a concerning -7.85. That negative figure is crucial as it shows he was underperforming compared to the average batter against those bowlers. The plan was clearly effective.
Now, when it came to slow left-arm orthodox bowlers, he managed to score 95 runs off 77 balls, with a strike rate of 123. There were control issues, dot balls were piling up, and the bowlers had a clear idea of where to target him. The pitch maps from that time tell the whole story. Good length, right on the stumps. Four dismissals in that one area. Bowlers were lining up to exploit that corridor. When you can dismiss one of the best powerplay batters in the world by repeatedly bowling the same delivery, you keep doing that, right?

The Decision Most Champions Never Make
Most elite batters, when they sense vulnerability, retreat into their strengths. They tend to leave the problem areas alone and hope nobody exploits them.
Interestingly, Rohit made the difficult choice of outsmarting the bowlers, which would later prove to be crucial in India's World cup win in 2024.
The centrepiece of that rebuild? The sweep shot.
Against spin bowlers in 2021-23, Rohit’s sweep was a minor part (18% of runs) of his arsenal. By 2024-26, it accounts for 28.19% of all his runs against spin in the powerplay. That’s a 10% jump. But the biggest revamp was the control percentage: 83.33%, which was around 65% before.
You do not play an attacking shot that contributes nearly a third of your runs against quality spin bowling with 83% control by accident. That takes hours in the nets, tweaking the trigger movement, adjusting the head position, recalibrating the risk-reward every single time.
The SLA Problem and how Rohit solved it
Let’s zoom in on slow left-arm orthodox bowlers, because this is where the story gets genuinely interesting and worth studying.

Against slow left arm bowlers alone, the sweep shot has given him a huge boost. His SR of 167 is 44 points higher than an avg batter against these kinds of bowlers.

8.2% runs in the square leg region in 2021-23 and that number has skyrocketed to 28% of runs in 2024-26. His reverse sweep and late cuts using the pace of the ball also have increased the runs he scores against these bowlers in third man region.
Now add this: that good-length delivery on the stumps that used to be his weakness? The one where four dismissals were logged in that single grid?
Rohit against spin when the ball was bowled on a good length on stump line: He used to play 15% of those balls in the square leg region in 2021-23. During 2024-26 this number has risen to 21.3% which is 8.3% more than an average Right-handed batter against spin on that very line and length.

The 2 key variations of left arm orthodox are both taken to the cleaners now by him. His performance against those 2 types of deliveries has risen significantly.

In 2021-23, when Rohit faced left-arm spin deliveries, he scored 106 runs off 100 balls. Strike rate of 106. Dot ball percentage of 40%. Balls per boundary of 8.33.
By 2024-26: 70 runs off 38 balls. Strike rate of 184.21. Dot ball percentage collapsed to 15.79%. Balls per boundary down to 3.… continue on X ↗
♥ 773 · ⟲ 279 · 👁 296.2KView on X ↗

Tony Fadell Argues Product Management and Marketing Should Be One Job

Former Apple engineer Tony Fadell argues that splitting product management and product marketing is a mistake, citing Steve Jobs and Greg Joswiak's customer empathy as the model for owning both the product and its story.

Original post · 1 min read
Most tech companies break out product management and product marketing into two separate roles: Product management defines the product and gets it built. Product marketing wires the messaging- the facts you want to communicate to customers- and gets the product sold. But from my experience that's a grievous mistake. Those are, and should aways be, one job.

There should be no separation between what the product will be and how it will be explained- the story has to be utterly cohesive from the beginning. Your messaging is your product. The story you're telling shapes the thing you're making.

I learned story telling from Steve Jobs. I learned product management from Greg Joswiak. Joz, a fellow Wolverine, Michigander, and overall great person, has been at Apple since he left Ann Arbor in 1986 and has run product marketing for decades. And his superpower- the superpower of every truly great product manager- is empathy. He doesn't just understand the customer. He becomes the customer.

So when Joz stepped into the world with his next-gen iPod to test it out, he fiddled with it like a beginner. He set aside all the tech specs- except one: battery life.

The numbers were empty without customers, the facts meaningless without context.

And, that's why product management has to own the messaging. The spec shows the features, the details of how a product will work, but the messaging predicts people's concerns and finds way to mitigate them.

- #BUILD Chapter 5.5 The Point of PMs
♥ 2.7K · ⟲ 255 · 👁 949.4KView on X ↗

Paul Solt Shares Workflow for Building Apps With Codex Sans Xcode

How I Build Apps With Codex Without Opening Xcode

iOS developer Paul Solt describes an Agent Skill called AppCreator and a Makefile-based workflow using xcbeautify so Codex can build, test, and run iPhone and Mac apps with clean pass/fail output.

Original post · 6 min read
X ArticleHow I Build Apps With Codex Without Opening Xcode
Do you want to build iOS or macOS apps with Codex?

I have a new Agent Skill that will help you make apps. Without this skill you're going to waste a lot of time. Let me explain.
I was building a Dangerous Spider app with Codex when I noticed the agent kept missing its own build and test failures. It was looking at the wrong status codes. It genuinely couldn't find the error and would say everything worked (when it didn't).
Xcode compiler build output is extremely verbose. Actual errors get buried in thousands of lines of text. Trying to find an error in Xcode build output is like finding a needle in a haystack.

When you layer in agents, you're wasting time and context by asking them to find the errors.
Agents need to know what worked and what didn't work. It needs to be pass/fail, so I created an agent-designed workflow to do just that.
Here's the 7-step workflow I use every day:
1. Make Xcode Projects Agent Friendly with AppCreator
I built an Agent Skill called AppCreator. Run it once, and it scaffolds a new Xcode project or retrofits an existing one. Now your project is agent-ready.

At the heart of it: a `Makefile` wraps the CLI `xcodebuild` commands using xcbeautify. Clean, readable output from Xcode app builds and tests. The agent sees what failed, fixes it, and moves on. No verbose output to search. It works for iPhone and Mac apps.
Download and install the AppCreator Skill to make your app project agent-friendly.
2. make Is the Only Build Command I Use
The skill is designed so that one command builds and runs your app:
make

Your agent knows how to work with Makefiles, so this is just the starting point. You can extend it however you want.
I ask agents to set the default action to "build-and-run", and to use special build scripts so the freshly built app is always relaunched, just like Xcode.
In Codex CLI, type "make" to check the agent's work.
In the Codex app, set up a custom run action: Click on the Play button and set it to: make.

Finding the setting later requires a few more steps: Go to Environments → Project → View → Edit Local Actions → Actions → Set Action Script to `make`.

With a Makefile, you have a fast, repeatable way to build and run your apps (via the up-arrow on the CLI or the Run button in the Codex app).
3. Just Talk to It
If you want to get good results, you need to actively steer your agent.
Agents are good, but they'll do things you don't want, and the only way to prevent that is to steer the ship as they work.
I frequently double-check the work and redirect when I see agents doing the wrong thing.
I use Wispr Flow to talk out loud — describe what I want, how it should behave, what needs to change. After an agent hands back work, I start Wispr Flow as I play-test the new changes. This gives me the chance to talk through what works and what doesn't, and then, when I'm done, I can paste the transcript directly into the Codex app.
Plan mode is helpful for sparking thoughts about how features and edge cases. However, I have found that the plan mode isn't good enough.

Instead, I would use plan mode to help you think through the feature as a starting point. Use it to spark discussion so you can refine which features you actually build.
Software is nuanced, and if you take it feature by feature, you're going to get better software in the end.
4. Tests Keeps Agents Accountable
Without tests, agents can get sloppy. Tests allow agents to catch their own mistakes.
Ask Codex to write unit tests as it builds. Your goal is fast tests. UI tests are helpful for verification, but they can be annoying and slow to run. I like to have agents use UI tests to catch errors that are impossible to test with unit tests alone.

My recommendation is that you separate your tests into at least two targets:
make test
make ui-test
When an agent runs UI tests, it takes over your machine, which can be disruptive if you need to do anything else (This is where having a second Mac can be useful).
UI tests will slow everything down, so it's best to test them only at hand-off points. When your agent has wrapped up one task or several tasks. Don't do full regression testing for incremental work; instead, ask the agent to test only the smallest subset to verify their work and remain fast.
Have your agents read this article from Peter Steinberger, @steipete: Running UI Tests on iOS With Ludicrous Speed. Using that, you can help keep your test suite from becoming a bottleneck.
5. Log Runtime Results
Another indispensable tool is the use of logs and app artifacts. Have your agent add logging to your app. This gives it another tool for seeing what went wrong in real time.
Agents can stream development logs to a file or read them in real time to fix problems that are not immediately obvious.

When my agent repeatedly fails to complete a task properly, I know it's time to introduce more detailed logs.
Just ask your agent:
Please add verbose logs around XYZ so that you can see what is happening and fix the… continue on X ↗
♥ 810 · ⟲ 76 · 👁 467.5KView on X ↗

Open-Multi-Agent Framework Reimplements Claude Code Orchestration Patterns

Open-Multi-Agent Framework Reimplements Claude Code Orchestration Patterns

Ivan Burazin highlights an open-source, model-agnostic multi-agent framework built from scratch by a former PM after the Claude Code source leak, with in-process orchestration deployable to serverless, Docker, or CI/CD.

Original post · 1 min read
After the Claude Code source code leak, a former PM extracted its multi-agent orchestration system into an open source model agnostic framework.

He studied the architecture, focused on the multi-agent orchestration layer (the coordinator that breaks goals into tasks, team system, message bus, task scheduler with dependency resolution), and reimplemented these patterns from scratch as a standalone open source framework without infringing on Anthropic's code.

The result is what @JackChen_x calls an "open-multi-agent." Unlike claude-agent-sdk, which spawns a CLI process per agent, this runs entirely in-process and can be deployed anywhere (serverless, Docker, CI/CD)

Check it out: github.com/JackChen-me/open-multi-agent
♥ 3.4K · ⟲ 538 · 👁 555.9KView on X ↗
AI7/10

Levelsio Warns Vibe-Coded Apps Pose Production Security Risks

Developer @levelsio says vibe coding into production is dangerous and plans to revoke database access and run apps with minimal privileges. He quotes a post reporting that LLMs hallucinate package names about 18-21% of the time, enabling 'slopsquatting' attacks.

Original post · 1 min read
Okay honestly this makes vibe coding into production very dangerous, you guys were all right

I think what I'll do is cut off all access to DBs and run it as a user with almost no privileges
Basel Ismail @BaselIsmail
URGENT PSA - New supply chain attack vector that I found WILD > AI LLMs hallucinate package names roughly 18-21% of the time.

Hackers have started pre-registering those hallucinated names on PyPI and npm with malicious payloads; they call it "slopsquatting"

You can only imagine what's next
♥ 1.6K · ⟲ 72 · 👁 430.8KView on X ↗

Cathryn Open-Sources Ops Toolkit for Keeping OpenClaw Running

Cathryn releases openclaw-ops, a set of open-source scripts that repair gateway issues, watch and restart the gateway, check configuration health, scan for security gaps, and audit third-party ClawHub skills.

Original post · 1 min read
🦞 openclaw-ops

tired of babysitting your OpenClaw?

I just open-sourced my ops layer. It fixes the gateway, exec approvals, broken crons, stuck sessions, channel issues, and security gaps.

• heal.sh — one-shot fix for the most common gateway issues (auth, exec approvals, crons, stuck sessions)
• watchdog.sh — runs every 5 min, restarts gateway if down, escalates after 3 failures
• watchdog-install.sh — installs the watchdog as a macOS LaunchAgent so it survives reboots
• check-update.sh — detects version changes, explains what config broke and why; --fix to auto-apply
• health-check.sh — declarative URL/process checks for gateway-adjacent services and workers
• security-scan.sh — config hardening score (0–100), drift detection, credential scan
• skill-audit.sh — static audit for third-party ClawHub skills before you install them

basically everything I built for myself since January to stop me from tearing my hair out 💀

link below 🫶
Cathryn @cathrynlavery
🦞 Openclaw update fix

If your agents are hitting exec approval walls after the latest update, the fix is three settings:

In exec-approvals.json defaults:
- security: "full"
- ask: "off"
- askFallback: "full"

In openclaw.json:
- tools.exec.security: "full"
- tools.exec.strictInlineEval: "false"

Then restart gateway.

The allowlist wildcard * alone isn't enough. There's a second policy layer that gates complex commands independently.
♥ 559 · ⟲ 44 · 👁 104.0KView on X ↗

Cursor Launches Version 3 Built for Agent-Written Code

Cursor Launches Version 3 Built for Agent-Written Code▶

Cursor announces Cursor 3, pitched as simpler and more powerful and designed for a world where code is written by agents while retaining the depth of a development environment. The post is a video announcement.

Original post · 1 min read
We’re introducing Cursor 3. It is simpler, more powerful, and built for a world where all code is written by agents, while keeping the depth of a development environment.
♥ 10.1K · ⟲ 972 · 👁 2.9MView on X ↗

Ashu Garg Argues Decision Traces Will Reshape Enterprise Software

Google's 20-year secret is now available to every enterprise

Ashu Garg and Jaya Gupta argue that SaaS multiples are compressing as AI commoditizes features, and that enterprises can build durable compounding loops by capturing decision traces rather than just end-state records.

Original post · 11 min read
X ArticleGoogle's 20-year secret is now available to every enterprise
Why decision traces will reshape B2B the way behavioral data reshaped B2C
Our latest thinking on context graphs, developed with my partner @JayaGup10.

Consumer platforms built one of the most powerful business models of the last two decades around a compounding loop: every user interaction became a signal that improved the system. Netflix, Meta, Amazon, TikTok, and Google did not just record outcomes. They instrumented behavior with extraordinary granularity—what you clicked, what you ignored, what you hovered over, what you abandoned, what brought you back—and fed those signals into systems that learned. That loop—capture, learn, improve, capture again—became one of the great compounding assets of the internet era.
Enterprise software has never had an equivalent loop. Not because enterprise decisions are less frequent, but because they were harder to observe.
Consumer systems operate inside controlled interfaces where a single user acts within a product the company fully owns. Enterprise decisions are fundamentally different: they are multiplayer negotiations across sales, finance, legal, operations, security, and management—each carrying different incentives, different authority, and different constraints. Sales wants velocity. Finance wants margin. Legal wants precedent control. These decisions are negotiated, not merely clicked. To date, enterprises have lacked instrumentation of the reasoning that connected action to outcome.
B2C companies have been compounding behavioral signals for two decades. B2B companies largely have not. Now, for the first time, that is starting to change.
The old model is breaking
SaaS multiples have compressed because AI is commoditizing the feature layer that justified premium pricing. When an LLM can generate a competent first draft of almost any workflow, the value of “better UI on a known process” collapses — and companies whose moats were features, not data, are the ones being marked down. They built workflows but never built compounding loops
The question is what replaces features as the durable source of enterprise value. The answer is the compounding loop that enterprise software never had, built not on behavioral traces, but on decision traces.
What enterprise software actually captured, and what it missed
Enterprise systems were built to record end state, not reasoning. A discount field tells you the final number, not why that number was justified. A redlined contract tells you the final clause, not which fallback positions were rejected along the way. A resolved ticket tells you the incident is closed, not why one escalation path was chosen over another. Decision traces sit in that missing layer between event and outcome. A context graph is what happens when that layer becomes structured, queryable, and connected across systems, actors, and time.
The relevant signals were also sparse, fragmented, and embedded inside human workflows rather than captured as first-class telemetry. Enterprise decisions happened partly in a meeting, partly in someone’s head, partly in an email thread, partly in a side conversation, and partly inside systems that did not talk to one another.
And there was no reason to store it. Decision data was treated as process exhaust—ephemeral, disposable—because no system existed that could learn from it. Even when fragments were captured, they rarely compounded. Companies had transcripts, email threads, comments, and approvals, but no practical way to extract structured decision artifacts from them, connect them across systems, and link them to outcomes. The raw material existed in pieces, but the loop did not.
What’s changed
Enterprise work now lives on instrumentable surfaces. Work has become distributed and asynchronous. Decisions increasingly get made in comment threads, document suggestions, ticket histories, approval flows, and call recordings. Reasoning that once lived only in someone’s head now leaves an increasingly rich trail in the workflow itself.
Language models make the unstructured data computable. For years, companies had transcripts, chat logs, document comments, and ticket histories, but these were mostly searchable, not learnable. Now an LLM can extract decision artifacts from them.. Language models do not eliminate the need for structure or evaluation, but they make it possible to turn previously inert collaboration data into something a system can reason over.
Agents create decision checkpoints automatically. This is the most important shift. Agents propose actions inside workflows, which humans approve, modify, or escalate. An agent drafts a pricing proposal; the sales rep adjusts the discount from 25% to 30% and adds a note: “competitive pressure from Vendor X, need to match their offer.” That edit is a decision trace.
The model’s proposal is a structured prior—what the system thought was right. The human’s modification is the judgment signal—what actually matters that the model missed. As agents insert themselves into more… continue on X ↗
♥ 669 · ⟲ 88 · 👁 453.8KView on X ↗

Jaya Gupta Says Enterprise Decision Traces Can Mirror Consumer Data Loops

Google's 20-Year Secret Is Now Available to Every Enterprise

Jaya Gupta's essay, co-developed with Ashu Garg, contends that enterprise software lacks the behavioral-signal feedback loops consumer platforms enjoy and that capturing reasoning behind decisions could create one.

Original post · 11 min read
X ArticleGoogle's 20-Year Secret Is Now Available to Every Enterprise
Consumer platforms built one of the most powerful business models of the last two decades around a compounding loop: every user interaction became a signal that improved the system. Netflix, Meta, Amazon, TikTok, and Google did not just record outcomes. They instrumented behavior with extraordinary granularity, what you clicked, what you ignored, what you hovered over, what you abandoned, what brought you back and fed those signals into systems that learned. That loop: capture, learn, improve, capture again - became one of the great compounding assets of the internet era.
Enterprise software has never had an equivalent loop. Not because enterprise decisions are less frequent, but because they were harder to observe.
Consumer systems operate inside controlled interfaces where a single user acts within a product the company fully owns. Enterprise decisions are fundamentally different: they are multiplayer negotiations across sales, finance, legal, operations, security, and management with each carrying different incentives, different authority, and different constraints. Sales wants velocity. Finance wants margin. Legal wants precedent control. These decisions are negotiated, not merely clicked. To date, enterprises have lacked instrumentation of the reasoning that connected action to outcome.
B2C companies have been compounding behavioral signals for two decades. B2B companies largely have not. Now, for the first time, that is starting to change.

The old model is breaking
SaaS multiples have compressed because AI is commoditizing the feature layer that justified premium pricing. When an LLM can generate a competent first draft of almost any workflow, the value of "better UI on a known process" collapses — and companies whose moats were features, not data, are the ones being marked down. They built workflows but never built compounding loops
The question is what replaces features as the durable source of enterprise value. The answer is the compounding loop that enterprise software never had, built not on behavioral traces, but on decision traces.
What enterprise software actually captured, and what it missed
is what happens when that layer becomes structured, queryable, and connected across systems, actors, and time.Enterprise systems were built to record end state, not reasoning. A discount field tells you the final number, not why that number was justified. A redlined contract tells you the final clause, not which fallback positions were rejected along the way. A resolved ticket tells you the incident is closed, not why one escalation path was chosen over another. Decision traces sit in that missing layer between event and outcome. Acontext graph
The relevant signals were also sparse, fragmented, and embedded inside human workflows rather than captured as first-class telemetry. Enterprise decisions happened partly in a meeting, partly in someone's head, partly in an email thread, partly in a side conversation, and partly inside systems that did not talk to one another.
And there was no reason to store it. Decision data was treated as process exhaust—ephemeral, disposable—because no system existed that could learn from it. Even when fragments were captured, they rarely compounded. Companies had transcripts, email threads, comments, and approvals, but no practical way to extract structured decision artifacts from them, connect them across systems, and link them to outcomes. The raw material existed in pieces, but the loop did not.
What's changed
Enterprise work now lives on instrumentable surfaces. Work has become distributed and asynchronous. Decisions increasingly get made in comment threads, document suggestions, ticket histories, approval flows, and call recordings. Reasoning that once lived only in someone's head now leaves an increasingly rich trail in the workflow itself.
Language models make the unstructured data computable. For years, companies had transcripts, chat logs, document comments, and ticket histories, but these were mostly searchable, not learnable. Now an LLM can extract decision artifacts from them.. Language models do not eliminate the need for structure or evaluation, but they make it possible to turn previously inert collaboration data into something a system can reason over.
Agents create decision checkpoints automatically. This is the most important shift. Agents propose actions inside workflows, which humans approve, modify, or escalate. An agent drafts a pricing proposal; the sales rep adjusts the discount from 25% to 30% and adds a note: "competitive pressure from Vendor X, need to match their offer." That edit is a decision trace.
The model's proposal is a structured prior, what the system thought was right. The human's modification is the judgment signal, what actually matters that the model missed. As agents insert themselves into more workflows, more judgment is forced to become explicit through edits, approvals, exceptions, and overrides. The instrumentation is no longer o… continue on X ↗
♥ 658 · ⟲ 120 · 👁 355.6KView on X ↗

Alfred Lin Revisits 1997 Prediction That Amazon Would Kill Walmart

Alfred Lin Revisits 1997 Prediction That Amazon Would Kill Walmart

Alfred Lin acknowledges his 1997 prediction that Amazon would kill Walmart was wrong, noting Walmart is now about 30 times larger, and lists failed e-commerce firms, rising acquisition costs and the value of physical presence as overlooked factors.

Original post · 1 min read
In 1997, I declared that Amazon would kill Walmart. Today, Walmart is 30 times larger than it was 30 years ago. The world was messier than the story:

- E-commerce companies also failed
- Customer acquisition costs online kept rising
- Certain categories had persistent try-before-you-buy dynamics
- Physical presence created brand equity that digital alone could not

What I should have asked: what would have to be true for this story to be wrong?
Alfred Lin @Alfred_Lin
Beware of Simple Narratives — Simple narratives can guide action and unify thinking, but they often obscure more than they reveal.

We've been taught to tell simple narratives. They are catchy and memorable. Let's be honest. They
♥ 433 · ⟲ 47 · 👁 127.7KView on X ↗

Alibaba Open-Sources Flyai Skill Bringing Travel Search to Coding Agents

Alibaba Open-Sources Flyai Skill Bringing Travel Search to Coding Agents

Tom Dörr shares a GitHub repository from Alibaba, flyai-skill, which adds travel search capabilities inside AI coding agents. The project is an open-source agent skill that developers can contribute to.

Original post · 1 min read
Travel search inside AI coding agents

github.com/alibaba-flyai/flyai-skill
github.comGitHub - alibaba-flyai/flyai-skill: fly ai agent skillfly ai agent skill. Contribute to alibaba-flyai/flyai-skill development by creating an account on GitHub.
♥ 495 · ⟲ 47 · 👁 29.0KView on X ↗

Job Search Coach Pitches Referral-First Strategy and $49 AI Job Search System

Job Search Coach Pitches Referral-First Strategy and $49 AI Job Search System

Aakash Gupta, who says he has placed candidates at OpenAI, Anthropic, Google and others, argues that referrals and tailored prototypes beat cold applications. He promotes a $49 system of 18 Claude Code skills for resumes, interview prep and networking.

Original post · 2 min read
I've placed people at OpenAI, Anthropic, Google, Meta AI, Databricks, and Stripe in the last year. They all did the same thing that 90% of job searchers skip.

They built referral paths before submitting a single application.

The average cold application callback rate in 2026 is around 2-4%. With a warm intro, it's 5x that. Every candidate I coached who got an offer at a top company had a referral on file before the resume went in. Every single one.

But here's what most people get wrong about networking for jobs. They send "I'd love to pick your brain" to strangers. One message, no follow-up, silence forever. The people who land offers send 25 personalized connection requests per week, rotate across target companies, follow up on day 3, 7, and 14, and ask for the referral only after building context.

The resume game broke too. I tested every paid AI resume tool on the market ($20-40/mo). They all do one of two things: invent experience you don't have (which gets you blacklisted when the interviewer checks) or swap keywords on a generic template (which recruiters now spot in a 6-second scan).

The move almost nobody makes: a 1-pager analyzing the company's product + a working prototype of your recommendation. 90 minutes. I've seen this land interviews where cold applications failed completely. The catch is specificity. If it could've been written for any company, a hiring manager told me it's actually a negative signal.

I spent 6 months building a system that automates all of this. 18 Claude Code skills. Resume tailoring from your real experience only. Interview prep with insider data from 250 companies. Mock interviews that compound (after 5-6 real interviews, the system knows your weakest question types). Networking sequences. Negotiation.

$49 once. 45-minute setup. Then 20 minutes a day.

Honest caveat: the system is only as good as your experience library. If you skip the 10-minute setup where you load your real career history, every output will be generic. The input is the bottleneck, not the tool.

Full deep dive: news.aakashg.com/p/job-search-os
♥ 426 · ⟲ 46 · 👁 126.0KView on X ↗