Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Uber Details Software Factory Cost Model Across Agent Layers

Running a Software Factory Efficiently at Uber Scale

Uber Engineering publishes an article by Uday Kiran on its software factory, reporting over 70% of pull requests attributed to agents, 3,600 agent skills, and a cost equation showing cost per 1,000 requests down about 34% from peak.

Original post · 12 min read
X ArticleRunning a Software Factory Efficiently at Uber Scale
Post author: @udaykiran

Introduction
AI tools are now embedded in every phase of software development at Uber. More than 70% of pull requests are attributed to local or cloud agents. Engineers have built over 3,600 agent skills across the software development life cycle, and executed more than 30K agent skill executions per day.
At the AI Engineer 2026 conference, we shared our vision for the Software Factory and the building blocks and managed agents we are building across the lifecycle. As we progress on that vision, a growing share of sessions aren’t initiated by humans, but by automated managed agents handling code review, self-healing CI failures, completing E2E PRs with visual validation, triaging on-call alerts, debugging incoming bugs, and handling a variety of code maintenance tasks with human reviews/escalations.
As shown in Figure 1, from February to Aug 2026, weekly active users across all agentic offerings across all our employees (engineers & non-engineers) grew 7x, and weekly agentic requests grew 9.4x. Meanwhile, our total AI spend has relatively stabilized since April due to optimizations across the board.

Since adoption, workload mix, and model upgrades are all continuously changing, isolating our own optimization gains means holding one model fixed, since behavior shifts with every upgrade and model family. We did that from February to July: cost per 1,000 model requests is down almost 34% from its peak, and cost per session is down 52% from its June peak.

This blog walks through how we think about our software factory: the four layers agent sessions run in, the cost equation we use to decompose spend, how we measure each term, and how we optimize those terms across every layer.
All pricing and vendor metrics in this comparison are based on publicly available information, with cost efficiency gains driven by routing our internal Uber workloads more intelligently within standard tier-pricing. While specific cost reductions we measure are unique to our environment and your mileage may vary depending on your codebase, team size, and agent workflows, the methodology of benchmarking real work and optimizing for accuracy and cost is universally applicable.
The Software Factory and Its Cost Equation
Four Layers of Agent Usage
We organize AI usage into four layers, from the most specialized to the most general. As shown in Figure 3, the higher the layer, the more control we have over cost, quality, and model selection.

The Cost Equation
Across any of the layers above, we can decompose the cost of an agentic session into the following terms, which we could measure and optimize independently.

The first two terms represent adoption & engagement, which we want to keep growing across our overall user base, whether users use it interactively or agents handle tasks on their behalf. The three middle terms provide opportunities for optimization: the work the agent does on its own behalf, on top of the request an engineer actually made. That is where most of our effort goes. This includes mechanisms that help agents plan faster, reduce unwanted turns or errors, optimize input tokens, and more.
How We Measure
Below is the full set of metrics we track weekly and monthly that enable us to forecast & plan our efforts short-term and long-term.

Optimization Levers
In the following sections, we detail the key levers we used to optimize each part of the cost equation. Some of these levers affect one or more rows in the cost equation.

Optimizing Price / Token
The vendor sets the token price. We pick which model runs which workload. Across all our managed agents’ layers, we pick the model that’s most Pareto efficient for that workload. For us, Pareto efficient means cost/completed task, output quality, and model reliability.
Benchmark-Driven Model Selection
Model selection happens in four steps, the same for every managed agent we run.
Build a benchmark out of the agent’s real work.
Run the agent on a harness that serves any model, frontier or open-weight, behind one interface.
Move to whatever is Pareto optimal, and keep moving. The frontier shifts every few weeks.
Looking ahead, we continually refine our workload performance by leveraging aggregated insights from our managed agents to test and deploy various model routing strategies.
For example, we use uReview, which handles AI code review for all pull requests. We built its benchmark from real pull requests with known bugs and graded them easy, medium, and hard. We score precision, recall, and F1 against those bugs, plus cost per review, latency, timeouts, and noise. As shown in Figure 5, switching models improved our F1 while dramatically reducing cost/PR. In the figure, the dashed line is the Pareto frontier. Everything below and left of it is beaten by something cheaper or better.

Using thousands of real-world PRs across our large monorepos, we internally also have an Uber SWE Benchmark that runs frontier and open-weight models across differe… continue on X ↗
♥ 4.8K · ⟲ 777 · 👁 2.7MView on X ↗

More in Agents & Dev Tools

Developer Rebuilds Seven Adobe Apps in Rust Using Opus 5.5

Peter Yang highlights a developer who reimplemented seven Adobe apps, including Photoshop, Premiere and Lightroom, in Rust with Claude Opus 5.5 and open-sourced them. The developer believes they can match Adobe's features within months, against Adobe's $840 yearly all-apps plan.

Original post · 1 min read
It's insane to watch AI blow apart closed source software and games.

4 examples from the past month:

1. 7 of Adobe's biggest apps, including Photoshop, Premiere, and Lightroom, have been partially rebuilt in Rust with Opus 5.5 and open sourced. It's still early, but the developer thinks they can match Adobe's features within months. Adobe's all-apps plan costs $840/year.
Miguel Ángel Durán @midudev
Todos los productos de Adobe reimplementados desde cero, gratuitos y de código abierto

→ getartcraft.com/apps
♥ 50 · ⟲ 2 · 👁 10.7KView on X ↗

Vercel's Guillermo Rauch Explains Turborepo's Migration From Go to Rust

Guillermo Rauch says Vercel moved Turborepo from Go to Rust, a migration that was controversial internally due to human costs. He argues that with AI agents the calculus has changed, so what is best for humans is no longer necessarily best for business.

Original post · 1 min read
DHH is fundamentally right about Rust. For context, Vercel has been undergoing a Rust-ification (carcinization, technically 🦀) for a while.

One of the first projects we migrated was Turborepo, from Go to Rust¹. The migration completed, but the RoI was actually quite controversial internally.

While Rust was in our eyes better for low-level OS access, something crucial for a build system like Turbo, the human migration costs were very sustantive.

Go is very fast. It's beautifully designed. It's easy to iterate on. We were very conflicted about the migration, because it was *humans* writing the code, *even if we knew Rust was a better choice*.

The calculus has now changed. What's "best for humans" is no longer necessarily "best for business".

FWIW, it's also quite unlikely that Rust is the end-all-be-all toolchain. I'm quite certain there's greener pasture ahead, because Rust itself was designed before the 'supersonic tsunami' of agents hit.

¹ https​://vercel.com/blog/how-turborepo-is-porting-from-go-to-rust
♥ 3.2K · ⟲ 152 · 👁 352.7KView on X ↗

Integer Multiplication Algorithm Bound Tightened Repeatedly With Astra

A post reports that a user running ChatGPT Astra in a loop is repeatedly breaking records for integer multiplication algorithms. It quotes an update to OpenAI problem #109 that tightens the constant from 2^-182 to 2^-59, a roughly 500,000-fold improvement over the previous result.

Original post · 1 min read
This guy has 6.1 Astra running in a loop and is breaking the record for integer multiplication algorithms every few hours lmaooooo.
Doug Colkitt @0xdoug
We are publishing an update to OpenAI problem #109 Integer multiplication) with another substantial further tightening:

κ = 2⁻⁵⁹ (from OpenAI’s original κ = 2⁻¹⁸²)

Approximately 500 thousand fold improvement over our previous result and a 2¹²³ fold improvement over the original OAI result.

The latest redesigned the finite network to share intermediate computations and scratch space, then tightened the recursion and Gaussian estimates.
♥ 4.1K · ⟲ 118 · 👁 167.4KView on X ↗

Boris Cherny Says Prompting Claude Should Feel Like Talking to a Coworker

Boris Cherny explains his approach to prompting Claude, advising users to give clear goals, specify effort level and verification steps rather than relying on heavy scaffolding.

Original post · 1 min read
I am surprised that people are surprised this is how I prompt Claude.

Talk to Claude the way you would a coworker. There's no secret to prompting. There's no need to be overly scaffolded or prescriptive for most tasks -- give Claude a goal, and it will figure it out.

Back in the Sonnet 3.5 days, your prompt mattered a lot. Nowadays, it's much more important to communicate to the model:

1. What you want it to do
2. How much effort you want it to spend
3. How it should verify that it did the right thing
Boris Cherny @bcherny
Prompt
♥ 12.1K · ⟲ 720 · 👁 1.1MView on X ↗

Eric Raymond Highlights Open-Source Rust Clone of Photoshop Built via LLM

Eric S. Raymond shares the photocraft GitHub project, a clean-room open-source reimplementation of Photoshop that he says was likely generated by decompiling the app, converting it to a spec and prompting an LLM for Rust. He argues this threatens closed-source software.

Original post · 1 min read
This is the doom I predicted a few days ago, coming for Photoshop. A clean-room open-source reimplementation.

No prizes for guessing that they decompiled Photoshop to source code, processed that to some kind of non-code specification language, then fed the spec to an LLM with an instruction to generate Rust.

Adobe just got nuked. And closed source is dead, dead, dead.

github.com/storytold/photocraft
♥ 16.2K · ⟲ 1.4K · 👁 3.7MView on X ↗

Nat Eliason Details Fourteen Ways His Bot Setup Automates Work

Nat Eliason lists fourteen functions of his bot setup, including a chief-of-staff agent that drafts emails, specialist agents per work lane, and cloud coding agents that open pull requests from Linear issues. He notes GrokBot as a substantial improvement over his previous OpenClaw setup.

Original post · 2 min read
Things my @bot setup does that still blow my mind:

1. A Chief of Staff who opens the day pulling open loops from email & tasks and suggesting things it can knock out before 7am.

2. After every meeting, decisions get folded into Notion, Linear, and Todoist — not left rotting in Granola

3. Every email starts as a draft. The CoS bot scans my email every ~2hr and drafts replies to nearly everything — including checking my cal for availability and finding requested attachments / links

4. A specialist for each lane: curriculum, engineering, coaching, hiring, content, ops, and one for every single piece of software

5. Routines that keep running while I’m offline (e.g. monitoring Sentry errors in our apps and proactively fixing things)

6. Group rooms where 2–4 bots share one project thread instead of me copy-pasting context

7. Cloud coding agents that pick up Linear issues and open PRs after running the list of open work by me EoD — then squash-merge to main when it’s done

8. Meeting prep briefs pulled from Granola + Notion before I walk in

9. A growing shareable knowledge base in Notion + a GitHub repo that we update daily based on what happens at school

10. Student progress look-up across Expertise, Followers, and CoFounder without inventing numbers — chat anytime to see where a student is on their business work

11. Mentor Mind that coaches me on how to hold the bar without inventing doctrine

12. Todoist as a central task list where it logs things it’s blocked on for me, or from meetings / emails — and I can paste links into chat to direct it how to solve them

13. Engineering work is automatically tracked in Linear so my and the product teams’ bots don’t collide with each other

14. Presentations spun up in Gamma / Claude Design without me opening a slide tool

15. Plaud / live capture → notes the bots can actually act on

Probably more but these were the immediate ones we thought of.
Nat Eliason @nateliason
GrokBot feels like absolute magic at this point, a meaningful leg up on my previous OpenClaw etc. setups.

And with how easy it is to setup, there's really no excuse now.
♥ 2.1K · ⟲ 207 · 👁 522.0KView on X ↗