Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Baseten Details Engineering Behind Fastest GLM-5.2 API

How we built the world’s fastest API for GLM-5.2

Baseten describes how it built an API serving GLM-5.2 above 280 tokens per second, using custom inference, NVFP4 quantization, KV-aware routing, disaggregated inference and multi-token prediction. The post positions the open MIT-licensed model as comparable to frontier models at 70-80% lower cost.

Original post · 9 min read
X ArticleHow we built the world’s fastest API for GLM-5.2
GLM-5.2 is the biggest news in open models since DeepSeek-R1.

It’s easy to see why. GLM-5.2 delivers comparable performance to GPT 5.5 and Opus 4.8 at a fraction of the cost, generally 70-80% less expensive on a pure token basis (use our calculator to estimate savings for your workload).
But a model has to be more than just smart and inexpensive. To be useful in production, a model needs to be fast, reliable, and available at scale. Delivering on the promise of frontier open intelligence requires exceptional inference.
Accordingly, we built the world’s fastest API for GLM-5.2, currently serving over 280 tokens per second as measured by Artificial Analysis.

We achieved this performance by leveraging a number of techniques across the entire inference process by:
Updating our custom inference engine to implement shared DSA for the GLM-5.2 architecture.
Running and calibrating an in-house NVFP4 quantization from the original FP8 weights that demonstrates equivalent quality on agentic benchmarks like BFCL.
Ensuring high KV cache hit rates via KV-aware routing built with NVIDIA Dynamo tools for lower prefill burden and improved TTFT on requests with repeated prefixes.
Achieving a 2x higher TPS for observed workload shapes by running disaggregated inference built with the NVIDIA Dynamo toolkit.
Improving TPS further via speculation by implementing support for GLM-5.2 Multi-Token Prediction heads.
You can experience this performance for yourself with GLM-5.2 on Baseten Model APIs. We also have GLM-5.2 available as a dedicated deployment for high-volume workloads.


GLM-5.2 Overview
GLM-5.2 by Z.ai is a 744B parameter frontier LLM that excels at agentic tasks (especially coding) and supports up to a 1 million token context window. It uses a similar architecture to its predecessor, GLM-5.1: mixture of experts (40B active parameters), non-thinking and thinking modes, and a fully open MIT license. While GLM-5.2 shares a lot in common with GLM-5.1, it now uses shared DSA weights, which we implemented support for in our customized runtime engine.

GLM-5.2 has great benchmark scores, but by now AI builders know that there is more to a model’s utility than its performance on standard evals. In practice, GLM-5.2 meets or exceeds the capabilities suggested by its benchmarks. It's a genuinely great model for writing code, operating agents, and other frontier language model tasks.
High-quality NVFP4 quantization for Blackwell GPUs
We run our model APIs on NVIDIA Blackwell GPUs with a customized inference engine within the Baseten Inference Stack. The selected runtime uses NVFP4 weights for maximum performance. From the original FP8 weights, we performed an in-house quantization to NVFP4 using NVIDIA ModelOpt. NVFP4 is a 4-bit floating point data format by NVIDIA that uses dual scale factors to retain high dynamic range and preserve model quality.
In our calibration and testing of the quantized model, we focused on ensuring that GLM-5.2 performs faithfully on common patterns for agents. On the BFCL function calling benchmark, we observed roughly equivalent performance between the native FP8 weights and our NVFP4 quantization, with scores across runs within the margin of error for the benchmark.
NVFP4 quantization improves performance on both time to first token and tokens per second by unlocking faster tensor cores and reducing burden on VRAM bandwidth.

Cache-aware routing with NVIDIA Dynamo
GLM-5.2 is particularly well suited for long context requests and complex agentic tasks. These workloads generally have very long input sequences. By re-using KV cache between requests, we can skip expensive prefill for shared sequences.
We generally talk about KV cache re-use in the context of time to first token (TTFT). However, reasoning models like GLM-5.2 generally care more about time to first answer token (TTFAT), which combines TTFT with some TPS for the reasoning sequence.

This chart shows that of the 7.9 second average to generate the first answer token, 7.1 of those seconds were spent generating reasoning tokens versus only 0.8 seconds spent processing the input sequence.
Still, bringing the TTFT down to 800 ms is important for the overall responsiveness and throughput of the system. In large-scale production deployments, KV cache is split across various independent replicas. We use tools from NVIDIA Dynamo to route incoming requests.

Exact cache hit rates on a multi-tenant API depend on the exact traffic profile at any given time. Thus far, we’re observing high hit rates across fairly heterogeneous traffic, which reduces load on prefill and improves end-to-end performance.
Prefill-decode disaggregation with NVIDIA Dynamo
One of the highest-impact optimizations we made to our performance is disaggregating prefill and decode for GLM-5.2.
There are two distinct phases of LLM inference:
Prefill: The compute-bound process that processes the input sequence, builds the KV cache, and generates the first output token. Prefill… continue on X ↗
♥ 1.5K · ⟲ 141 · 👁 549.5KView on X ↗

More in Agents & Dev Tools

Developer Rebuilds Seven Adobe Apps in Rust Using Opus 5.5

Peter Yang highlights a developer who reimplemented seven Adobe apps, including Photoshop, Premiere and Lightroom, in Rust with Claude Opus 5.5 and open-sourced them. The developer believes they can match Adobe's features within months, against Adobe's $840 yearly all-apps plan.

Original post · 1 min read
It's insane to watch AI blow apart closed source software and games.

4 examples from the past month:

1. 7 of Adobe's biggest apps, including Photoshop, Premiere, and Lightroom, have been partially rebuilt in Rust with Opus 5.5 and open sourced. It's still early, but the developer thinks they can match Adobe's features within months. Adobe's all-apps plan costs $840/year.
Miguel Ángel Durán @midudev
Todos los productos de Adobe reimplementados desde cero, gratuitos y de código abierto

→ getartcraft.com/apps
♥ 50 · ⟲ 2 · 👁 10.7KView on X ↗

Vercel's Guillermo Rauch Explains Turborepo's Migration From Go to Rust

Guillermo Rauch says Vercel moved Turborepo from Go to Rust, a migration that was controversial internally due to human costs. He argues that with AI agents the calculus has changed, so what is best for humans is no longer necessarily best for business.

Original post · 1 min read
DHH is fundamentally right about Rust. For context, Vercel has been undergoing a Rust-ification (carcinization, technically 🦀) for a while.

One of the first projects we migrated was Turborepo, from Go to Rust¹. The migration completed, but the RoI was actually quite controversial internally.

While Rust was in our eyes better for low-level OS access, something crucial for a build system like Turbo, the human migration costs were very sustantive.

Go is very fast. It's beautifully designed. It's easy to iterate on. We were very conflicted about the migration, because it was *humans* writing the code, *even if we knew Rust was a better choice*.

The calculus has now changed. What's "best for humans" is no longer necessarily "best for business".

FWIW, it's also quite unlikely that Rust is the end-all-be-all toolchain. I'm quite certain there's greener pasture ahead, because Rust itself was designed before the 'supersonic tsunami' of agents hit.

¹ https​://vercel.com/blog/how-turborepo-is-porting-from-go-to-rust
♥ 3.2K · ⟲ 152 · 👁 352.7KView on X ↗

Integer Multiplication Algorithm Bound Tightened Repeatedly With Astra

A post reports that a user running ChatGPT Astra in a loop is repeatedly breaking records for integer multiplication algorithms. It quotes an update to OpenAI problem #109 that tightens the constant from 2^-182 to 2^-59, a roughly 500,000-fold improvement over the previous result.

Original post · 1 min read
This guy has 6.1 Astra running in a loop and is breaking the record for integer multiplication algorithms every few hours lmaooooo.
Doug Colkitt @0xdoug
We are publishing an update to OpenAI problem #109 Integer multiplication) with another substantial further tightening:

κ = 2⁻⁵⁹ (from OpenAI’s original κ = 2⁻¹⁸²)

Approximately 500 thousand fold improvement over our previous result and a 2¹²³ fold improvement over the original OAI result.

The latest redesigned the finite network to share intermediate computations and scratch space, then tightened the recursion and Gaussian estimates.
♥ 4.1K · ⟲ 118 · 👁 167.4KView on X ↗

Boris Cherny Says Prompting Claude Should Feel Like Talking to a Coworker

Boris Cherny explains his approach to prompting Claude, advising users to give clear goals, specify effort level and verification steps rather than relying on heavy scaffolding.

Original post · 1 min read
I am surprised that people are surprised this is how I prompt Claude.

Talk to Claude the way you would a coworker. There's no secret to prompting. There's no need to be overly scaffolded or prescriptive for most tasks -- give Claude a goal, and it will figure it out.

Back in the Sonnet 3.5 days, your prompt mattered a lot. Nowadays, it's much more important to communicate to the model:

1. What you want it to do
2. How much effort you want it to spend
3. How it should verify that it did the right thing
Boris Cherny @bcherny
Prompt
♥ 12.1K · ⟲ 720 · 👁 1.1MView on X ↗

Eric Raymond Highlights Open-Source Rust Clone of Photoshop Built via LLM

Eric S. Raymond shares the photocraft GitHub project, a clean-room open-source reimplementation of Photoshop that he says was likely generated by decompiling the app, converting it to a spec and prompting an LLM for Rust. He argues this threatens closed-source software.

Original post · 1 min read
This is the doom I predicted a few days ago, coming for Photoshop. A clean-room open-source reimplementation.

No prizes for guessing that they decompiled Photoshop to source code, processed that to some kind of non-code specification language, then fed the spec to an LLM with an instruction to generate Rust.

Adobe just got nuked. And closed source is dead, dead, dead.

github.com/storytold/photocraft
♥ 16.2K · ⟲ 1.4K · 👁 3.7MView on X ↗

Nat Eliason Details Fourteen Ways His Bot Setup Automates Work

Nat Eliason lists fourteen functions of his bot setup, including a chief-of-staff agent that drafts emails, specialist agents per work lane, and cloud coding agents that open pull requests from Linear issues. He notes GrokBot as a substantial improvement over his previous OpenClaw setup.

Original post · 2 min read
Things my @bot setup does that still blow my mind:

1. A Chief of Staff who opens the day pulling open loops from email & tasks and suggesting things it can knock out before 7am.

2. After every meeting, decisions get folded into Notion, Linear, and Todoist — not left rotting in Granola

3. Every email starts as a draft. The CoS bot scans my email every ~2hr and drafts replies to nearly everything — including checking my cal for availability and finding requested attachments / links

4. A specialist for each lane: curriculum, engineering, coaching, hiring, content, ops, and one for every single piece of software

5. Routines that keep running while I’m offline (e.g. monitoring Sentry errors in our apps and proactively fixing things)

6. Group rooms where 2–4 bots share one project thread instead of me copy-pasting context

7. Cloud coding agents that pick up Linear issues and open PRs after running the list of open work by me EoD — then squash-merge to main when it’s done

8. Meeting prep briefs pulled from Granola + Notion before I walk in

9. A growing shareable knowledge base in Notion + a GitHub repo that we update daily based on what happens at school

10. Student progress look-up across Expertise, Followers, and CoFounder without inventing numbers — chat anytime to see where a student is on their business work

11. Mentor Mind that coaches me on how to hold the bar without inventing doctrine

12. Todoist as a central task list where it logs things it’s blocked on for me, or from meetings / emails — and I can paste links into chat to direct it how to solve them

13. Engineering work is automatically tracked in Linear so my and the product teams’ bots don’t collide with each other

14. Presentations spun up in Gamma / Claude Design without me opening a slide tool

15. Plaud / live capture → notes the bots can actually act on

Probably more but these were the immediate ones we thought of.
Nat Eliason @nateliason
GrokBot feels like absolute magic at this point, a meaningful leg up on my previous OpenClaw etc. setups.

And with how easy it is to setup, there's really no excuse now.
♥ 2.1K · ⟲ 207 · 👁 522.0KView on X ↗