Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Edition of Monday, June 22, 2026

4 stories

Baseten Details Engineering Behind Fastest GLM-5.2 API

How we built the world’s fastest API for GLM-5.2

Baseten describes how it built an API serving GLM-5.2 above 280 tokens per second, using custom inference, NVFP4 quantization, KV-aware routing, disaggregated inference and multi-token prediction. The post positions the open MIT-licensed model as comparable to frontier models at 70-80% lower cost.

Original post · 9 min read
X ArticleHow we built the world’s fastest API for GLM-5.2
GLM-5.2 is the biggest news in open models since DeepSeek-R1.

It’s easy to see why. GLM-5.2 delivers comparable performance to GPT 5.5 and Opus 4.8 at a fraction of the cost, generally 70-80% less expensive on a pure token basis (use our calculator to estimate savings for your workload).
But a model has to be more than just smart and inexpensive. To be useful in production, a model needs to be fast, reliable, and available at scale. Delivering on the promise of frontier open intelligence requires exceptional inference.
Accordingly, we built the world’s fastest API for GLM-5.2, currently serving over 280 tokens per second as measured by Artificial Analysis.

We achieved this performance by leveraging a number of techniques across the entire inference process by:
Updating our custom inference engine to implement shared DSA for the GLM-5.2 architecture.
Running and calibrating an in-house NVFP4 quantization from the original FP8 weights that demonstrates equivalent quality on agentic benchmarks like BFCL.
Ensuring high KV cache hit rates via KV-aware routing built with NVIDIA Dynamo tools for lower prefill burden and improved TTFT on requests with repeated prefixes.
Achieving a 2x higher TPS for observed workload shapes by running disaggregated inference built with the NVIDIA Dynamo toolkit.
Improving TPS further via speculation by implementing support for GLM-5.2 Multi-Token Prediction heads.
You can experience this performance for yourself with GLM-5.2 on Baseten Model APIs. We also have GLM-5.2 available as a dedicated deployment for high-volume workloads.


GLM-5.2 Overview
GLM-5.2 by Z.ai is a 744B parameter frontier LLM that excels at agentic tasks (especially coding) and supports up to a 1 million token context window. It uses a similar architecture to its predecessor, GLM-5.1: mixture of experts (40B active parameters), non-thinking and thinking modes, and a fully open MIT license. While GLM-5.2 shares a lot in common with GLM-5.1, it now uses shared DSA weights, which we implemented support for in our customized runtime engine.

GLM-5.2 has great benchmark scores, but by now AI builders know that there is more to a model’s utility than its performance on standard evals. In practice, GLM-5.2 meets or exceeds the capabilities suggested by its benchmarks. It's a genuinely great model for writing code, operating agents, and other frontier language model tasks.
High-quality NVFP4 quantization for Blackwell GPUs
We run our model APIs on NVIDIA Blackwell GPUs with a customized inference engine within the Baseten Inference Stack. The selected runtime uses NVFP4 weights for maximum performance. From the original FP8 weights, we performed an in-house quantization to NVFP4 using NVIDIA ModelOpt. NVFP4 is a 4-bit floating point data format by NVIDIA that uses dual scale factors to retain high dynamic range and preserve model quality.
In our calibration and testing of the quantized model, we focused on ensuring that GLM-5.2 performs faithfully on common patterns for agents. On the BFCL function calling benchmark, we observed roughly equivalent performance between the native FP8 weights and our NVFP4 quantization, with scores across runs within the margin of error for the benchmark.
NVFP4 quantization improves performance on both time to first token and tokens per second by unlocking faster tensor cores and reducing burden on VRAM bandwidth.

Cache-aware routing with NVIDIA Dynamo
GLM-5.2 is particularly well suited for long context requests and complex agentic tasks. These workloads generally have very long input sequences. By re-using KV cache between requests, we can skip expensive prefill for shared sequences.
We generally talk about KV cache re-use in the context of time to first token (TTFT). However, reasoning models like GLM-5.2 generally care more about time to first answer token (TTFAT), which combines TTFT with some TPS for the reasoning sequence.

This chart shows that of the 7.9 second average to generate the first answer token, 7.1 of those seconds were spent generating reasoning tokens versus only 0.8 seconds spent processing the input sequence.
Still, bringing the TTFT down to 800 ms is important for the overall responsiveness and throughput of the system. In large-scale production deployments, KV cache is split across various independent replicas. We use tools from NVIDIA Dynamo to route incoming requests.

Exact cache hit rates on a multi-tenant API depend on the exact traffic profile at any given time. Thus far, we’re observing high hit rates across fairly heterogeneous traffic, which reduces load on prefill and improves end-to-end performance.
Prefill-decode disaggregation with NVIDIA Dynamo
One of the highest-impact optimizations we made to our performance is disaggregating prefill and decode for GLM-5.2.
There are two distinct phases of LLM inference:
Prefill: The compute-bound process that processes the input sequence, builds the KV cache, and generates the first output token. Prefill… continue on X ↗
♥ 1.5K · ⟲ 141 · 👁 549.5KView on X ↗

Meta's $900M Cred Investment Signals WhatsApp Payments Push

Meta's $900M Cred Investment Signals WhatsApp Payments Push

Sugandha argues Meta's $900M investment in Indian fintech Cred, alongside CEO Kunal Bahl joining to lead WhatsApp, points to plans to expand WhatsApp into a full payments product. The analysis cites India's UPI volumes and microtransaction growth as the rationale.

Original post · 2 min read
It’s not that complicated. WhatsApp may seem like a global product but effectively it’s not. India is WhatsApp’s only “viable” market, with both the numbers and the consumer habits to make it a possible lifeboat for Meta’s otherwise flailing position as a tech company. Kunal himself has commented previously on India being the world’s DAU farm (which is true).

WhatsApp has already maxxed out its business product in India (no other country’s consumer or regulatory bodies would permit or tolerate the level of spam India deals with on WhatsApp) as well as its ads product (there are ads even between stories now, ffs, in a private messaging app).

The only lever that is yet to be maxxed out is its payments product which launched in India a few years back. Considering the growth of India’s digital microtransaction economy and corresponding consumer habits, it’s tempting to consider that WhatsApp has the chance to outdo every payment product in the region.

All of this narrows down the executive search quite a bit. The $900M investment is not only for Kunal, it’s for the intellectual property he brings about India’s fintech (a headache for global executives everywhere) and Indian consumer habits, from the homegrown CRED. It’s actually a small price considering UPI hit ~230B transactions last year, 33% increase y-o-y. I believe the microtransaction economy is projected to reach $600B in less than a decade. If that’s even fractionally true, it’s a small price for a strong hire. About half of that $900M is going to be new fuel for the company (which will certainly buy CRED good runway), the rest helps investors get an exit.

I surmise WhatsApp is planning to become a fullblown payments product. Messaging + microtransactions has anyway been the trend in Indian consumer products. Bad news for the local fintech startup economy. Worse news for the consumer, imho.

It may be time for someone to build the next messaging app for friends and families. It’s been a while.
Sheel Mohnot @pitdesi
Very interesting- single person acquihire sorta

Meta invests $900M in Indian Fintech Cred at $4.5B valuation (mix of primary and secondary)

CEO @kunalb11 steps down, joins Meta to lead WhatsApp (from India??? It is WhatsApp’s biggest market by a long shot) twitter.com/jbahrdestefano/status/206907856571…
♥ 859 · ⟲ 120 · 👁 149.9KView on X ↗
AI6/10

Bryan Johnson Praises Midjourney's Whole-Body Scanner

Bryan Johnson calls the Midjourney scanner revolutionary after using it at the unveiling, arguing its fast, low-cost whole-body imaging could create routine health baselines. He says he plans to add weekly scans to his data and AI-driven health analysis.

Original post · 6 min read
The Midjourney scanner is revolutionary. There’s a bullish case that exceeds the most optimistic takes.

I was at the unveiling and used the scanner myself. I personally want to experiment with a weekly whole body Midjourney scan to add to my 1.5 billion data points and let my AI and doctors start connecting the dots.

Most of the early commentary has focused on the wrong questions: “is it as good as MRI?” and “what about false positives?” These are legitimate concerns, but they miss the bigger shift.

The more important question is: what does fast, low cost, safe whole body imaging unlock?

Let’s start with measurement.

A speedometer tells you how fast you are going. A fuel/battery gauge tells you when to stop. A thermostat tells you what to wear. The stock price tells you how much money you’ve made or lost. We measure what we care about.

Except, oddly, for our bodies, which are among the least measured things in our lives. Most people have more data on their favorite sports team, bank account, and social media performance than their body. The future will think we were crazy for this.

The first law of medicine is to do no harm. Our current system has harm baked into it.

+ an undiagnosed condition progressing silently is harm
+ a doctor who can’t easily get a patient screened preventively is harm
+ having no baseline to compare against when something shows up is harm

Our preventive net is narrow and inconsistent. Late stage diagnoses that could have been caught earlier remain common. Midjourney’s technology won’t eliminate that overnight, but it points toward a future where routine wholebody baselines become normal rather than exceptional.

Midjourney can help flip harm-by-default into a new expectation for our health infrastructure: almost no one will ever again be blindsided by a late-stage, life-threatening diagnosis that could have been caught earlier reasonably and cost-effectively.

Some examples of what earlier structural visibility enables:

+ breast cancer caught while localized has a ~99% five year survival rate. Once it has spread distantly, that drops to around 32%.

+ an abdominal aortic aneurysm kills more than 8 in 10 people when it ruptures. A single ultrasound finds the aorta in 99 percent of people, and screening cuts aneurysm deaths by a third to a half.

Midjourney’s technology will not do it all on its own. Its full angle, water immersion approach works around bone rather than seeing through it, and routes bowel gas to image the full abdominal cross section. Yet two real limits remain: air filled lungs stay a blind spot even here, and the brain is out of reach behind the skull, beyond the torso and legs this scanner covers.

That is fine, and they may improve these areas over time. Midjourney doesn’t need to do it all in order for it to be one of the biggest things to hit medicine in a long time.

Let’s look at where specifically Midjourney may be useful to each of us. We’ll start with where we get data today:

1) Blood draws tell us what is happening chemically.
2) Wearables tell us how the body is functioning.
3) Imaging tells us what is happening structurally.

The third layer, soft tissue, is the one we have never been able to access easily. MRI is great, but it is expensive, intimidating, and slow.

Midjourney's technology excels with soft tissue. Here are three places it could be game changing. There are many more.

1. Metabolic health - fatty liver is one of the earliest structural signs of metabolic dysfunction. It’s strongly linked to insulin resistance, type 2 diabetes, and cardiovascular risk. Being able to track visceral fat, muscle fat infiltration, and liver fat over time could give a much clearer picture than blood markers alone. Over 88% of Americans are metabolically unhealthy.

2. Endocrine tissue - the same metabolic patterns often cluster with thyroid issues, PCOS, and hypogonadism. Ultrasound can directly image the thyroid and ovarian structures. Fat tissue itself is an endocrine organ, so tracking it structurally adds another useful data layer.

3. Soft tissue + multiomics - new proteomic aging clocks can already predict risk for many chronic diseases from blood proteins. These molecular models could become significantly more powerful when combined with actual structural imaging data. The two are complementary, not competitive.

The real advantage: baseline + longitudinal tracking

The biggest unlock isn’t a single scan. It’s having a baseline followed by regular follow-ups. A one off scan in a moment of concern turns every finding into a potential crisis. Without context, you have no idea whether something is new, stable, or changing. With baseline + repeated measurement, the question changes from “what is this?” to “is this changing?” Most incidental findings stay stable. The dangerous ones tend to grow or evolve. Trajectory is often more informative than any single image or timepoint.This is why false positives become more manageable with frequent, low-friction imaging.

Midjourney has a difficult road ahead. Building robust, clinically validated medical hardware and software is extremely hard. Regulatory, technical, and adoption challenges shouldn’t be understated. Also, David is doing this for the right reasons and he’s well positioned financially to push through the difficulty.

On the horizon

We are moving quickly into a future where we will have continuous biological measurement. It will be all around us, a lot of it invisible and autonomous. Measurement will be in our gyms, beds, homes, clothing, offices, cars, glasses, and wearables. It will also be inside of us, in tissue and circulating in our blood vessels. This moves us from managing crises to preventing them. But this future will not just show up. We need bold builders like David and his team, willing to do the hard work.
Midjourney @midjourney
A technical dive inside our new "Midjourney Scanner"
♥ 4.9K · ⟲ 386 · 👁 616.9KView on X ↗

Dhilip Subramanian Switches From Wispr Flow to Open-Source FluidVoice

Dhilip Subramanian reports dictating heavily with paid tool Wispr Flow, then moving to FluidVoice, an open-source local voice tool for Mac that needs no API key. He says he cancelled his paid plan and recommends it to Mac users.

Original post · 1 min read
I've dictated almost everything for 6 months with Wispr Flow. 44,414 words, 161 wpm, top 0.1% of users.

Last week I tried FluidVoice. Open source, runs local on my Mac, corrects as I speak with no API key, and handles slang better than I expected.

Cancelled my paid plan. If you're on a Mac, this one's for you: altic.dev/fluid

@ALTIC_DEV
♥ 6.1K · ⟲ 253 · 👁 1.8MView on X ↗