Wednesday, October 7, 2026ArchiveSearchAsk the paper

The Computomatix Times

All the posts fit to save — curated from @computomatix's bookmarks & likes on X

Edition of Monday, June 29, 2026

6 stories

AI8/10

Explainer Breaks Down Prefill Versus Decode in LLM Inference

Explainer Breaks Down Prefill Versus Decode in LLM Inference

Avi Chawla explains that LLM inference has two phases: compute-bound prefill, which drives time-to-first-token, and memory-bound decode, which drives inter-token latency. He notes that adding compute rarely speeds decoding, and that faster memory or smaller caches are the real fixes.

Original post · 3 min read
Prefill & decode in LLM inference.

Have you ever noticed that the first token from an LLM always takes a moment to appear? But the subsequent tokens stream out smoothly?

That pause isn't a network lag, but rather it's a structural property of how LLMs fundamentally work.

Inference happens in two phases that share the same model and the same code path, but the workload looks completely different in each, with different bottlenecks.

> Prefill stage starts when you submit a prompt.

The model processes every input token in one parallel pass, computing Q, K, and V for all of them at once.

Attention runs as a matrix multiplication, and the GPU chips run at high utilization, doing fast math.

Prefill is compute-bound, and the metric that captures it is time-to-first-token (TTFT).

> Decode stage starts once the first token is out.

To generate the next one, the model only computes Q, K, and V for that single new token, because everything before it is already cached.

So the model loops one token per forward pass, multiplying a single query against the cached keys instead of a full matrix. This makes the inference fast due to the tiny computation.

But the GPU still has to load every weight and every cached entry from memory to do that tiny computation, so the bottleneck flips and compute sits idle while memory bandwidth becomes the limiting factor.

Decode is memory-bound, and the metric that captures it is inter-token latency (ITL).

GPU utilization peaks during prefill and drops sharply during decode because memory, not compute, is the bottleneck in the second phase.

Throwing more compute at a slow-streaming model often does nothing because the fix for memory-bound workloads is faster memory or a smaller cache, not more FLOPs.

Long contexts feel disproportionately slow because the KV cache grows with every token, and every decode step has to read all of it.

But maintaining the cache is an important optimization since it makes decoding viable.

- Without KV cache, every new token would force a recomputation of attention over the entire growing sequence.
- With KV cache, the cache is built once during prefill, then grows by exactly one entry per decode step, with existing entries reused rather than recomputed.

The cache lives in GPU memory and grows linearly with sequence length, so a 13B model roughly requires 1 MB per token, which means a 4K context consumes 4 GB of VRAM on the cache alone.

The entire field is now optimizing around this constraint with quantized caches, sliding windows, grouped-query attention, and PagedAttention, while DeepSeek's V4 series goes further and redesigns attention itself so the cache stays small from the start.

The practical takeaway is that when someone says their model feels slow, the first question is whether it's slow to start or slow to stream.

Slow to start means prefill and a compute bottleneck, while slow to stream means decode and a memory bottleneck.

The article below is a first-principles guide to LLM inference that walks through everything between your prompt and the streamed response, covering tokenization, embeddings, attention, the prefill and decode split, KV caching, and quantization.

It will give you a complete mental model of how inference actually works under the hood.

Read it below.
Avi Chawla @_avichawla
How LLM Inference Works, Clearly Explained. — Every generate() call to an LLM runs two distinct computational phases on the same GPU:
prefill (processing the prompt) is compute-bound
while decode (generating tokens one at a time) is memory-bound.
♥ 648 · ⟲ 109 · 👁 55.6KView on X ↗

Cardiologist Says AI Is Shifting Power From Doctors to Patients

Cardiologist Afshine Emrani argues AI tools are moving medical diagnosis into patients' hands. He cites OpenAI's o3 diagnosing rare pediatric diseases in an NEJM-published study, a WashU blood-marker biological age calculator, and AI-enhanced CT angiography detecting inflamed arteries.

Original post · 4 min read
I'm a cardiologist. I've spent twenty years as the person patients trust to interpret their bodies. And I need to tell you something that most physicians won't say out loud:

AI is about to change the power dynamic between you and your doctor. Forever.

Four days ago, OpenAI's o3 model diagnosed 18 children with rare diseases that the best human specialists at Boston Children's Hospital couldn't solve — some after nearly twenty years of searching. Published in the New England Journal of Medicine.

Two weeks ago, WashU researchers proved that nine routine blood markers can calculate your biological age — and predict cancer risk years before any tumor forms. A free calculator. Available to anyone.

Last month, AI-enhanced coronary CT angiography detected inflamed arteries in patients whose standard stress tests said "normal." Patients who would have gone home reassured and wrong.
The pattern is unmistakable. The tools that used to require a specialist, a referral, a three-month wait, and a $400 copay are migrating into your phone, your bloodwork portal, and your own hands.

And I'm watching something in my practice I never expected.
Patients are walking in more informed than some of the residents I trained. They've run their PhenoAge score. They know their ApoB. They've read the study about Lp(a) before I've had time to bring it up. They come with questions so specific that the conversation starts at a level it took me years of training to reach.
This used to threaten physicians. It shouldn't. It should liberate us.
Because here's the truth about the old model: a 15-minute appointment where your doctor runs a basic metabolic panel, glances at the numbers, says "looks fine," and sends you home — that model was never good enough. It was just all we had. It missed 75% of future heart attacks. It caught cancer late. It told women with microvascular disease they had anxiety. It filed children with rare diseases as "unsolvable."

AI doesn't replace the physician. I've said this before and I mean it — the human moment, the clinical judgment, the hand on the shoulder when the diagnosis lands — that's irreplaceable.

But AI does something the old model never could: it gives you the ability to see inside your own biology with a depth and speed that was impossible a decade ago. To track your own numbers. To calculate your own biological age. To bring data to your doctor that elevates the conversation from "am I sick?" to "where exactly am I heading, and what do we do about it?"

The patient who walks in with their ApoB, their Lp(a), their hsCRP, their PhenoAge calculation, and a list of questions from the latest research — that patient doesn't threaten me.

That patient is the easiest person in my practice to keep alive.
Because they've already done the one thing most patients never do: they stopped waiting for permission to understand their own body.

I went into medicine because I wanted to help people live longer. What I've learned is that the patients who live longest are the ones who took ownership — not of my job, but of their own data, their own questions, and their own decisions.

The tools are here. The research is published. The calculators are free. The blood tests cost less than a dinner out.

You don't need to wait for your annual physical to find out what's happening inside you. You don't need permission to understand your own biology. And you don't need to accept "looks fine" from anyone — including me — when the science offers a deeper answer.

The revolution isn't coming. It's in your pocket. In your patient portal. In the published studies you can read yourself.

The only question left is whether you'll use it — or keep waiting for someone to tell you it's time.
Your body. Your data. Your life.

Take ownership. Your future self is counting on it.
♥ 2.9K · ⟲ 631 · 👁 774.9KView on X ↗

Inference.net Gateway Lets Teams Test GLM 5.2 Without Production Risk

Catalyst by Inference.net - Inference.net Documentation

Sam Hogan describes how Inference.net's Gateway mirrors live traffic to GLM 5.2, generates evals with an RLM, and notifies teams when switching is safe. He claims a 90% token cost saving, with setup described in the linked documentation.

Original post · 1 min read
Want to try GLM 5.2 in production but worried how it might change your product?

Don’t worry, we got you:

1. Install Inference Gateway (docs.inference.net)
2. Keep sending traffic to your current provider
3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours.
4. Gateway starts mirroring live traffic to GLM 5.2 to run evals. Traffic is only mirrored - you’re still using your old provider in prod.
5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch.
6. Switch model identifier in your code to “glm-5.2”

Congrats, you just saved 90% on your monthly token bill, and you own your LLM stack end to end.
docs.inference.netCatalyst by Inference.net - Inference.net DocumentationFetch the complete documentation index at: /llms.txt Use this file to discover all available pages before exploring further. Catalyst is a platform for understa
♥ 1.5K · ⟲ 68 · 👁 537.7KView on X ↗

Developer Reports Strong Results From New Codex Development Workflow

Developer Reports Strong Results From New Codex Development Workflow

Paul Solt says his new Codex workflow exceeded expectations, producing eight features ready for release in his app after some early trial and error. He credits Dimillian, emanueledpt and steipete for inspiration.

Original post · 1 min read
My NEW Codex workflow is better than I expected.

8 new features ready for release in my app.

Took a few attempts to figure out the workflow and some bugs. Feels like the future.

Thanks @Dimillian @emanueledpt and @steipete for the inspiration.
♥ 566 · ⟲ 25 · 👁 176.0KView on X ↗

Trader Advises Buying QQQ Between 10:15 and 11:30 After Gap Fill

Trader Advises Buying QQQ Between 10:15 and 11:30 After Gap Fill

Prof reports that QQQ fell from 718 to 705 before closing at highs near 724, and advises buying in the 10:15 to 11:30 window. The post is a short, self-congratulatory trading tip with no independent analysis.

Original post · 1 min read
$QQQ went from 718 to 705 before closing the day at highs @ 724.

My first post: Avoid buying
My second post: This is where you buy
This post: Happy that I helped.

My third advice: If you're looking at buying, look between 10:15 to 11:30. That window will work more times than not.
Prof @TheProfInvestor
Second advice: This is where you buy.

The gap up: traced all the way down, filled.

Stocks you liked 30 mins ago are 5% lower now twitter.com/TheProfInvestor/status/20715891000…
♥ 301 · ⟲ 13 · 👁 61.5KView on X ↗

Trader Says Chasing Gap-Ups Hurts as Stocks Fall 5 Percent

Trader Says Chasing Gap-Ups Hurts as Stocks Fall 5 Percent

Prof argues that after a gap up traced down and filled, stocks liked 30 minutes earlier are now 5 percent lower, so buyers should wait for support. The post is brief trading commentary with two chart photos.

Original post · 1 min read
Second advice: This is where you buy.

The gap up: traced all the way down, filled.

Stocks you liked 30 mins ago are 5% lower now
Prof @TheProfInvestor
Solid advice that will save you a lot of money:

There is zero reason to be chasing a gap up when indices are below a declining 21EMA.

Either you buy when indices hit support (like they did last week) or you wait for a structure to form.

( Reclaim 21EMA + put a higher low )

Chasing gaps in a downtrend hurts more than it rewards.
♥ 361 · ⟲ 15 · 👁 169.9KView on X ↗