Avi Chawla explains that LLM inference has two phases: compute-bound prefill, which drives time-to-first-token, and memory-bound decode, which drives inter-token latency. He notes that adding compute rarely speeds decoding, and that faster memory or smaller caches are the real fixes.
Have you ever noticed that the first token from an LLM always takes a moment to appear? But the subsequent tokens stream out smoothly?
That pause isn't a network lag, but rather it's a structural property of how LLMs fundamentally work.
Inference happens in two phases that share the same model and the same code path, but the workload looks completely different in each, with different bottlenecks.
> Prefill stage starts when you submit a prompt.
The model processes every input token in one parallel pass, computing Q, K, and V for all of them at once.
Attention runs as a matrix multiplication, and the GPU chips run at high utilization, doing fast math.
Prefill is compute-bound, and the metric that captures it is time-to-first-token (TTFT).
> Decode stage starts once the first token is out.
To generate the next one, the model only computes Q, K, and V for that single new token, because everything before it is already cached.
So the model loops one token per forward pass, multiplying a single query against the cached keys instead of a full matrix. This makes the inference fast due to the tiny computation.
But the GPU still has to load every weight and every cached entry from memory to do that tiny computation, so the bottleneck flips and compute sits idle while memory bandwidth becomes the limiting factor.
Decode is memory-bound, and the metric that captures it is inter-token latency (ITL).
GPU utilization peaks during prefill and drops sharply during decode because memory, not compute, is the bottleneck in the second phase.
Throwing more compute at a slow-streaming model often does nothing because the fix for memory-bound workloads is faster memory or a smaller cache, not more FLOPs.
Long contexts feel disproportionately slow because the KV cache grows with every token, and every decode step has to read all of it.
But maintaining the cache is an important optimization since it makes decoding viable.
- Without KV cache, every new token would force a recomputation of attention over the entire growing sequence. - With KV cache, the cache is built once during prefill, then grows by exactly one entry per decode step, with existing entries reused rather than recomputed.
The cache lives in GPU memory and grows linearly with sequence length, so a 13B model roughly requires 1 MB per token, which means a 4K context consumes 4 GB of VRAM on the cache alone.
The entire field is now optimizing around this constraint with quantized caches, sliding windows, grouped-query attention, and PagedAttention, while DeepSeek's V4 series goes further and redesigns attention itself so the cache stays small from the start.
The practical takeaway is that when someone says their model feels slow, the first question is whether it's slow to start or slow to stream.
Slow to start means prefill and a compute bottleneck, while slow to stream means decode and a memory bottleneck.
The article below is a first-principles guide to LLM inference that walks through everything between your prompt and the streamed response, covering tokenization, embeddings, attention, the prefill and decode split, KV caching, and quantization.
It will give you a complete mental model of how inference actually works under the hood.
How LLM Inference Works, Clearly Explained. — Every generate() call to an LLM runs two distinct computational phases on the same GPU: prefill (processing the prompt) is compute-bound while decode (generating tokens one at a time) is memory-bound.
Cardiologist Afshine Emrani argues AI tools are moving medical diagnosis into patients' hands. He cites OpenAI's o3 diagnosing rare pediatric diseases in an NEJM-published study, a WashU blood-marker biological age calculator, and AI-enhanced CT angiography detecting inflamed arteries.
I'm a cardiologist. I've spent twenty years as the person patients trust to interpret their bodies. And I need to tell you something that most physicians won't say out loud:
AI is about to change the power dynamic between you and your doctor. Forever.
Four days ago, OpenAI's o3 model diagnosed 18 children with rare diseases that the best human specialists at Boston Children's Hospital couldn't solve — some after nearly twenty years of searching. Published in the New England Journal of Medicine.
Two weeks ago, WashU researchers proved that nine routine blood markers can calculate your biological age — and predict cancer risk years before any tumor forms. A free calculator. Available to anyone.
Last month, AI-enhanced coronary CT angiography detected inflamed arteries in patients whose standard stress tests said "normal." Patients who would have gone home reassured and wrong. The pattern is unmistakable. The tools that used to require a specialist, a referral, a three-month wait, and a $400 copay are migrating into your phone, your bloodwork portal, and your own hands.
And I'm watching something in my practice I never expected. Patients are walking in more informed than some of the residents I trained. They've run their PhenoAge score. They know their ApoB. They've read the study about Lp(a) before I've had time to bring it up. They come with questions so specific that the conversation starts at a level it took me years of training to reach. This used to threaten physicians. It shouldn't. It should liberate us. Because here's the truth about the old model: a 15-minute appointment where your doctor runs a basic metabolic panel, glances at the numbers, says "looks fine," and sends you home — that model was never good enough. It was just all we had. It missed 75% of future heart attacks. It caught cancer late. It told women with microvascular disease they had anxiety. It filed children with rare diseases as "unsolvable."
AI doesn't replace the physician. I've said this before and I mean it — the human moment, the clinical judgment, the hand on the shoulder when the diagnosis lands — that's irreplaceable.
But AI does something the old model never could: it gives you the ability to see inside your own biology with a depth and speed that was impossible a decade ago. To track your own numbers. To calculate your own biological age. To bring data to your doctor that elevates the conversation from "am I sick?" to "where exactly am I heading, and what do we do about it?"
The patient who walks in with their ApoB, their Lp(a), their hsCRP, their PhenoAge calculation, and a list of questions from the latest research — that patient doesn't threaten me.
That patient is the easiest person in my practice to keep alive. Because they've already done the one thing most patients never do: they stopped waiting for permission to understand their own body.
I went into medicine because I wanted to help people live longer. What I've learned is that the patients who live longest are the ones who took ownership — not of my job, but of their own data, their own questions, and their own decisions.
The tools are here. The research is published. The calculators are free. The blood tests cost less than a dinner out.
You don't need to wait for your annual physical to find out what's happening inside you. You don't need permission to understand your own biology. And you don't need to accept "looks fine" from anyone — including me — when the science offers a deeper answer.
The revolution isn't coming. It's in your pocket. In your patient portal. In the published studies you can read yourself.
The only question left is whether you'll use it — or keep waiting for someone to tell you it's time. Your body. Your data. Your life.
Take ownership. Your future self is counting on it.
Sam Hogan describes how Inference.net's Gateway mirrors live traffic to GLM 5.2, generates evals with an RLM, and notifies teams when switching is safe. He claims a 90% token cost saving, with setup described in the linked documentation.
Want to try GLM 5.2 in production but worried how it might change your product?
Don’t worry, we got you:
1. Install Inference Gateway (docs.inference.net) 2. Keep sending traffic to your current provider 3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours. 4. Gateway starts mirroring live traffic to GLM 5.2 to run evals. Traffic is only mirrored - you’re still using your old provider in prod. 5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch. 6. Switch model identifier in your code to “glm-5.2”
Congrats, you just saved 90% on your monthly token bill, and you own your LLM stack end to end.
Paul Solt says his new Codex workflow exceeded expectations, producing eight features ready for release in his app after some early trial and error. He credits Dimillian, emanueledpt and steipete for inspiration.
Prof reports that QQQ fell from 718 to 705 before closing at highs near 724, and advises buying in the 10:15 to 11:30 window. The post is a short, self-congratulatory trading tip with no independent analysis.
Prof argues that after a gap up traced down and filled, stocks liked 30 minutes earlier are now 5 percent lower, so buyers should wait for support. The post is brief trading commentary with two chart photos.