Manthan Gupta Applies Karpathy's Autoresearch Idea to LLM Inference
Manthan Gupta describes building Auto-Inference-Optimiser, an open repo where an AI coding agent hill-climbs on MLX inference speed on Apple Silicon under a locked evaluation harness. The article focuses on the failed experiments and fake wins behind benchmark gains.
Original post · 10 min read
You see the final graph. You see the +12% or the "runs 2x faster now" claim. What you usually do not see is the graveyard of bad ideas behind it. The settings that looked promising but were just noise. The optimizations that made throughput better by quietly making the model worse. The fake wins that only happened because the benchmark got easier.
So I built a small repo called Auto-Inference-Optimiser (star the repository!) to study exactly that.
The idea is simple: lock the evaluation, open one file for experimentation, and let an AI coding agent hill-climb on inference speed forever on Apple Silicon.
The most interesting part was not that it achieved a speedup. It was the kind of speedup it achieved, what it failed to improve, and what that indicates about inference engineering on real hardware.
Let's get into it.
Why I Built This
I care a lot about inference right now.
Not in the abstract "LLMs are cool" sense. I mean the actual production questions: where latency comes from, what batching buys you, what prompt processing costs, how KV cache decisions change throughput, and where the hardware wall starts pushing back.
There is a lot of content online about training. There is also a lot of content online about agents. But there is still not enough content that combines both instincts: build a tight experimental harness, let the agent search inside it, and use that process to learn something real about inference.
This repo was my way of doing that.
It is inspired by Karpathy's Autoresearch, but pointed at a different layer of the stack. Instead of searching over training code on a GPU box, this one searches over an MLX inference pipeline on a Mac because I am GPU poor (please sponsor a GPU).
What The Repo Actually Does
At a high level, the repo turns "make inference faster" into a bounded optimization problem.
The structure is intentionally small:
That boundary is doing most of the work.
prepare.py is read-only. It fixes the benchmark model, the prompts, the warmup behavior, the averaging logic, and the quality gates. The agent cannot "win" by quietly changing the test.
inference.py is the search surface. That is where the agent is allowed to touch sampling, prefill step size, prompt formatting, and the general generation path.
program.md tells the agent how to behave:
That is the core harness.
And I like this design a lot because it bakes in three things that most autonomous coding demos hand-wave away:
Reversibility - bad ideas are cheap to discard.
Observability - every run leaves behind metrics and logs.
Constraints - the agent is not allowed to optimize by moving the goalposts.
The Evaluation Is The Real Product
The truth is that the most important file in this repo is not inference.py. It is prepare.py.
That file fixes the benchmark around a small Apple Silicon friendly model, runs warmups, averages across multiple runs, and evaluates five different prompt types:
explanation
long-context summarization
reasoning
creative generation
code generation
That already makes the benchmark better than a lot of speed demos, because decode-heavy and prefill-heavy cases behave differently.
But the more important choice is the quality gate.
This repo does not let the agent optimize only for tokens/sec. It requires two checks to pass:
avg_perplexity has to stay below a threshold
sanity_check has to stay above a threshold
That second gate matters a lot.
Perplexity is useful, but it is still a model-internal metric. It can tell you that outputs are becoming unstable or degenerate, but it does not fully tell you whether the answer is still usable. So the repo also checks for concrete task-level correctness: did the train speed answer contain 48? Did the transformer explanation mention the right ideas? Did the LCS prompt actually return something that looks like Python code?
This is one of my favorite design choices in the whole project.
Because if you do not defend quality explicitly, an optimization harness will absolutely "improve" your system by making it worse.
What Actually Worked
After the optimization runs, the pattern was surprisingly clear.
Here is the short version:
But the more interesting part is where they came from.
1. Argmax sampling was the biggest win
On the Qwen run, setting sampling to greedy decoding gave the largest gain: about +10.8% generation throughput.
On the Llama run, it was also the best keep: about +2.6%.
That tells you something important: sampling overhead is not free. Top-p decoding is doing real work every token, and if your objective is pure throughput, removing that work can matter more than a lot of fancier ideas.
Of course, there is a trade-off.
You get deterministic output and lose diversity. So this is not a universal recommendation for every product. But as an inference lesson, it is very clean: sometimes the fastest path is just doin… continue on X ↗