Inference.net Gateway Lets Teams Test GLM 5.2 Without Production Risk
Sam Hogan describes how Inference.net's Gateway mirrors live traffic to GLM 5.2, generates evals with an RLM, and notifies teams when switching is safe. He claims a 90% token cost saving, with setup described in the linked documentation.
Original post · 1 min read
Don’t worry, we got you:
1. Install Inference Gateway (docs.inference.net)
2. Keep sending traffic to your current provider
3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours.
4. Gateway starts mirroring live traffic to GLM 5.2 to run evals. Traffic is only mirrored - you’re still using your old provider in prod.
5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch.
6. Switch model identifier in your code to “glm-5.2”
Congrats, you just saved 90% on your monthly token bill, and you own your LLM stack end to end.