Jev Model Forecasts Booking Outcomes From 2,029 Real AI Receptionist Calls
Muratcan Koylan reports a zero-shot experiment in which the Jev model analyzed structural features of 2,029 phone calls without audio or transcripts. It reached an AUC of 0.78 at the halfway point and ranked calls correctly 94% of the time near the end.
Original post · 1 min read
No transcripts or audio; it never heard a word. Our AI receptionist's calls were reduced to pure structure, meaning turns, tool calls, workflow stages and timing.
During the calls, Jev made 38,012 turn-level forecasts at 118 ms median latency, reviewed every call with five typed questions and produced 10,145 answers in 26 seconds with 256 requests in flight.
The experiment was zero-shot, with no fine-tuning or examples from our data. We compared Jev's forecasts with what actually happened in the EHR.
By the halfway point, Jev could meaningfully separate calls that would book from those that wouldn't (AUC 0.78), and near the end it ranked them correctly 94% of the time.
Even though Jev over-focused on visible errors our agent usually overcomes, it's still pretty incredible that it analyzed thousands of real calls in seconds for only $3.



