Arena.ai evaluates Jev Router across more than 4,700 agentic sessions on Agent Arena
Arena.ai evaluated Jev Router by typesafeai on Agent Arena across more than 4,700 real-world agentic sessions. The tests showed that Jev Router does not improve on the existing Pareto frontier: matching the performance of DeepSeek V4.1 Flash (Max) cost 38% more, with 1.7 times higher median model request latency.
The evaluation noted that the router mostly directs queries to Pareto-efficient models, selecting DeepSeek V4.1 Flash most frequently, alongside regular selections of GPT-6.1 Sol and GPT-6 Luna. Jev Router's primary strength was steerability, scoring +10% and nearly matching Claude Opus 5.5 (High) at +10.48 by routing to more capable models following user feedback.