WHITE PAPER · 2026.10.01
Japanese Performance Ranking of Judgment-Specialized Models (Jev and Others)
the original and 10 open models on the same 15 questions
Public AI leaderboards are built from English questions. We gave TypeSafe AI's Jev and ten free open models built in its image the same 15 Japanese questions, and report the ranking, the two kinds of failure, self-hosting versus paying, and a checklist for testing in-house.
Contents
- 1The short version: the English leaderboard fell apart in Japanese
- 2The same 15 questions, asked the same way, once
- 3One open model in ten matched the original
- 4The English leaderboard did not predict Japanese
- 5Two kinds of failure, two kinds of pain
- 6The AI that never stops anything lets risky posts through quietly
- 7Most false alarms came from a single misreading
- 8Add more questions, and the gap at the top disappeared
- 9Self-host, or pay the original?
- 10Cheaper than trusting a leaderboard — a 12-step check you can run yourself
- 11How far to trust this ranking
Get the PDF
Fill in the form and we will email you a download link.
Optional
AI developers and providers are not eligible. To check requests, the entered details are sent to TypeSafe AI (USA). See our privacy policy and terms.