We build evaluation for the domains where standard metrics quietly lie to you: dialect, culture, law, medicine, agent behavior. Places where a model can sound fluent and still be wrong, and where that gap is exactly what breaks in production. Our newest one is public. It grades Arabic cultural and sociolinguistic knowledge. The finding: Implicit cultural reasoning, not vocabulary, is where models fail. And so do the automated graders. If you're using an LLM to judge quality, it's most confident in exactly the spots it's most wrong. How we build it: Prompts and rubrics written by subject-matter experts. Penalty-weighted scoring, so it's clear what has to be present and what counts as an error. Human ground truth first, automated judges second. Reporting that names who's accurate, who's lenient, and where each model breaks. We covered 103 expert-authored prompt and rubric pairs across Egyptian and Iraqi Arabic, grading frontier models both as answerers and as judges against native-speaker ground truth. If you're shipping into a high-stakes or multilingual market and need evaluation that holds up to expert scrutiny, we'll run it for your model, your language, your domain. Explore the results: https://lnkd.in/gP_PtKjk Paper: https://lnkd.in/gqmHfVCW Benchmark your model: https://perle.ai/benchmark #LLMEval #AIEvaluation #ArabicNLP #ModelBenchmarking
This is such an important distinction. A model can sound completely fluent and still get the underlying cultural context wrong. That’s exactly where evaluation needs to go deeper.
A strong benchmark reminder: models can sound fluent yet still fail where cultural context matters most making expert-grounded evaluation essential for trustworthy AI.