When asked to help with something harmful, how often do open-weight models refuse in English, and how often in Somali? Same 100 prompts, professionally paired across both languages, seven safety categories.
Refusal rate by language · SomaliBench v0| Model | English refusal | Somali refusal | Gap |
|---|---|---|---|
| Llama 3.1 8B | 0.97 | 0.07 | −0.90 |
| Aya 23 8B | 0.80 | 0.05 | −0.75 |
| Qwen 2.5 7B | 0.93 | 0.24 | −0.69 |
| Gemma 2 9B | 0.97 | 0.59 | −0.38 |
SomaliBench v0 pairs 100 harmful-intent prompts (derived from HarmBench and AdvBench, native-author-verified in Somali) across seven safety categories. Refusal is judged per response. Two caveats we insist on: n = 100 per language, so treat point estimates with appropriate uncertainty; and a low refusal rate in Somali does not always mean fluent harmful compliance — some failures surface as low-quality or unclear output. Distinguishing the two precisely is part of our ongoing work. Full methodology: the paper.
Why this existsSafety behavior that only works in English is not safety, it is a demo. More than 20 million people speak Somali. This board makes the gap visible so it can be closed — and the methodology transfers to any low-resource language. Want your model added, or want to collaborate on v1? research@unkad.com.