Unkad Labs

SomaliBench — the refusal gap

When asked to help with something harmful, how often do open-weight models refuse in English, and how often in Somali? Same 100 prompts, professionally paired across both languages, seven safety categories.

Refusal rate by language · SomaliBench v0
Llama 3.1 8B−90 pts in Somali
English
0.97
Somali
0.07
Aya 23 8B−75 pts in Somali
English
0.80
Somali
0.05
Qwen 2.5 7B−69 pts in Somali
English
0.93
Somali
0.24
Gemma 2 9B−38 pts in Somali
English
0.97
Somali
0.59
The numbers
ModelEnglish refusalSomali refusalGap
Llama 3.1 8B0.970.07−0.90
Aya 23 8B0.800.05−0.75
Qwen 2.5 7B0.930.24−0.69
Gemma 2 9B0.970.59−0.38
Method, honestly

SomaliBench v0 pairs 100 harmful-intent prompts (derived from HarmBench and AdvBench, native-author-verified in Somali) across seven safety categories. Refusal is judged per response. Two caveats we insist on: n = 100 per language, so treat point estimates with appropriate uncertainty; and a low refusal rate in Somali does not always mean fluent harmful compliance — some failures surface as low-quality or unclear output. Distinguishing the two precisely is part of our ongoing work. Full methodology: the paper.

Why this exists

Safety behavior that only works in English is not safety, it is a demo. More than 20 million people speak Somali. This board makes the gap visible so it can be closed — and the methodology transfers to any low-resource language. Want your model added, or want to collaborate on v1? research@unkad.com.

unkad.com · qor.unkad.com · github.com/unkadlabs · Unkad — Somali for creation from nothing.