Standards

How a word gets in, and how you know it is right

What we collect, who may verify it, what our numbers do and do not mean, and how to reuse the data.

The one rule

Only contribute in a language you actually speak. If you are unsure of a word, leave it for someone who is not. An honest gap is more useful to us than a confident guess, because a guess is very hard to find again later.

How an entry is made

Four stages. An entry can sit at any of them.

  1. 01

    Contributed

    Someone who speaks the language supplies a word. Either by filling a gap we asked for, or by adding one we did not know to ask about. Nothing is machine-generated in an indigenous language, ever.

  2. 02

    Bridged

    Every entry carries at least one bridge translation, English or Kiswahili. This is enforced by the database, not by a form. It is what makes the corpus a connected graph rather than 37 separate word lists.

  3. 03

    Attested

    Reviewers with rights in that specific language vouch for it, or dispute it. Two affirmations, at least one from a verified first-language speaker, with no outstanding dispute, meets the bar for publication.

  4. 04

    Voiced

    Speakers record it. We keep many recordings per word rather than one, because a speech model learns from variation across age, region and device. One perfect recording teaches it nothing.

What the statuses mean

Our public figures count only the first of these. An entry nobody has checked is not a verified entry, however good it looks.

  • Verified

    Publicly visible and counted. Reviewed, bridged, and not flagged for spelling.

  • Pending

    A person submitted it and it is waiting for review. Not public.

  • Seeded

    Imported in bulk from published sources, with a placeholder definition. Not public, not counted, waiting for a human. Roughly 1,200 entries are in this state and we would rather say so than pretend otherwise.

  • Needs orthography review

    An old import damaged the spelling. Held back until a speaker of that language confirms the correct form, because guessing at orthography is the one thing this project must not do.

4,712
Indigenous entries
36
Languages
1,199
Awaiting curation
358
Awaiting spelling

Who may verify what

Authority is granted per language, not globally. A linguist working on Dholuo has no standing over Kipsigis, and a Kipsigis first-language speaker with no degree has the highest standing there is on Kipsigis.

Native speaker

Grew up speaking the language. The strongest single voucher.

Heritage speaker

Family language, not spoken daily. Weighted below first-language.

Linguist

Academic or professional training, verified against an affiliation.

Student

Studying the language. Contributions welcome, weighted lowest.

Institution

A university, archive or language body.

A claim is not a credential until a moderator has checked it. Nobody can vouch for an entry they contributed themselves. That is enforced by the database, not by policy.

What belongs here

Yes

  • Everyday words, as people actually say them
  • Greetings, idioms, proverbs, set phrases
  • Regional and dialect variants, saying which
  • Farming, fishing, craft and ceremonial vocabulary
  • Words falling out of use, marked as such

No

  • Words you are guessing at
  • Invented or joke words
  • Slurs and hate speech
  • Promotional content
  • Knowledge a community has asked not to publish

Some knowledge is not any individual's to publish. Where an entry carries ceremonial or otherwise restricted meaning, a community can ask for it to be limited or removed, and that request is honoured regardless of the licence.

Recordings and consent

Your voice is personal data. We record only with explicit consent, stored against the exact wording you agreed to, and you can withdraw at any time from your profile. That removes your recordings from the corpus and deletes the audio.

Consent is split into separate permissions rather than one blanket agreement. You choose publishing on the site, training speech models, and redistribution under the corpus licence separately. Being credited by name is optional and separate. You must be 18 or over.

We ask where you learned the language, your age band and your voice. Not out of curiosity. A speech model trained mostly on one kind of voice from one region works badly for everyone else.

Using the data

The corpus is licensed Creative Commons Attribution 4.0 International (CC BY 4.0). Use it for anything, including commercially, as long as you credit the contributors and their communities.

We chose attribution over ShareAlike deliberately. Kenyan languages are missing from most commercial language technology, and the point of this project is for them to appear in it. A ShareAlike clause would oblige every such system to adopt our licence, which in practice means most would carry on without these languages.

Cite it as

LughaKonnect. 2026. An open corpus of Kenyan languages. https://lughakonnect.co.ke. Licensed CC BY 4.0.

Every entry records where it came from, so a claim in this corpus can be traced. Material under a ShareAlike licence is never merged into the corpus, precisely so the licence above stays true.

Ready?

The fastest way to help is to fill a gap. We will show you meanings your language does not have yet, one at a time.