About the Bengali LKC

The Bengali Local Knowledge Core (Ben-LKC) is a role-based, web-accessible platform for crowdsourced lexical data entry, multi-tier validation, dispute resolution, and export of structured data fully compatible with the Universal Knowledge Core (UKC) Lexical Markup Framework (LMF) XML schema.

The UKC, developed at the University of Trento, is a large-scale multilingual lexico-semantic knowledge base extending the Princeton WordNet model with language-independent concepts spanning 335 languages. Bengali — despite its ~230 million native speakers — currently lexicalises only a small fraction of the UKC concept space. This platform exists to close that gap, one validated word at a time.

The current lexicon snapshot

19,270

Lexical entries

22,904

Word senses

12,480

Synsets

2,471

Polysemous words

Seeded from livelanguage-ben v1.0, published by KnowDive - University of Trento. Sources include IndoWordNet, Wiktionary, CogNet, MorphyNet and Princeton WordNet.

Seven roles, one pipeline

Platform Architect

Full system governance: users, tokens, configuration, anti-bot limits, export control and audit logs.

System Moderator

Operational queue governance: clears anti-bot holds, breaks arbitration ties, applies panel corrections, direct approval.

Lexical Contributor

Enters new lexical entries via Quick Input, manages drafts, and holds formal dispute rights to the arbitration panel.

Lexical Editor

Primary validation: approves, flags for correction, or rejects contributor submissions with mandatory reasoning.

Chief Lexical Editor

Final linguistic validation. Only this role (or a binding panel verdict) promotes an entry to VALIDATED.

Arbitration Curator

Votes on escalated disputes. Verdicts require a 2/3 quorum and a majority; ties go to the System Moderator.

Visitor (unauthenticated)

Word search, gloss search and dictionary browse over validated entries — with no internal identifiers exposed.

Concept-anchored by design

Every synset carries an Inter-Lingual Index (ILI) code that anchors it to the UKC's language-independent Concept Core. The platform never modifies or creates ILI codes — it only references existing ones, so every validated Bengali word becomes instantly interoperable with the vocabularies of hundreds of other languages when merged back into the UKC.

অবদান রাখতে চান?

Want to contribute? Access tokens are issued by the Platform Architect — or try the demo tokens on the login page.

Go to Login