Open Source

Projects

Libraries, datasets, and models built to advance Persian AI, all freely available under permissive licenses.

Python Library

Shekar

A high-performance Persian NLP library providing tokenization, normalization, POS tagging, NER, embeddings, spell checking, sentiment analysis, and dependency parsing, all in one package.

Normalization Tokenization POS & NER Embeddings
Speech Dataset

Neyshekar

A large-scale open Persian speech dataset collected via community crowdsourcing. Version 6.0 holds 62,279 recordings totalling 99.02 hours of native Persian speech with predefined speaker-disjoint splits, for ASR, TTS, and representation learning.

ASR TTS 99+ hours CC0 1.0