Neyshekar

A Large-Scale Open Persian Speech Dataset

Neyshekar is an open, community-driven Persian speech dataset collected via a web-based crowdsourcing platform at ney.shekar.io. Recordings are provided by volunteer contributors and paid voice actors, all native Persian speakers. Each release is a stable snapshot enabling reproducible research and consistent benchmarking. Released entirely under CC0 1.0 Universal. Free for any use, including commercial.

Dataset Statistics: v6.0 (Latest)

99.02h
Total audio duration
62,279
Total recordings
5.72s
Average clip duration
701,621
Total tokens
29,535
Vocabulary size
15,222
Informal samples (24.44%)

Speaker gender distribution: 33,735 female recordings and 28,076 male recordings.

Informal-register samples are identified using the Shekar rule-based InformalClassifier, so colloquial and formal speech can be filtered or balanced during training.

Predefined Splits

v6.0 is the first release to ship predefined splits. Splits are assigned per speaker, so a single recorder never straddles two sets and the evaluation splits stay speaker-disjoint from train. This makes results directly comparable across papers without each team inventing its own partition.

Train
58,244
93.5% · 91.99 h
Validation
1,886
3.0% · 3.14 h
Test
2,149
3.5% · 3.88 h

Named Entity Distribution

Entities were identified automatically using the Shekar NER model, providing rich linguistic metadata for downstream tasks. The counts below are measured on the v4.1 snapshot (12,443 entities in total).

LOC
4,817
Locations
DAT
3,616
Dates
PER
2,156
Persons
ORG
1,588
Organizations
EVE
266
Events

Why Neyshekar?

Open & CC0 Licensed

Released under CC0 1.0 Universal, the most permissive open license. Use it for any purpose, including commercial products and proprietary models, without attribution requirements.

Native Speakers Only

All recordings are from native Persian speakers: a mix of volunteer contributors and professional voice actors, ensuring natural prosody and authentic pronunciation.

Speaker-Disjoint Splits

Since v6.0, every release ships predefined train, validation, and test splits assigned per speaker, so no recorder appears in more than one set and benchmark numbers reflect true generalisation.

Reproducible Releases

Each version is a stable, versioned snapshot on Zenodo with a DOI, enabling reproducible experiments and consistent benchmarks across publications.

Rich Metadata

Every release includes transcriptions, speaker gender, duration statistics, vocabulary counts, register labels, and NER-tagged entity annotations from the Shekar NLP pipeline.

Community-Driven Growth

The dataset grows continuously through community contributions at ney.shekar.io. Each new release captures the latest snapshot with incremental improvements.

Balanced Gender Coverage

v6.0 includes 33,735 female and 28,076 male recordings, a balanced split essential for building gender-robust speech models.

Formal & Informal Speech

About 24% of v6.0 (15,222 samples) is colloquial Persian, labelled by the Shekar InformalClassifier, so models can be trained or evaluated on everyday spoken register rather than formal reading style alone.

Use Cases

Neyshekar is designed for a broad range of Persian speech research and engineering tasks:

Automatic Speech Recognition (ASR) Text-to-Speech (TTS) Speech Representation Learning Speaker Identification Voice Activity Detection Low-Resource Language Modeling Acoustic Model Training End-to-End Speech Models Pronunciation Modeling Multilingual Speech Research

Version History

Neyshekar is released incrementally. Each version is archived on Zenodo with a permanent DOI.

Note on v4. A number of clips in the v4 release (May 14, 2026) have misaligned audio–transcript pairs. Training or evaluating on that snapshot may introduce label noise and lead to unreliable results. If you have already downloaded v4, discard it and re-download v4.1 or later.

v6.0 September 1, 2026
62,279 recordings 99.02 h audio 701,621 tokens 29,535 vocab 15,222 informal predefined splits
v5.0 July 16, 2026
50,026 recordings 79.22 h audio 565,797 tokens 27,250 vocab 17,491 informal
v4.1 June 15, 2026
40,008 recordings 63.03 h audio 456,268 tokens 26,758 vocab 12,443 entities
v3 March 23, 2026
30,019 recordings 45.71 h audio 331,714 tokens 23,972 vocab
v2 January 15, 2026
20,020 recordings 29.08 h audio 208,472 tokens 20,853 vocab
v1 December 29, 2025
10,044 recordings 14.42 h audio 103,757 tokens 15,224 vocab

License & Terms of Use

CC0 1.0 Universal

The Neyshekar dataset is released under the CC0 1.0 Universal (Public Domain Dedication) license. It may be used, modified, and redistributed for any purpose, including commercial use, without restriction or attribution.

Note: Any attempt to identify or uncover the identity of individual speakers in the dataset is strictly prohibited.

Frequently Asked Questions

What is Neyshekar?

Neyshekar is a large-scale open Persian speech dataset built through community crowdsourcing. It contains 62,000+ recordings totalling 99+ hours of native Persian speech, along with transcriptions, gender labels, formal/informal register labels, NER-tagged entity metadata, and predefined speaker-disjoint splits.

How many hours of speech does Neyshekar contain?

Version 6.0 (the latest, released September 2026) contains 99.02 hours of audio across 62,279 recordings averaging 5.72 seconds each. The dataset has grown from 14.42 hours in v1 to 99+ hours in v6.0.

Does Neyshekar include informal Persian speech?

Yes. v6.0 contains 15,222 informal-register samples (24.44% of the dataset), labelled using the Shekar rule-based InformalClassifier. This supports training and evaluating models on colloquial spoken Persian rather than formal reading style alone.

Which version should I use?

Use v6.0, the latest release, which also ships the predefined splits. Avoid v4: it contains misaligned audio–transcript pairs and was superseded by v4.1. If you downloaded v4, discard it and re-download.

Where can I download the Neyshekar dataset?

Download it from Zenodo: doi.org/10.5281/zenodo.18073632. The collection platform is at ney.shekar.io.

Can I use Neyshekar for commercial products?

Yes. The CC0 1.0 Universal license places the dataset in the public domain. You can use, modify, and redistribute it for any purpose, including building commercial ASR or TTS products, without restriction.

Who recorded the dataset?

Recordings were contributed by volunteer community members at ney.shekar.io and by paid professional voice actors. All speakers are native Persian speakers.

What tasks is Neyshekar designed for?

Neyshekar targets ASR, TTS, speech representation learning, speaker identification, voice activity detection, and other downstream Persian speech applications.

Does Neyshekar provide train/validation/test splits?

Yes, since v6.0. The release ships predefined splits — train (58,244 samples, 91.99 h), validation (1,886 samples, 3.14 h), and test (2,149 samples, 3.88 h). Splits are assigned per speaker, so a recorder never straddles two sets and the evaluation splits are speaker-disjoint from train.

How do I cite Neyshekar?

Cite the Zenodo record: Amirivojdan, A. (2026). Neyshekar: A Large-Scale Open Persian Speech Dataset (v6.0). Zenodo. DOI: 10.5281/zenodo.18073632.

Citation

If you use Neyshekar in your research, please cite:

@dataset{amirivojdan_neyshekar_2026,
  author    = {Amirivojdan, Ahmad},
  title     = {{Neyshekar: A Large-Scale Open Persian Speech Dataset}},
  year      = {2026},
  version   = {v6.0},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.18073632},
  url       = {https://doi.org/10.5281/zenodo.18073632}
}

Links