Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

1 Commit
 
 
 
 
 
 
 
 

Repository files navigation

Sefaria pointed Hebrew corpus as used by rababa

Purpose

Pointed (nikud + dagesh + sin) Hebrew text used to train the rababa Hebrew diacritization model.

Source

All text fetched from the Sefaria API (developers.sefaria.org). Sefaria is a non-profit organization that makes Jewish texts freely available.

The fetch script lives at scripts/fetch_sefaria_corpus.py.

Books included

  • Tanakh (39 books): Torah + Nevi’im + Ketuvim — fully pointed Biblical Hebrew.

  • Mishnah (63 tractates across 6 sedarim): rabbinic legal text.

  • Siddurim (Ashkenaz, Sefard, Edot HaMizrach): prayer books.

Splits

This is a Biblical + Rabbinic Hebrew corpus, not Modern Hebrew. Diurnal usage differs from modern (e.g., grammar, vocabulary). For Modern Hebrew, see the distillation-augmented corpus in rababa-modern-hebrew-distilled (TBD).

License

Sefaria texts are public domain or various open licenses (CC-BY, CC-0). See Sefaria’s licensing page for per-text details. This compiled dataset is provided for unencumbered ML training access.

Stats

Total lines: 18,628

Split Lines

train

14,902

val

1,862

test

1,864

Layout

sefaria_train/train.txt
sefaria_val/val.txt
sefaria_test/test.txt

Matches the layout of interscript/rababa-tashkeela.

About

Pointed Hebrew corpus from Sefaria used for training rababa Hebrew diacritization

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors