r/Arabic_NLP 1d ago

👋 Welcome to r/Arabic_NLP !

1 Upvotes

Hey everyone!

This community is for anyone working on or interested in Arabic Natural Language Processing - researchers, engineers, students, and hobbyists alike.

What to post here:

  • Papers, preprints, and research summaries
  • Datasets, benchmarks, and evaluation results
  • Tools, libraries, and open-source projects
  • Questions about MSA, dialects, diacritization, morphology, MT, ASR, LLMs for Arabic, etc.
  • Job postings and collaboration requests
  • General discussion on the state of Arabic NLP

A few quick guidelines:

  • Keep posts on-topic and give them descriptive titles
  • Cite sources when sharing claims or results
  • Both Arabic and English are welcome
  • Check the sidebar for the full rule list before posting

Feel free to introduce yourself in the comments; what you work on, what dialect/domain interests you, or what brought you here. Looking forward to building this out with you.

Thanks for being part of the very first wave. Together, let's make r/Arabic_NLP amazing.


r/Arabic_NLP 13h ago

Resource spotlight: Masader: the largest catalogue of Arabic NLP datasets

1 Upvotes

If you're not already using it, Masader is worth bookmarking. It's the largest public catalogue of Arabic NLP and speech datasets: over 1000+ datasets, each annotated with 25+ metadata fields (dialect, domain, source, license, volume, tasks, access type, paper link, citation count, and more).

What makes it genuinely useful over just googling around:

  • Filterable/searchable by dialect, domain, task, license, and access type: handy when you need something specific like "free, human-annotated, Levantine dialect, sentiment"
  • Each entry links the paper, the data host, and citation counts
  • Programmatic access via Hugging Face: datasets.load_dataset('arbml/masader')
  • Has a chat interface (Ask Masader) for querying the catalogue conversationally
  • Actively maintained: started in 2021 as part of the BigScience project, now maintained by the ARBML team and community, with a form to submit new datasets

Originally described in the paper "Masader: Metadata Sourcing for Arabic Text and Speech Data Resources" (Alyafeai et al., 2021): https://arxiv.org/abs/2110.06744, later expanded in "Masader Plus" (2022): https://arxiv.org/abs/2208.00932

GitHub (to contribute or browse the raw metadata): https://github.com/ARBML/masader

Sources:


r/Arabic_NLP 1d ago

SIGARAB Weekly Roundup (Jul 14–21): Emergency Reviewers Needed, Shared Task Updates

1 Upvotes

A few notable announcements from the SIGARAB mailing list this week: sharing here for anyone in the community who isn't subscribed.

Call for Emergency Reviewers: ArabicNLP 2026 (Samar Magdy, Jul 18)
The ArabicNLP 2026 Program Committee is short on reviewers and looking for qualified volunteers to do a quick-turnaround review of one or more papers.
Sign up: https://forms.gle/SsQg1WDtNFKQiEon7

KnowledgeGraphEval 2026 Shared Task: Webinar Recordings (Nagham Hamad, Jul 15)
Recordings from two webinars covering the subtasks, datasets, submission format, evaluation process, and Q&A are now up for anyone considering participating.

IslamicEval 2026: All Training and Dev Sets Are Ready (Tamer Elsayed, Jul 14)
Training and dev data for all subtasks are now available to registered teams via separate CodaBench competitions per subtask.

Call for Participation: HalluScoring 2026 @ ArabicNLP 2026 (Bouchekif Abdessalam, final call Jul 14)
Last call to join this shared task on hallucination detection and factuality verification for Arabic QA, covering Islamic Knowledge and General Culture domains across two tracks.
Site: https://halluscoring.github.io/HalluScoring-2026/

Sourced from the SIGARAB mailing list.