click on red text to expand and collapse information
General Information about arabiCorpus
- What is a corpus?
- A corpus is simply a large collection of words from a given language. These words can come from a variety of sources.
- A person searches from a corpus by entering a search string, which is simply a string of letters or characters.
- Corpora (plural of corpus) are often used by linguists to learn about some phenomenon of language.
- Anyone interested in a particular language, though, can use a corpus and learn from it.
- The benefit of a corpus is that it contains real language that was spoken or written by speakers of that language.
- This allows the researcher the ability to see how the language is used in actual context instead of just in theoretical examples.
- A student studying a foreign language can use a corpus to see how a particular word is most often used.
- A language teacher can identify more frequent words and teach those to students earlier on.
- A researcher can discover what spelling variations are used in different texts.
- These are just a few of the countless ways that someone can use a corpus.
- arabiCorpus is a program that allows you to search from a large corpus of Arabic words.
- You can use the Tutorial and Instructions below to see the different uses and benefits of this program.
- Important Things to Remember
- arabiCorpus allows you to search large, untagged Arabic corpora.
- 'Untagged' means that the words in the corpora have not been assigned to a particular a certain part of speech.
- There are ways to find the part of speech you are looking for but there may still be errors. See 'Select a Part of Speech Filter' in the Detailed Search Instructions section of Performing Searches in arabiCorpus below for more information about this.
- arabiCorpus searches for EXACTLY the string you type in (and nothing else).
- In connection with the comment above, the part of speech (POS) filters can help in finding what you are looking for, but they will not necessarily give you the exact results you want.
- The program simply searches for the string you type in with various possible attachments to that string. It does not actually know if the resulting word is a verb, noun, etc.
- arabiCorpus accepts most regular expression language in search strings.
- For more information on these types of searches, see the section on 'Searching with Regular Expressions' below.
- You can easily search for individual words and you have the ability to search for multiple words at once.
- You can search individual texts or combined corpora, which consist of multiple texts.
- Only some of the individual texts are available in Basic Search but all of them are available in Advanced Search.
- Avoid using the large combined corpora 'All' and 'All Newspapers' when searching for common words.
- Some filtering of the results is available after the search.
- There are various tools provided after a search has been performed to analyze the results and make them more meaningful.
- See the section 'Analyzing the Results of a Search' below for more information.
- Use the Tutorial below to help you with your initial searches and see the Instructions for additional details on using arabiCorpus.
- Basic Corpus Information
- arabiCorpus has corpora in five main categories or genres: Newspapers, Modern Literature, Nonfiction, Egyptian Colloquial, and Premodern.
- You will select a specific corpus each time you perform a search in arabiCorpus.
- The program gives you the ability to search combined corpora made up of multiple texts in a related genre.
- You can also search any text individually by using the Advanced Search mode.
- You can even search all of the texts at the same time. (Although this will often slow down your search).
- Avoid using the large corpora 'All' and 'All Newspapers' when searching for common words because it may cause problems for future searches.
- The names and word counts for the Basic Search corpora are found in this section. For a more detailed listing of the texts found in each corpus, see the 'Detailed Corpus Information' section.
- Note that you cannot add these amounts together for the total number, since the grouped corpora group and regroup in various configurations.
- The total number of words of the whole corpus is: 173,600,000.
- All Newspapers: 135,360,804
- Al-Masri Al-Yawm 2010: 13,880,826
- Ahram 1999: 15,892,001
- ShuruqColumns: 2,067,137
- AlGhad01: 19,234,228
- AlGhad02: 19,628,088
- Hayat 1997: 19,473,315
- Hayat 1996: 21,564,239
- Tajdid 2002: 2,919,782
- Watan 2002: 6,454,411
- Thawra: 16,153,918
- Modern Literature: 1,026,171
- Nonfiction: 27,945,460
- Islamic Discourse: 27,365,915
- Other Nonfiction: 579,545
- Egyptian Colloquial: 164,457
- Premodern: 9,127,331
- Adab Literature: 2,073,071
- Grammarians: 1,210,614
- Medieval Philosophy/Science: 1,576,860
- Hadith Literature: 3,624,346
- Quran: 84,532
- 1001Nights: 557,908
Detailed Corpus Information
- Newspapers
- The bulk of the corpora come from newspapers and each newspaper corpus can be searched individually in Basic Search. The following list shows the years and locations for the current newspapers:
- The 'All Newspapers' corpus has 135,360,804 words and consists of the following corpora:
- Al-Masri Al-Yawm 2010 from Egypt – 13,880,826 words (المصري اليوم من مصر)
- Al-Thawra from Syria – 16,153,918 words (الثورة من سوريا)
- At-Tajdid 2002 from Morocco – 2,919,782 words (التجديد من المغرب)
- Al-Watan 2002 from Kuwait – 6,454,411 words (الوطن من الكويت)
- Al-Ghad(01) 2010-11 from Jordan – 19,234,228 words (الغد من الأردن)
- Al-Ghad(02) 2010-11 from Jordan - 19,628,088 words (الغد من الأردن)
- Al-Ahram 1999 from Egypt – 15,892,001 words (الأهرام من مصر)
- Al-Hayat 1997 from London – 19,473,315 words (الحياة من لندن)
- Al-Hayat 1996 from London – 21,564,239 words (الحياة من لندن)
- Shuruq Columns Egypt – 2,067,137 words (الشروق من مصر)
- Remember that the 'All Newspapers' corpus is very large so you should avoid using it when searching for common words.
- Premodern
- There are also various Premodern texts, two of which can be searched individually in Basic Search and others are available in combined corpora based on genre. Any individual text may be searched using Advanced Search. This list details the various corpora in this genre:
- The total number of words in the 'Premodern' corpus is 9,127,331 words and all of the following texts mentioned in this section are part of it. There are also smaller combined corpora that are included in the Premodern corpus.
- The following corpora can be searched individually in Basic Search and are also found in the combined 'Premodern' corpus:
- Quran – 84,532 words (القرآن الكريم)
- 1001 Nights – 557,908 words (كتاب ألف ليلة وليلة)
- The 'Adab Literature' corpus, which has 2,073,071 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
- Book of Songs by Abu Al-Faraj Al-Isfahani - 1,523,874 words (كتاب الأغاني لأبو الفرج الأصفهاني)
- The Scrooges by Al-Jahiz - 48,241 words (البخلاء للجاحظ)
- Letters of Al-Jahiz - 167,001 words (رسائل الجاحظ)
- The Book of Animals by Al-Jahiz - 333,955 words (كتاب الحياون للجاحظ)
- The 'Grammarians' corpus, which has 1,210,614 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
- Ajarumiyya Grammar – 2,399 words (الكتاب للأجرومية)
- Ibnjinni Grammar - 192,748 words (الخصائص لأبو الفتح عثمان بن جني الموصلي)
- Jurjani Grammar - 77,653 words (أسرار البلاغة للجرجاني)
- Mubarrad Grammar - 168,789 words (المقتضب في اللغة للمبرد)
- Sayuti Grammar - 316,090 words (الكتاب: همع الهوامع فى شرح جمع الجوامع للإمام السيوطى)
- Thalabi Grammar - 101,757 words (سحر البلاغة وسر البراعة للثعالبي and فقه اللغة وسر العربية للثعالبي)
- Zamakhshari Grammar - 131,298 words (أساس البلاغة للزمخشري)
- Sibawaihi – 219,880 words (الكتاب لأبو بشر عمرو بن قنبر الملقب بسيبويه)
- The 'Medieval Philosophy and Science' corpus, which has 1,576,860 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
- Ihya by Al-Ghazali - 992,341 words (إحياء علوم الدين للغزالي)
- Incoherence of the Philosophers by Al-Ghazali – 47,307 words (تهافت الفلاسفة للغزالي)
- Medical Aphorisms by Maimonides – 81,581 words (كتاب الفصول في الطبّ لموسى بن ميمون)
- On Asthma by Maimonides – 17,904 words (مقالة في الربو لموسى بن ميمون)
- Middle Commentary by Averroes – 33,004 words (تلخيص كتاب النفس لأرسطو لابن رشد)
- Metaphysics by Avicenna – 90,760 words (الإلهيات من الشفاء لابن سينا)
- Issues in Fiqh by Ibn Hanbal - 72,589 words (مسائل أحمد ين حنبل رواية ابنه عبد الله)
- Prolegomenon by Ibn Khaldun - 241,374 words (مقدمة لابن خلدون)
- The 'Hadith Literature' corpus, which has 3,624,346 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
- Sunan by Al-Bukhari - 617,933 words (سنن للبخاري)
- Sunan by Abi Daud - 429,227 words (سنن لأبي داود)
- Sunan by Al-Tarmthi - 425,807 words (سنن للترمذي)
- Sunan by Al-Darmi - 210,909 words (سنن للدارمي)
- Sunan by Ibn Maja - 321,924 words (سنن ابن ماجة)
- Sunan by Muslim - 746,829 words (صحيح مسلم لمسلم بن الحجاج)
- Sunan by Al-Nasai - 871,717 words (السنن الكبرى للإمام النسائي)
- Modern Literature
- The next section of corpora comes from Modern Literature, primarily from novels. The countries represented with the novels are Egypt, Palestine, Algeria, Saudi Arabia, the Sudan, Syria, and Lebanon. Currently about half of the material in the Modern Literature corpus is from Egypt, about a fourth from Algeria, and lesser amounts from the rest of the countries. These are the novels found in the corpus as of Feb 2012:
- The 'Modern Literature' corpus has 1,026,171 words and consists of the following texts:
- ريم بسيوني: رائحة البحر with 38,485 words
- ريم بسيوني: مدبولي with 40,042 words
- إبراهيم عبد المجيد: لا أحد ينام في الاسكندرية with 98,580 words
- علاء الأسواني: عمارة يعقوبيان with 55,191 words
- علاء الأسواني: شيكاجو with 77,026 words
- خالد الخميسي: تاكسي with 28,391 words
- نجيب محفوظ: ميرامار with 35,579 words
- نجيب محفوظ: الكرنك with 13,742 words
- نجيب محفوظ: صدى النسيان with 6,276 words
- نجيب محفوظ: أصداء السيرة الذاتية with 11,542 words
- نجيب محفوظ: أولاد حارتنا with 104,684 words
- أحلام مستغانمي: ذاكرة الجسد with 71,065 words
- أحلام مستغانمي: عابر سرير with 53,925 words
- أحلام مستغانمي: فوضى الحواس with 57,418 words
- رجاء عبدالله الصانع: بنات الرياض with 56,173 words
- الطاهر وطار: الولي الطاهر يعود إلى مقامه الزكي with 19,271 words
- الطاهر وطار: الولي الطاهر يرفع يديه بالدعاء with 19,957 words
- التطاهر وطار: الحوات والقصر with 26,756 words
- الطيب صالح: عرس الزين with 16,283 words
- إدوار الخرات: ترابها زعفران with 20,475 words
- لطيفة الزيات: الشيخوخة وقصص أخرى with 19,556 words
- يحيى حقي: قصص ليحيى حقي with 9,409 words
- إلياس خوري: مملكة الغرباء with 19,790 words
- غسان كنفاني: أم سعد with 6,623 words
- غسان كنفاني: عائد إلى حيفا with 11,326 words
- غسان كنفاني: مسرحية الباب with 9,263 words
- نجاة حالو: سر الحياة with 11,211 words
- سعد الله ونوس: مغامرة رأس المملوك جابر with 17,659 words
- تميم صائب: لا تفقأ عينيك يا أوديب with 7,607 words
- غادة السمان: ختم َلذاكرة بالشمع الأحمر with 4,070 words
- علي سالم: أولادنا في لندن with 16,865 words
- أمجد ناصر: حيث لا تسقط الأمطار with 41,931 words
- Egyptian Colloquial
- There is a small combined corpus with some Egyptian Colloquial data.
- The 'Egyptian Colloquial' corpus has 164,457 words and consists of the following corpora:
- The Egyptian play Awaladna fi Landan – 16,865 words
- Material from the Egypt Chat website – 140,234 words
- An interview with Hosni Mubarak – 7,358 words
- It is important to note that these texts have some colloquial and a lot of Fusha and mixed Fusha and colloquial. The literature and the newspapers also contain some colloquial, as that is the nature of Arabic.
- Nonfiction
- The last group of corpora is in the Nonfiction category, which includes a large amount of material from an Islamic Discourse web-site along with some literary criticism, other scholarly (and not so scholarly) works, some political speeches, and some official UN and other diplomatic documents. These are the corpora found in this genre:
- The 'Nonfiction' corpus has 27,945,460 words and contains all of the following texts, which are also divided into smaller subsections:
- The 'Islamic Discourse' corpus, which has 27,365,915 words, is a subsection of the 'Nonfiction' corpus and contains the following corpus:
- Material from Sayd.net (an Islamic Discourse web-site) - 27,365,915 words (صيد الفوائد)
- The 'Other Nonfiction' corpus, which has 579,545 words, is a subsection of the 'Nonfiction' corpus and contains the following texts:
- Landed from the Sky by Anis Mansour – 40,322 words (الذين هبطوا من السماء لأنيس منصور)
- The Absent Truth by Farag Fouda – 33,234 words (الحقيقة الغائبة لفرج فودة)
- We March Forward edited by Dr. Hala Esbanyuli – 11,113 words (نسير إلى الأمام لدكتور هالة اسبانيولي)
- Myth of Delusion by Muhammad Khalil Al-Hakaymah – 46,959 words (أسطورة الوهم لمحمد خليل الحكايمه)
- Milestones by Sayyid Qutb – 38,842 words (معالم في الطريق لسيد قطب)
- Black Book of Capitalism translated by Dr. Anton Hamsy – 145,362 words (الكتاب الأسود للرأسمالية ترجم الدكتور أنطون حمصي)
- American Time from New York to Kabul by Muhammad Hasanayn Haykal – 73,778 words (الزمن الأمريكي من نيويورك إلى كابول لمحمد حسنين هيكل)
- Al-Naba’ Magazine – 5,204 words (مجلة النبأ)
- UN Resolution 1636 – 1,870 words (نص القرار 1636 للأمم المتحدة)
- Media Arabic in Egypt by Farouq Shousha – 3,990 words (اللغة العربية في الإذاعة والتلفاز والفضائيات في جمهورية مصر العربية للاستاذ فاروق شوشة)
- Sharm Al-Sheikh Resolution 1999 – 1,852 words (نص مذكرة شرم الشيخ من أيلول 1999)
- Sana Speech by Hosni Mubarak – 1,110 words (نص خطاب السيد رئيس الجمهورية في قمة تجمع صنعاء اديس ابابا)
- Dramatic Structure by Sabah Al-Anbari - 23,976 words (البناء الدرامي لصباح الأنباري)
- Dialogue with an Atheist by Mustafa Mahmoud – 12,068 words (حوار مع صديقي الملحد للدكتور مصطفى محمود)
- Yusuf Idris Literary Criticism by Dr. Abir Salama – 54,549 words (نداهة الكتابة: نصوص مجهولة فى إبداء يوسف إدريس للدكتور عبير سلامة)
- In Pre-Islamic Poetry by Taha Hussein - 35,199 words (في الشعر الجاهلي لطه حسين)
- Political Necessity Volume 5 - 15,178 words (الضرورة السياسية)
- Place Aesthetics in Contemporary Arab Literary Criticism by Abdullah Abu Heif - 6,090 words (جماليات المكان في النقد الأدبي العربي المعاصر لعبد الله أبو هيف)
- Letters from Kanafani to Ghada - 18,253 words (رسائل غسان كنفاني إلى غادة السمان)
- Interview with Muammar Al-Gaddafi - 2,714 words (نص حديث الأخ قائد لاثورة لمعمر القذافي)
- New Military Service Law - 524 words (نص قانون خدمة العلم الجديد)
- Interview with Hosni Mubarak – 7,358 words (نص مقابلة الرئيس حسني مبارك مع قناة العربية)
- All
- The 'All' corpus consists of all the texts mentioned and has a total of 173,600,000 words.
- Remember that searching in the larger corpora such as 'All' or 'All Newspapers' will lead to much longer wait times.
- If you combine a larger corpus with a search that will yield hundreds of thousands of results, the machine may balk and not return anything.
- You should only use the 'All' corpus to search for rare or unusual forms. Searching for common words can lead to inaccurate results and may cause problems in subsequent searches.
- In general, the 'All' category is mainly appropriate for researchers looking for broad quantitative or comparative data.
- The smaller categories are usually going to be more appropriate for pedagogical data, or data for students, since it will come in a quantity that you can understand and deal with.
Performing Searches in arabiCorpus
- Basic Searching Instructions
- Type a search word or phrase into one of the two boxes on the left.
- Type it WITHOUT short vowels and without prefixes like the definite article.
- You can either type a transliteration of the word into the ‘latin chars’ box, using the DT transliteration system OR type Arabic script into the ‘arabic chars’ box.
- You can click ‘transliteration help’ to see a chart of the DT transliteration system. Note that there is a one-to-one correspondence for Latin and Arabic characters.
- Your computer must be set up to type Arabic script if you use the ‘arabic chars’ box.
- Choose a POS filter for your search. If you don’t want to use a specific filter, choose ‘string.'
- You must choose one of the POS filters for every search.
- Select a corpus to search from the dropdown list of corpora.
- Only use the 'All' and 'All Newspapers' corpora when searching for rare or unusual forms.
- Press 'submit.'
- If you have not entered a search string or chosen a corpus or POS filter, a little red message will appear near the 'corpus' dropdown box telling you what you have forgotten and no search will be performed.
- Detailed Searching Instructions
- Enter a Search String
- Type a word, phrase, or regular expression into one of the leftmost boxes (not both).
- Click on ‘transliteration help’ to see the transliteration chart.
- Note that there is a one-to-one correspondence between Arabic and transliterated letters.
- To type Arabic, make sure your computer is set up to type Arabic, and use the regular computer input method.
- When entering a search string, a space means a space. If you type two or more words it will look for the exact phrase.
- When searching for a noun, input the noun without the definite article (Al ال).
- When searching for a verb, input the past and present huwa form of the verb separated by a comma.
- For more information on searching for different verb forms, see the 'Searching for verbs' section under 'Important Points on Searching.'
- To find all the examples of what you are looking for, you must be aware of Arabic spelling variations.
- The program automatically removes initial hamzas, meaning that typing L أ or A ا or E إ or M آ will end up finding any of these possibilities.
- For example, a search for ATfAl اطفال will result in both ATfAl اطفال and LTfAl أطفال.
- Other variations, however, are not automatically accounted for by the program. For example:
- Final yaa' varies with final yaa' without the two dots. For instance, you will find both lbnAne لبناني and lbnAny لبنانى in the corpus.
- So, if you search for yly يلي you will find only that form, and no occurrences of yle يلى and vice versa.
- If you want to find both yly يلي and yle يلى at the same time, you have to build it into your search (see 'Important Points on Searching' below).
- Select a Part of Speech Filter
- Part of Speech Filters (Read this before reading the choices below)
- You must select a part of speech filter for every search you perform.
- Any part of speech filter (POS) other than ‘string’ will filter your results according to certain rules:
- The results will contain the search string alone, with only a space or punctuation before or after.
- The results will also contain the search string with the suffixes and/or prefixes that go with the chosen POS filter.
- All other instances of the string will be filtered out (those that are not the bare form or with the known suffixes and prefixes).
- Click on the different POS filters below to see how they will affect your search.
- String
- The ‘string’ POS filter will not apply any filter to the results.
- EVERYTHING that contains your search string will be returned, no matter what is before or after that string.
- For example, if you type ktb كتب, you will get EVERY example of that string in the corpus, including the following:
- Alktb الكتب
- ktbtm كتبتم
- fyktb فيكتب
- AlmktbQ المكتبة
- ystktb يستكتب
- ktbrcAt كتبرعات
- Note that this does NOT mean it will find all examples with a particular root.
- Noun
- Choose the ‘noun’ POS filter and the program will accept the bare search string you typed in along with the following attachments:
- The conjunctions wa- و and fa- فـ
- The attached prepositions ka- كـ and bi- بـ and li- لـ
- The definite article (ال)
- The pronoun endings (هن|كن|نا|هم|تم|تما|هما|ك|كي|ها|ه|ي)
- The alif tanwiin at the end marking indefinite accusative
- The dual ending markers
- For example, if you type in ktb كتب, it will accept:
- ktb كتب
- Alktb الكتب
- wAlktb والكتب
- llktb للكتب
- bAlktb بالكتب
- ktbh كتبه
- ktbnA كتبنا
- ktbhn كتبهن
- bktbhA بكتبها
- The program would also accept ktbAn if it found it, since it assumes you are typing in a singular, and it looks for possible dual endings added to it.
- But it will not accept (it will filter out) results that cannot possibly be nouns or that add more characters than a normal suffix/prefix, such as the following:
- ktbtm كتبتم
- fyktb فيكتب
- ystktb يستكتب
- ktbrcAt كتبرعات
- If you type a noun that ends with the feminine marker (Q ة), the program knows to allow forms that have a t ت instead when there are pronoun endings.
- For example, if you type ktAbQ كتابة it will accept:
- ktAbQ كتابة
- AlktAbQ الكتابة
- ktAbthA كتابتها
- bktAbtnA بكتابتنا
- The fact that you have filtered your results with the 'noun' POS filter does NOT mean that all your results will be nouns. It only means that, given the morphological ambiguity of Arabic, they COULD be nouns.
- For example, if you type in ktb كتب and choose ‘noun,’ some of the results will unambiguously be nouns, such as:
- Alktb الكتب
- bktbhA بكتبها
- while others will simply be ambiguous:
- ktb كتب (could be 'books' or 'he wrote')
- ktbnA كتبنا (could be 'our books' or 'we wrote')
- ktbh كتبه (could be 'his books' or 'he wrote it')
- As mentioned, it WILL filter out forms that are unambiguously NOT nouns, which can be very helpful.
- Choosing a filter, therefore, reduces the number of false hits, but does not eliminate them entirely.
- REMEMBER that this tool CANNOT and DOES NOT try to overcome the ambiguities of Arabic morphology. If a form is inherently ambiguous, you WILL find it even though, in context, it is NOT what you are looking for!
- Remember that the words in the corpora are not tagged for part of speech. So, in the end, YOU decide if the word is a noun, verb, etc.
- Note: Remember, it only searches for what you type in. If you type in a singular form, it will only search for that. If you want to find the plural forms, you must search for them separately or at the same time in a combined search.
- Adj
- Choose the 'adj' POS filter and the program will accept the bare search string you typed in along with the following attachments:
- The conjunctions wa- و and fa- فـ
- The definite article (ال)
- But NOT the prepositions or the pronoun endings
- The alif tanwiin at the end marking indefinite accusative
- The masculine and feminine dual endings
- The feminine marker (Q ة) (Note: the 'noun' POS filter will NOT do this)
- Note that the program will only search for all forms if you input the masculine form, so do not add the feminine marker unless you ONLY want those results.
- For example, if you type in jmyl جميل, it will accept:
- jmyl جميل
- Aljmyl الجميل
- jmylQ جميلة
- wAljmylQ والجميلة
- But it will not accept (it will filter out) results that cannot possibly be adjectives, such as:
- jmylhA جميلها
- bAljmyl بالجميل
- If you want to try to find forms like jmylhA جميلها and bAljmyl بالجميل, you need to search for jmyl جميل as 'noun', not as 'adj'.
- Alternatively, you could search directly for jmylhA جميلها and bAljmyl بالجميل as strings, bypassing the POS filters.
- Remember that choosing 'noun' or 'adj' or any other POS filter does not mean that the results will actually BE that form or even that you WANT to find that form. It only means that you want it to allow that particular set of prefixes and suffixes through the search filter.
- As with nouns, if you want to find the plurals as well you must explicitly include them. It does not happen automatically.
- Adv
- Choose the 'adv' POS filter and the program will accept the bare search string you typed in and the conjunctions wa- و and fa- فـ
- Nothing else will be accepted
- This category is handy for adverbs, but is also useful when you are searching for a specific form and not an entire conjugation.
- For example, if you want to find the noun ktb كتب only when it has the definite article preceded by b بـ (and not all the other possible forms of ktb كتب), type bAlktb بالكتب and choose 'adv'. It will accept the following:
- bAlktb بالكتب
- fbAlktb فبالكتب
- wbAlktb وبالكتب
- Nothing else will be accepted.
- The same technique works if you want to find a specific verb form (nqwl نقول) as opposed to all the conjugations of that verb.
- Remember, choosing the 'adv' POS filter does not necessarily mean you think the word is an adverb. It means that you want the filter to cut out everything except the specific form you typed in, plus that form with wa- و or fa- فـ.
- Choosing 'adv' is almost the opposite of choosing 'string'. 'String' accepts every occurrence of the string in the corpus, no matter what else surrounds it, and 'adv' accepts basically only what you typed in as a complete, isolated word and nothing else.
- Verb
- Choose the 'verb' POS filter and the program will accept the bare search string you typed in along with the following attachments:
- The conjunctions wa- و and fa- فـ
- The particle li- لـ
- The future prefix sa- سـ
- The colloquial prefixes Ha- حـ and bi- بـ
- The pronoun endings
- The perfect and imperfect verb conjugation suffixes and prefixes
- IMPORTANT NOTE: The program assumes you will type in the masculine singular (huwa) form of both the past and the present tense, separated by a comma. However, if you only type the first, it will usually guess the second correctly.
- For example, if you type in ktb,yktb كتب،يكتب, it will accept:
- ktb كتب
- ktbt كتبت
- wktbwA وكتبوا
- ktbwhA كتبوها
- ktbth كتبته
- yktb يكتب
- fsyktbwn فسيكتبون
- lnktbhA لنكتبها
- It will also accept some fairly ludicrous forms like:
- flLktbny فلأكتبني
- nktbnA نكتبنا
- This happens because the program more or less mechanically applies the rules without asking what the resulting forms mean.
- The suffixes and prefixes are applied to the exact string(s) you type in and nothing else. (You will not necessarily find other forms of the verb).
- The program tries to handle double, hollow, defective, assimilated and hamzated verbs correctly, but its analyses are not perfect, so you should check the search strings it actually uses on the Summary Page.
- The program will usually guess when you put in one of these irregular forms but it may guess incorrectly. For more accurate searches of this type, see the 'Searching for verbs' section under 'Important Points on Searching.'
- The program will alert you when it has made a guess so that you will know to check which search strings it used.
- Choose a Corpus
- Select one of the corpora from the dropdown list labeled 'corpus.'
- arabiCorpus gives you the ability to search from a selection of Arabic corpora in a variety of genres and topics.
- In Basic Search, only the newspapers, the Quran, and 1001 Nights can be searched individually and the other texts are grouped according to genre into larger combined corpora. You can search any text individually by going to the Advanced Search mode and selecting one of the texts in the dropdown list.
- In addition to choosing what type of text you want to search, it is important to realize that choosing larger corpora, especially the corpus combinations, will often make your ‘wait’ time MUCH longer.
- Using the 'All' or 'All Newspapers' corpora when searching for common words may also cause problems in subsequent searches. The extra long search time may cause a search to "carry over" to the next search, resulting in multiple searches appearing in one set of results. Only use these large corpora when searching for rare or unusual forms.
- Therefore, if you are searching for an uncommon word, using the combinations can be valuable, since this will give you more examples.
- On the other hand, if you are searching for a common word, using the combinations can be unhelpful; aside from taking a long time for the search, it will return so many results (possibly 10s of thousands) that you will not be able to deal with them efficiently.
- In general, the ALL category is mainly appropriate for researchers looking for broad quantitative or comparative data.
- The smaller categories are usually going to be more appropriate for pedagogical data, or data for students, since it will come in a quantity that you can understand and deal with.
- You should be aware that EgyptChat includes some colloquial and a lot of fusha and mixed fusha and colloquial, while the literature and even the newspapers do include some colloquial.
- For more information and a full listing of all the works that are found in each corpus, please see the 'Detailed Corpus Information' section.
- Important Points on Searching
- Searching for verbs
- There are two different ways to search for various verb forms in the corpora.
- In Basic Search, the program does a large amount of guessing and figuring to create the actual search strings it uses to search the corpus.
- In Advanced Search, you create the exact search strings the program uses. It doesn't try to second guess you. This gives you more control, but it also allows you to do searches that are relatively ridiculous.
- To learn about the different POS filters available in Advanced Search and how to use them, please see the Advanced Search Instructions by switching to Advanced Search mode.
- The method for searching for verbs in Basic Search is similar to that of Advanced Search. The main difference is that Basic Search has only one verb POS filter whereas Advanced Search has multiple.
- You should find similar results in either mode but Advanced Search can better predict what you want. If you are not seeing the results you expect, you may want to try Advanced Search. (Remember, however, that Advanced Search is still not perfect).
- You can use the 'verb' POS filter in Basic Search to find a variety of verb forms. The program will usually guess which form you want. Consider the following possibilities:
- For verbs in which the perfect and imperfect stems are identical, simply type in the past tense huwa form.
- Search words would look like this: ktb or prb
- For verbs in which the perfect and imperfect stems are different, type in both the huwa past and huwa present tense stems. The 'y' on the imperfect stem is optional and will always be deleted when the search is actually performed.
- Search words would look like this: Lkrm,ykrm or wSl,ySl.
- Form IV, VII, VIII and X verbs, Form I assimilated verbs, and Form II verbs with initial hamza are among the verbs with a different stem in the perfect and imperfect.
- For verbs in which the perfect and imperfect both have two stems, you must type in four forms. In the case of hollow verbs, the stems that must be typed in are the perfect-long form, perfect-short, imperfect-long, and imperfect-short. In the case of doubled verbs, the stems must be the perfect form with only one of the doubleds, the perfect with both, the imperfect with only one, and the imperfect with both.
- Use this type of search exclusively for hollow and doubled verbs.
- Search words would look like this: qAl,ql,qwl,ql and AHtl,AHtll,Htl,Htll
- Again, the 'y' on the imperfect forms is always optional and will work either way.
- For defective verbs, you only need to type in the huwa past and huwa present tense stems.
- Search words would look like this: bne,ybny or lqy,ylqe.
- This type of search may be better in Advanced Search because 'verbd' sometimes better predicts the variations of defective verbs.
- Basic Search will always make guesses for what type of search you intended even if you do not input the search correctly.
- The advantage to using Advanced Search is that you know exactly what the program will be searching for when you input a search string.
- Either way, you should check the Summary Results page to see if any alerts have appeared to tell you when a search was input incorrectly or to see what guessing the program performed.
- Searching with or without vowels
- The program automatically strips the vowels and kashidas from both the search string and the text before it searches.
- In order to perform searches with vowels, you must use Advanced Search and check the box labeled 'include vowels.'
- For more information on these types of searches, see the Advanced Search Instructions.
- Searching for words with hamzas
- Words hamzas present a special challenge for all search engines because texts are incredibly inconsistent.
- To help remedy part of this problem, the program automatically searches for all initial hamza possibilites (A ا and L أ and E إ and M آ) when any one of them is typed in.
- Consider the following examples as to why this feature is included:
- Without this feature, the program would only search for what you type in.
- So, if you searched for LTfAl أطفال, it would allow:
- LTfAl أطفال
- LTfAlh أطفاله
- AlLTfAl الأطفال
- But NOT:
- ATfAl اطفال
- ATfAlh اطفاله
- AlATfAl الاطفال
- If you typed in ATfAl اطفال, the opposite would happen, and you wouldn't get any of the examples with the hamza.
- The only way that you could guarantee that you get both is to type both within square brackets: [AL] [اأ]
- Square brackets are the regular expression way of indicating that anything within them can go in that position.
- If you typed [AL]TfAl, the program would accept:
- LTfAl أطفال
- LTfAlh أطفاله
- AlLTfAl الأطفال
- ATfAl اطفال
- ATfAlh اطفاله
- AlATfAl الاطفال
- Sometimes there is a hamza on the alif at the beginning of form VIIIs, so to get them all you would have to type [AEL]xtAr, etc.
- To help eliminate the extra confusion, the program automatically looks for all hamza possibilities.
- If you DO want to search for a specific form, i.e. if you are looking only for LTfAl أطفال and NOT ATfAl اطفال, then you can click on the 'differentiate initial hamzas' box on the Advanced Search page.
- In that case, it will search only for the exact hamza sequence you type in and nothing else.
- See the Advanced Search Instructions for more on this feature.
- The other issue is that texts can also be inconsistent with hamzas within words. This issue cannot be resolved automatically by the program, so you will have to account for the variability in your search.
- For example, both AlMn الآن and AlAn الان can be found in the corpora.
- If you wanted to find all examples, you would have to search for Al[MA]n ال[آا]ن.
- As always, you need to think of the different variations possible if you want to find all of something.
- Searching for other spelling variations (alif maqsuuras, yaa's)
- You must consider other spelling variations that may exist in the various texts found on the program.
- For example, the Ahram 1999 corpus (which came from the internet) uses the letter yaa' for the alif maqsuura, although not consistently.
- So if you want to find ALL examples of the word mqhe مقهى in the Ahram 1999 corpus, you must type: mqh[ey]
- Likewise, most of the newspapers sometimes leave the dots off words that are supposed to be yaa's. So you must search for them also with [ey] if you want to get everything.
- There are many other such variations you need to take into account to be able to find all of anything.
- For example, the Ahram rarely leaves a space between the two parts of lA bd لا بد, while the Hayat fairly consistently puts the parts together (lAbd لابد).
- Searching for more than one item at once
- If you search for more than one item at once, it will sort the results all together.
- This can be handy, for example, if you want all examples of a noun and its plural.
- To search for two items at once, type in the items separated by a comma. Be sure NOT to add extra spaces.
- To find 'book' and 'books', type: ktAb,ktb كتاب،كتب and choose 'noun.'
- To find examples of two different ways to say 'during', type xlAl,[LA]VnAC خلال،[أا]ثناء and choose 'adv.'
- If the words you type in go with different POS filters (say 'adv' and 'noun'), the results will be unpredictable.
- It is not possible to search for more than one word at once when the 'verb' POS filter is chosen, since the comma separated list is used in verbs to list the various stems of the single verb.
- This does not mean you cannot search for more than one verb at once; you just cannot use the 'verb' POS filter.
- You would, however, need to be skilled at using regular expressions to find the kind of results you want.
- Searching for phrases
- Typing a space will cause the program to look for a space.
- This means that if you type two or more words separated by a space, it will look for the entire phrase.
- Most of the POS filters make no sense when searching for phrases. However, sometimes they can be made to be helpful.
- Normally you would choose 'string' to simply find the phrase no matter what else is around it.
- If you want to find the exact phrase but want to allow wa- و and fa- فـ, then choose 'adv.'
- If you want to allow the last word in the phrase to have pronoun endings, choose 'noun.'
- If you type: lA gbAr clyh لا غبار عليه and choose 'adv' it will allow:
- lA gbAr clyh لا غبار عليه
- flA gbAr clyh فلا غبار عليه
- wlA gbAr clyh ولا غبار عليه
- and nothing else.
- Searching with Regular Expressions
- Using regular expressions can greatly increase the power of your search.
- Some aspects of the regular expression language don't work yet, so be patient. Most parts do work however.
- You can use the backslash characters that mean specific things:
- \w = any word character
- \s = any space character
- \b = a word boundary
- You can also use quantifiers with these:
- \w? = 0 or 1 word characters
- \w+ = 1 or more word characters in a row
- \w* = 0 or more word characters in a row
- A? = either an A or not
- You can use square brackets to indicate a list of things, one of which can go there, and you can use quantifiers with the list:
- [hny] = either a haa' or a nuun or a yaa'
- [hny]+ = any combinations of haa', nuun and yaa' in a row
- [aui~]? = either a fatha, damma, kasra, shadda or nothing
- You can use ^ at the beginning of this list to make it a list of things that cannot go there:
- [^AL] = anything except A or L
- You can use ^ (outside of the brackets) and $ to indicate the beginning and end of a word, respectively:
- ^br (with the filter 'string' chosen) means find every word that begins with br
- At$ (with the filter 'string' chosen) means find every word that ends with At
- You need to think carefully about how your regular expression is going to interact with the POS filter you have chosen. For example:
- If you type \bktAb and choose the 'adv' filter, the only thing it will allow is:
- ktAb
- since the 'adv' filter already allows no suffixes, and only allows the wa- و and fa- فـ prefixes, which are cut out by your \b restriction.
- If you type \bktAb and choose the 'noun' filter, however, you will get the bare stem and the word with pronoun endings.
- If you want complete control, choose 'string', and only what you cut out 'by hand' will be cut out.
- Indicate alternation between whole forms using parentheses and the vertical bar. You must match your parentheses or you will get an error.
- To have the program find mdrs, mdrswn, mdrsyn, or mdrsAt, type mdrs(wn|yn|At)? or alternatively mdrs([wy]n|At)?
- Using parentheses and the vertical bar in combination with square brackets and square brackets with the carrot (^) can be a VERY powerful way to design precise searches.
- Using Arabic Script in regular expressions
- Regular expressions work in Arabic script as well as they do in English (\w = \و etc.)
- However, most browsers display the backslashes and parentheses in bizarre and unpredictable ways. If you type the search in carefully, it should work, but if you try to read it after you have typed it in, or on the summary page, you will probably have trouble.
- Cutting out results by hand
- Sometimes you know that the ambiguous morphology of a particular form is going to give you thousands of false hits that you don't want to look through to find what you are really looking for.
- For example, if you search for all forms of the verb bne ybny بنى يبني, you will get many examples of the imperative form Abn ابن (almost all of which are really 'son') and of the third person feminine past tense form bnt بنت (many of which are really 'girl')
- You may choose to do the search with those forms deleted so that you will have an easier time going through the remaining citations looking for something specific.
- To cut out forms that otherwise would be found by a search, type the search, a space, one or two dashes, a space, and then a regular expression that indicates what you don't want
- You can use a vertical bar to cut out multiple things.
- To cut out the exact form Abn ابن and nothing else (choose 'verb'), type: bne,ybny -- ^Abn$
- This will cut out Abn ابن but not Abny ابني.
- To cut out bnt بنت no matter what is after it (choose 'verb'), type: bne,ybny -- ^bnt
- This will cut out bnt بنت and also bntk بنتك, etc.
- To cut out any form that has the sequence bnt بنت anywhere in it (choose 'verb') type: bne,ybny -- bnt
- To cut out both Abn ابن and bnt بنت at the same time (choose 'verb'), type: bne,ybny -- bnt|Abn
- To cut out bnt بنت only when it is alone or preceded by wa- or fa-, type: bne,ybny -- ^[wf]?bnt$
- Hopefully you get the idea. Like searching with regular expressions in general, this can be a very powerful tool in helping refine your searches.
- Searching for punctuation
- This tool currently does not have a way to search for punctuation.
Analyzing the Results of a Search
- Basic Information about Results
- The results come back after waiting for a few seconds.
- If you search for a single noun or adjective in a single corpus, it should come back in about 10 seconds
- If you search for multiple nouns, or verbs, or search in the combined corpora, the results will take longer.
- If you search for a very common word that produces tens of thousands of hits or more, the program will sometimes balk because the database it is using will limit what can be inserted.
- If the search takes too long and returns no results, the best solution is to log out of the program and log back in. You may have to wait a few minutes before the program returns back to normal.
- The results of a search are saved temporarily in a database, and you can access them by clicking on the words that appear in the red bar.
- Summary Page
- This page comes up first automatically after a search, and you can return to it by clicking on 'summary' in the red bar.
- This page gives you summary information about your search:
- The search string you typed in, in both English and Arabic scripts
- The search strings actually used by the program after its 'figuring'
- The corpus you searched
- The time it took the search engine to do the basic search (this will generally be less than the actual time experienced by you, since it doesn't include the time it takes for the server to receive your request or to serve the results back to you)
- The POS filter you chose
- The POS filter that was actually used
- The number of 'hits'
- What that number translates to in terms of words per 100,000 words of your corpus
- The latter bit of information is useful for comparative purposes. The corpora are of vastly different sizes, so the actual number per corpus may be misleading, but the number per 100,000 words can be compared.
- This page also displays various alerts according to the search terms you entered. Usually, these alerts will appear when you have used the 'verb' POS filter and will inform you of any guessing the program has performed and how you might need to change your search.
- It is important to read these alerts so you will know what the program is searching for and how to change your search if you do not find the results you are looking for.
- Citations Page
- Click on 'citations' in the red bar to see the 'hits' with the 10 words before and the 10 words after.
- By default, these citations are sorted by the word that appears directly before the word you searched for.
- Note that this sorting is done by the whole word, not by the root, so AlktAb الكتاب, ktAb كتاب, and wAlktAb والكتاب are nowhere near each other.
- The word directly before is repeated at the left hand side of the page so you can quickly glance through the citations and notice patterns, collocations, etc.
- If you would rather see the citations sorted by the word directly after the 'hit', click on the red sentence near the top of the page that says 'sort by word after.'
- The citations are shown 100 at a time.
- If your search returned more than 100 citations, they will be organized into 'pages' which you can access by clicking on the red numbers at the top.
- The far right column on the Citations page shows you the subsections of the corpus you searched.
- If you searched a single corpus, the subsection column will indicate what subsection of the corpus the example comes from.
- The newspapers have self-labeled subsections, so sometimes you need to figure out what the codes mean, but they are related to the various sections of the newspaper.
- If you have searched a combined corpus, the subsection column will tell you which specific corpus the example came from.
- If you want to see the exact reference of your example, click on the red 'subsection' heading and it will change to 'reference.'
- The references in many of the corpora are basically incomprehensible and not very helpful, but the references to some, like the Quran and the novels, are useful.
- The references in the Quran are to Sura and verse.
- The references in the novel are to chapter, section of chapter (divided by stars), and paragraph number.
- The references in 1001 Nights are to Nights and paragraphs (but in an odd kind of way, which I'll explain if you ask me by e-mail).
- If you want to see more context than the 10 words before and the 10 words after, click on the number at the far left of the citation.
- This will bring up a separate window that will display the whole verse from the Quran, the whole paragraph from the novel and 1001 Nights, or the whole article from the newspapers.
- You may have to use your browser's Arabic script search function to find the place in the larger context where your item appears, since it does not highlight it.
- Subsections Page
- Click on 'subsections' in the red bar to see the totals for the various subsections of the corpus you searched.
- Again, if you searched a single corpus, those subsections will be those defined for that corpus.
- Some corpora, particularly the non-news ones, have no subsections defined.
- If you searched a combined corpus, the subsections will be the names of the corpora in that combined corpus.
- The results in this section are ordered from most frequent to less frequent in terms of number.
- The frequency per 100,000 for each subsection is also given.
- This page allows you to compare, say, Premodern Arabic with Modern, or to compare newspapers from various places.
- Type lA bd لا بد and choose All Newspapers, for example, click on subsections, and you will find that the Ahram uses it much less frequently than the Hayat.
- Type lAbd لابد, though, and choose All Newspapers and you will find the opposite (the Ahram apparently typically types the phrase without a space, and the Hayat does the opposite).
- Word Forms Page
- Words Before/After Page
- Click on 'words before/after' in the red bar to see a list of the common words directly before and directly after the 'hit' ordered by frequency.
- Examining these lists is a good way to scope out the main usages and collocations of the word you searched for.
- Examining the lists can also help you identify structures that you had not intended to search for that you want to cut out.
- If you want to see only the citations with that particular word before or after, click on a word in the list.
- A separate window will open showing you just those citations.
- Collocates Page
- Click on 'collocates' in the red bar to see a list of the most common collocates that appear with your search term.
- For this program, collocates are defined as the most frequent word forms that appear up to four positions to the left or right of your search term.
- The program only lists those collocates that appear at least 4 times in the given corpus.
- You cannot click on the word to see the occurrences in context like with other results pages.
- Like the Words Before/After Page, this page allows you to see what words occur immediately before or after your search term, but it also allows you to see those words that commonly appear near your search term and not just right before or after.
Tutorial
- First Search: Look for a single noun
- Click on the red Instructions above the submit button to get these instructions in a separate window (otherwise they will go away as you follow them).
- Perform a simple search for a noun.
- Type mktb into the leftmost box. Do NOT type any vowels (i.e. do NOT type maktab). Do NOT type the definite article (do NOT type Almktb).
- Choose 'noun' from the POS list.
- Choose Ahram 1999 from the corpus list.
- Click on 'Submit'.
- Wait about 10 seconds.
- Analyze the results of your search.
- Examine summary of search that appears, and notice how many examples it found, and how frequent these are in the corpus (per 100,000 words).
- Click on 'citations' in the dark red bar.
- Scroll down and look at a few of the citations. Scroll back to the top.
- Note that there are about 40 pages of results. Click on page 25.
- Note that each example gives you the word in context with 10 words before and 10 after.
- Note that the examples are organized by the word that comes before (here fy), and that this word is also listed at the beginning of each line so it can be easily picked out.
- Click on 'sort by word after' to sort the examples by the word after instead.
- Click on page 25 again, and see what kind of information you get with this order.
- Click on one of the numbers at the left, and see a new window open with even more context. Close that window
- Click on 'subsections' in the dark red bar.
- Examine the frequencies and relative frequencies of this word in the various sections of the Ahram.
- Click on 'word forms' in the dark red bar.
- Notice the different forms in which this word was found: alone, with wa-, bi-, fa-, with the definite article, and with various pronoun endings. Notice which of these forms were more common and which relatively rare.
- Click on مكتبك in the middle of the second column to see the citations just for that word form in a separate window. Examine briefly and close the second window.
- Click on 'words before/after' in the dark red bar.
- Examine the most common words that come before our search word and the most common words that come after. Any surprises? Is this what you would have predicted?
- Click on التنسيق in the second column to see the 200 examples of this word coming after our search word in a new window. After examining for a few moments, close the new window.
- Click on 'summary' in the dark red bar to go back to the summary page.
- Try other single words
- Type jmyl into the leftmost box, choose the 'adj' POS filter, the Ahram 1999 corpus, and click submit.
- Go through the various choices (as under First Search) and see the differences. Note particularly under 'word forms' that you get examples with and without the feminine ending, but no examples with pronoun endings or prepositions.
- Now try running the same word on the same corpus but with the 'noun' POS filter chosen. Note under 'word forms' that you don't get the feminine forms, but you do get forms with prepositions and pronouns.
- Type tqrybA into the leftmost box, choose the 'adv' POS filter and the Ahram 1999 corpus and run through the various options once you get the results.
- Type ktAbhm into the leftmost box, choose the 'adv' POS filter and the Ahram 1999 corpus, and look to see what you got under word forms.
- 'adv' is a good POS filter when you are looking for a very specific form, and don't want to see other possible prefixes and suffixes.
- Type prb into the leftmost box, choose the 'verb' POS filter, the Ahram 1999 corpus, and run through the various options looking at the results. Note particularly the large number of forms under 'word forms'.
- Now run prb again, but this time choose the 'string' POS filter.
- Note under word forms that besides all the verb forms you got a moment ago, you are now getting a bunch of other things that happen to contain the sequence prb, including a bunch of typos.
- Account for variabilty in the corpus
- There are several ways to have the corpus look for alternates (either this OR that). The easiest involves square brackets.
- The square brackets tell the program to search for any one of the characters in the brackets.
- Type lbnAne into the leftmost box, choose the 'noun' POS filter, and the Ahram 1999 corpus.
- Notice that you get 20 hits that end with an alif maqsura. If you had searched for lbnAny and wanted to see all examples of use of this word, the program would have missed these 20.
- In general, if you want ALL examples of forms that end in y, you would need to use [ey] instead.
- The Ahram itself is a special case because for whatever reason they use the yaa' for the alif maqsuura much of the time. So if you search for mqhe (with 'noun') in the Ahram 1999 corpus you get zero results, but if you search for mqhy you get 158, and examining their citations reveals that they are meant to be mqhe.
- Therefore, if the Ahram is included in your search, you should use [ye] for words that end in y AND e.
- Now type msWwl into the leftmost box, choose the 'adj' POS filter and the Newspaper combined corpus. You will have to wait longer for the results because of the larger corpus.
- If you check out the subsections page, you will see that there are a huge number of hits in the other papers, but very few in the Ahram.
- Now type msYwl into the leftmost box, choose the 'adj' POS filter, and the Newspaper combined corpus.
- Notice in the subsections page that almost all the examples are from the Ahram, and very few from the others.
- If you wanted to catch ALL instances of this word in the Newspaper combined corpus, you would need to type ms[WY]wl.
- Now type tlfwn into the leftmost box, choose the 'noun' POS filter, and the Ahram 1999 corpus. Look how few the results are.
- Now type tlyfwn into the leftmost box, choose the 'noun' POS filter, and the Ahram 1999 corpus. Look how many the results are.
- If you want to find tlfwn and tlyfwn together you could type tly?fwn. The question mark tells the program the yaa' can either be there or not.
- The program is really just an automoton. It doesn't know what you want. It just looks for exactly what you type in. If you need it to pay attention to spelling variation, then you must tell it to do so.
- Moral: PAY ATTENTION to spelling variations in Arabic, and account for them in your searches, often by using square brackets or question marks.
- Search for more than one word at once
- Type mktb,mkAtb into the leftmost box. Note that there is no space after the comma. Choose the 'noun' POS filter and the Ahram 1999 corpus.
- Notice on the word forms page that the two words have been combined into a single search. You can string any number together with commas.
- Now try the following combinations:
- hjwm,hjmAt,hjwmAt,hjmQ,mhAjmQ (check out the word forms page to get an idea of relative frequency)
- bTbycQ AlHAl,TbcA,bAlTbc (ditto)
- hAtf,tlfwn,tlyfwn,mwbAyl,mHmwl,xlwy
- Filter results 'by hand'
- Type lbnAn into the leftmost box, choose the 'noun' POS filter, and Ahram 1999 corpus.
- Examine the word forms. Notice that you have quite a few examples of lbnAny, which the program got by putting the y suffix for 'my' on the end of the noun.
- However, if you click on this word on the word forms page and examine these citations, you will see that none of them mean 'my Lebanon'. The y here is a nisba adjective suffix.
- In other words, you have gotten a bunch of things you don't want because of the morphological ambiguity of Arabic.
- Let's say you want to look through all the citations with lbnAn, and you don't want to have to deal with all those lbnAny's.
- Type lbnAn -- y$ into the leftmost box, choose the 'noun' POS filter and Ahram 1999 corpus.
- The two dashes tell the program that whatever you type after them is NOT wanted (filter it out).
- The dollar sign means that this is at the end of the word.
- So this means, search for all examples of lbnAn with normal noun suffixes, but if you find one the ends in a y, filter it out.
- A look at the word forms list will convince you that it has now gotten rid of the unwanted lbnAny's.
- Try the following searches with and without the dash section and check the word forms page to see what it cuts out:
- mktb -- ^m (this will cut out all word forms that begin with a miim, i.e. all that don't have some other prefix)
- mktb -- h (this one will cut out all word forms that have a haa' in them, meaning that mktb with certain pronoun endings will be cut out.
- mktb -- h|y|n (this one will cut out all word forms with a haa', kaaf, yaa', or nuun, meaning that maktb with all but second person pronoun endings will be filtered out. Note that the vertical bar represents alternation.
- mktb -- h|y|n|km?A?$ (this one also cuts out the second person forms, but you have to be careful: if you just list kaaf by itself it will cut out everything, since kaaf is in the word itself. So you need the question marks and the dollar sign.
- A good strategy is to perform plain searches first, and then examine the word form list to see if there are any forms you would like to filter out by hand. If there are, create a 'dash' phrase to do the job.
- Note for those with a 'regular expression' background: You can consider the list after the dashes to be a regular expression for which the program automatically provides an outside set of grouping parentheses.
- Search for a phrase
- Type lA gbAr into the leftmost box, choose the 'adv' POS filter, and Hayat 1997.
- Check out some of the citations.
- Try these other phrases:
- Type (Al)?bnyQ (Al)?tHtyQ into the leftmost box, choose the 'string' POS filter and Hayat 1997.
- Type pyk bdwn rSyd into the leftmost box, choose the 'string' POS filter and Hayat 1997 (one hit only, check out the citation)
- Type ttHlb lh Al[LA]fwAh into the leftmost box, choose the 'string' POS filter and All (one hit only, check out the citation)
- Type \bwzArQ Al\w+ into the leftmost box, choose the 'string' POS filter and ShuruqColumns (to see a list of different ministries)
- Try AlwzArQ Al\w+ (obviously not as useful)
- Type mA lbV [LA]n into the leftmost box, choose the 'string' POS filter and ShuruqColumns.
- Try (mA lbV|lm ylbV) [LA]n with the 'string' POS filter and Hayat 1997 (be prepared to wait)
- You are basically limited only by your imagination. However, if you want to be sure to get the results you want, remember to account for hamza variations and other spelling variations.
- Also remember that some writers do not put spaces where others do (some write lAbd and others lA bd). You can capture this with lA ?bd (try it with ShuruqColumns).
- Compare results of different corpora
- Type hAtf into the leftmost box, choose the 'noun' POS filter and All Newspapers.
- Go to subsections and see what papers use it more commonly than others.
- Now type tly?fwn into the leftmost box, choose the 'noun' POS filter and All Newspapers.
- Compare the subsection results with the ones you got for hAtif.
- Type lqd into the leftmost box, choose the 'adv' POS filter and All.
- Note in subsections that the Quran and the novel have a much higher rate of use than the other corpora, and that the Ahram uses it more frequently that the Hayat.
- Type shl into the leftmost box, choose the 'adj' POS filter and Hayat 1997.
- Look in subsections and see which parts of the paper like this adjective more than others.
- Then try it with Ahram 1999 and see if the same pattern holds. Look at the words before words after page to see if you get any hints as to why this may be so.
- Try the same thing with hjwm (with 'noun') and see if the subsection patterns make sense to you.
Miscellaneous Info
- Recommended Browser
- You should be able to use this program with any of Google Chrome, Firefox, or Safari.
- Internet Explorer is the only browser that seems to cause some problems, so it is recommended to use one of the others mentioned.
- Note About Numbers
- Because of how the numbers were originally entered into many of the texts making up the corpora, some of the numbers appear reversed or backwards after performing a search (i.e. 1999 may actually appear as 9991). We were unable to fix this problem, so you will simply have to accept that the digits of numbers might be reversed in any particular case.
- Therefore, we do not recommend that you use this program to search for numbers.
- Warning
- This tool can easily become annoying if you are searching for a common word in a large corpus (you will have to wait a long time or the program may do nothing at all).
- Only use the large corpora 'All' and 'All Newspapers' when searching for rare or unusual words.
- When learning to use the tool, try a small corpus and search for slightly less common words.
- You may want to switch to Advanced Search mode and select an individual text from the corpus list.
- Common Errors to Avoid
- Typing a noun or adjective with the definite article (the tool automatically looks for nouns and adjectives with and without the article; if you type it with the article, it will only find examples with it).
- Choosing the wrong POS filter (looking for an adjective with 'noun').
- Choosing 'noun' when looking for a phrase (you should usually choose 'adv' or 'string').
- Searching for a common word in the 'All' or 'All Newspapers' combined corpora.
- Responses to Common Questions/Problems
- My results seem to be showing a mix of different searches. What is wrong?
- You have probably just performed a search that took a very long time or appeared to time out.
- If you used the large combined corpora 'All' or 'All Newspapers' to search for a common word, the search may have appeared to fail because it took so long. If you perform another search in this time, both searches will be running simultaneously and your results may show both searches.
- Avoid using these large corpora when searching for common words. Limit searches in these corpora to rare or unusual forms.
- The best way to manage this problem is to log out and wait a few minutes until performing another search.
- It is a good idea to perform a few simple searches with few results first to check when the search results have returned to normal.
- It is also possible that you were logged in as 'guest' at the same time someone else was using the site as 'guest.' This can cause the same issue.
- This is easy to solve! Simply register with your e-mail address and this problem will occur far less often. It is free and easy and will be much more beneficial for you.
- Why do I want to filter my results?
- If you are trying to find examples of a particular word or construction, and the program finds hundreds of results of something else that happens to match what you are looking for (because of the morphological ambiguity of Arabic), it can be very time consuming and annoying to search through the citations one by one looking for the ones you really want. If you can figure out a way to filter out the 'bad' ones, you can save yourself many hours.
- You want to take care of filtering yourself, and bypass the POS filters.
- Choose 'string' and type a regular expression that matches what the POS filters do, or which varies them.
- \bgbAr\b will find ONLY the word gbAr with no prefixes or suffixes.
- \b[wf]?gbAr\b allows gbAr, wgbAr and fgbAr and nothing else.
- \b[wf]?(Al)?gbAr\b allows gbAr, wgbAr and fgbAr with and without the definite article.
- \b[wf]?(Al)?gbAr(h|hA|k|y)?\b allows all of the above and the singular pronoun endings.
- \b[ytnLA]ktb[wyA]?[nA]?\b is one way to find present tense verb forms, here without any other prefixes or suffixes.
- The program can deal with quite complex expressions efficiently, as long as parentheses are matched, so you can get quite a bit of control over what you are searching for if you want it.
- You want to find all examples of the verb twj, but looking at the results you realize that you have many instances of twjh, which could conceivably mean 'he crowned him' but which 99% of the time is in fact another verb: tawajjaha 'to head for'.
- Choose 'verb' and search for twj -- h. This will delete all the twjh's. You may miss a couple of 'he crowned him's but it's worth it because you can now see that your results are mainly what you want and expected.
- You want to find all examples of the verb qAl, including passives.
- In Advanced Search, searching for qAl,ql,qwl,ql with 'verb4' will give you all the actives. Searching for qyl,ql,qAl,ql will give you all the passives. Searching for q[Ay]l,ql,q[wA]l,ql will give you both all sorted together. Try this on a smaller corpus like 1001 Nights first, since it generates a huge number of hits.
- You want to find all examples of the verb Lyd ('to support).
- In Basic Search, the program will often guess incorrectly when you search for this verb, thinking it is supposed to be a hollow verb.
- In Advanced Search, searching for Lyd using 'verb' will miss all the imperfects and all the perfects where the hamza was not typed on the alif.
- Searching for [LA]yd will get the hamza/alif variation, but will still miss the imperfects. To get those you need to choose 'verb2' instead and type: [LA]yd,Wyd i.e. the perfect stem and the imperfect stem with a comma in between. This should give you what you are looking for.
- You want to get comparative statistics on use of the verb forms yumkinu and tumkinu when they have a feminine verbal noun subject, but when you search for tmkn (using 'adv') you get overwhelmed with things that are really tamakkana, tamakkun, tumakkinu and the like.
- This one is a good example of the inherent ambiguity of much of Arabic morphology, particularly graphemically. To get a statistic you can rely on, particularly if you don't have the stomach to actually read through hundreds of citations, you probably need to do what I call a 'limited search.' Instead of looking for all instances of tmkn, do a search for [yt]mkn \w+(Q|t(h|hA|k|hm|hmA|nA)), which will give you a list of all the words that end in a taa' marbuuta or a taa' followed by a pronoun ending after one of these verbs, and look at the list of words before and after. Look down the list of words after and pick out (say) the top 25 (or 10 or 5) feminine verbal nouns that come directly after ymkn or tmkn, and make sure they 'feel' right (i.e. you are pretty sure that the nouns you are choosing really are likely to function as the subjects of these verbs).
- Then do a search for those forms only, once with ymkn before, and once with tmkn. You don't have an overall statistic, but you have a 'subset' statistic that you can trust. Sample search strings with 5 verbs I found would be:
- ymkn (AstfAd|tsmy|syTr|zyAd|mwAjh)(Q|t(h|hA|k|y|hm|hn|km|kn|nA|hmA|kmA))
- tmkn (AstfAd|tsmy|syTr|zyAd|mwAjh)(Q|t(h|hA|k|y|hm|hn|km|kn|nA|hmA|kmA))
- Comparing the totals you get on those two searches should give you a good start at figuring out the relative frequency.
Announcements
- 13 August 2012: A new Nonfiction corpus of material from an Islamic Discourse web-site has been added.
- 2 July 2012: The newspapers mentioned below have been changed. Please see the Corpus Information for the updated word counts. The next section of AlGhad will appear on the corpus shortly. Let us know if there are any questions.
- 20 June 2012: The word count for some of the Newspaper corpora will change in the coming weeks. When we collected the texts originally, we were aware that there were some duplicate articles but we retained them because that is how they were archived on the Newspaper web-sites. We recently decided to remove as many of the duplicate articles as possible from these texts. If you have been using word counts for statistical research, your results may be altered slightly after the change. Therefore, we plan to wait to make the change until July 1, 2012. If you urgently need more time to finish research with the existing data or have any questions or concerns, please e-mail D. Parkinson (dil@byu.edu) immediately. Otherwise, we will proceed with the change as planned. The Newspapers that will be affected are Ahram 1999, Thawra, and AlGhad01. Another large section of AlGhad will also be added at that time.
- 22 February 2012: Over the next few weeks, arabiCorpus will be undergoing various changes. We are hoping to eliminate bugs and make the program more user-friendly. We will allow you to continue using the site as we make these changes, so be aware that things may change as you are performing searches. We apologize for the inconvenience but we are grateful for your patience as we try to make this program better for everyone.
- 20 February 2012: The list of all the texts found in the various corpora has been updated.
- 18 February 2012: Several premodern Arabic grammar texts have been added to the Premodern combined corpus and a Grammarians corpus has been created.
- 16 February 2012: A glitch causing an error in the frequency appearing on the Subsections page after performing a search has been resolved.
- 14 February 2012: Instructions have been edited and rewritten to reflect the changes that have taken place in the program. (This process is still ongoing).