click on red text to expand and collapse information
General Information about arabiCorpus
  • What is a corpus?
    • A corpus is simply a large collection of words from a given language. These words can come from a variety of sources.
    • A person searches from a corpus by entering a search string, which is simply a string of letters or characters.
    • Corpora (plural of corpus) are often used by linguists to learn about some phenomenon of language.
      • Anyone interested in a particular language, though, can use a corpus and learn from it.
    • The benefit of a corpus is that it contains real language that was spoken or written by speakers of that language.
    • This allows the researcher the ability to see how the language is used in actual context instead of just in theoretical examples.
      • A student studying a foreign language can use a corpus to see how a particular word is most often used.
      • A language teacher can identify more frequent words and teach those to students earlier on.
      • A researcher can discover what spelling variations are used in different texts.
      • These are just a few of the countless ways that someone can use a corpus.
    • arabiCorpus is a program that allows you to search from a large corpus of Arabic words.
    • You can use the Tutorial and Instructions below to see the different uses and benefits of this program.
  • Important Things to Remember
    • arabiCorpus allows you to search large, untagged Arabic corpora.
      • 'Untagged' means that the words in the corpora have not been assigned to a particular a certain part of speech.
      • There are ways to find the part of speech you are looking for but there may still be errors. See 'Select a Part of Speech Filter' in the Detailed Search Instructions section of Performing Searches in arabiCorpus below for more information about this.
    • arabiCorpus searches for EXACTLY the string you type in (and nothing else).
      • In connection with the comment above, the part of speech (POS) filters can help in finding what you are looking for, but they will not necessarily give you the exact results you want.
      • The program simply searches for the string you type in with various possible attachments to that string. It does not actually know if the resulting word is a verb, noun, etc.
    • arabiCorpus accepts most regular expression language in search strings.
      • For more information on these types of searches, see the section on 'Searching with Regular Expressions' below.
    • You can easily search for individual words and you have the ability to search for multiple words at once.
    • You can search individual texts or combined corpora, which consist of multiple texts.
      • Only some of the individual texts are available in Basic Search but all of them are available in Advanced Search.
      • Avoid using the large combined corpora 'All' and 'All Newspapers' when searching for common words.
    • Some filtering of the results is available after the search.
      • There are various tools provided after a search has been performed to analyze the results and make them more meaningful.
      • See the section 'Analyzing the Results of a Search' below for more information.
    • Use the Tutorial below to help you with your initial searches and see the Instructions for additional details on using arabiCorpus.
  • Basic Corpus Information
    • arabiCorpus has corpora in five main categories or genres: Newspapers, Modern Literature, Nonfiction, Egyptian Colloquial, and Premodern.
    • You will select a specific corpus each time you perform a search in arabiCorpus.
      • The program gives you the ability to search combined corpora made up of multiple texts in a related genre.
      • You can also search any text individually by using the Advanced Search mode.
      • You can even search all of the texts at the same time. (Although this will often slow down your search).
        • Avoid using the large corpora 'All' and 'All Newspapers' when searching for common words because it may cause problems for future searches.
    • The names and word counts for the Basic Search corpora are found in this section. For a more detailed listing of the texts found in each corpus, see the 'Detailed Corpus Information' section.
      • Note that you cannot add these amounts together for the total number, since the grouped corpora group and regroup in various configurations.
    • The total number of words of the whole corpus is: 173,600,000.
      • All Newspapers: 135,360,804
      • Al-Masri Al-Yawm 2010: 13,880,826
      • Ahram 1999: 15,892,001
      • ShuruqColumns: 2,067,137
      • AlGhad01: 19,234,228
      • AlGhad02: 19,628,088
      • Hayat 1997: 19,473,315
      • Hayat 1996: 21,564,239
      • Tajdid 2002: 2,919,782
      • Watan 2002: 6,454,411
      • Thawra: 16,153,918
      • Modern Literature: 1,026,171
      • Nonfiction: 27,945,460
      • Islamic Discourse: 27,365,915
      • Other Nonfiction: 579,545
      • Egyptian Colloquial: 164,457
      • Premodern: 9,127,331
      • Adab Literature: 2,073,071
      • Grammarians: 1,210,614
      • Medieval Philosophy/Science: 1,576,860
      • Hadith Literature: 3,624,346
      • Quran: 84,532
      • 1001Nights: 557,908
Detailed Corpus Information
  • Newspapers
    • The bulk of the corpora come from newspapers and each newspaper corpus can be searched individually in Basic Search. The following list shows the years and locations for the current newspapers:
      • The 'All Newspapers' corpus has 135,360,804 words and consists of the following corpora:
        • Al-Masri Al-Yawm 2010 from Egypt – 13,880,826 words (المصري اليوم من مصر)
        • Al-Thawra from Syria – 16,153,918 words (الثورة من سوريا)
        • At-Tajdid 2002 from Morocco – 2,919,782 words (التجديد من المغرب)
        • Al-Watan 2002 from Kuwait – 6,454,411 words (الوطن من الكويت)
        • Al-Ghad(01) 2010-11 from Jordan – 19,234,228 words (الغد من الأردن)
        • Al-Ghad(02) 2010-11 from Jordan - 19,628,088 words (الغد من الأردن)
        • Al-Ahram 1999 from Egypt – 15,892,001 words (الأهرام من مصر)
        • Al-Hayat 1997 from London – 19,473,315 words (الحياة من لندن)
        • Al-Hayat 1996 from London – 21,564,239 words (الحياة من لندن)
        • Shuruq Columns Egypt – 2,067,137 words (الشروق من مصر)
      • Remember that the 'All Newspapers' corpus is very large so you should avoid using it when searching for common words.
  • Premodern
    • There are also various Premodern texts, two of which can be searched individually in Basic Search and others are available in combined corpora based on genre. Any individual text may be searched using Advanced Search. This list details the various corpora in this genre:
      • The total number of words in the 'Premodern' corpus is 9,127,331 words and all of the following texts mentioned in this section are part of it. There are also smaller combined corpora that are included in the Premodern corpus.
      • The following corpora can be searched individually in Basic Search and are also found in the combined 'Premodern' corpus:
        • Quran – 84,532 words (القرآن الكريم)
        • 1001 Nights – 557,908 words (كتاب ألف ليلة وليلة)
      • The 'Adab Literature' corpus, which has 2,073,071 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
        • Book of Songs by Abu Al-Faraj Al-Isfahani - 1,523,874 words (كتاب الأغاني لأبو الفرج الأصفهاني)
        • The Scrooges by Al-Jahiz - 48,241 words (البخلاء للجاحظ)
        • Letters of Al-Jahiz - 167,001 words (رسائل الجاحظ)
        • The Book of Animals by Al-Jahiz - 333,955 words (كتاب الحياون للجاحظ)
      • The 'Grammarians' corpus, which has 1,210,614 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
        • Ajarumiyya Grammar – 2,399 words (الكتاب للأجرومية)
        • Ibnjinni Grammar - 192,748 words (الخصائص لأبو الفتح عثمان بن جني الموصلي)
        • Jurjani Grammar - 77,653 words (أسرار البلاغة للجرجاني)
        • Mubarrad Grammar - 168,789 words (المقتضب في اللغة للمبرد)
        • Sayuti Grammar - 316,090 words (الكتاب: همع الهوامع فى شرح جمع الجوامع للإمام السيوطى)
        • Thalabi Grammar - 101,757 words (سحر البلاغة وسر البراعة للثعالبي and فقه اللغة وسر العربية للثعالبي)
        • Zamakhshari Grammar - 131,298 words (أساس البلاغة للزمخشري)
        • Sibawaihi – 219,880 words (الكتاب لأبو بشر عمرو بن قنبر الملقب بسيبويه)
      • The 'Medieval Philosophy and Science' corpus, which has 1,576,860 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
        • Ihya by Al-Ghazali - 992,341 words (إحياء علوم الدين للغزالي)
        • Incoherence of the Philosophers by Al-Ghazali – 47,307 words (تهافت الفلاسفة للغزالي)
        • Medical Aphorisms by Maimonides – 81,581 words (كتاب الفصول في الطبّ لموسى بن ميمون)
        • On Asthma by Maimonides – 17,904 words (مقالة في الربو لموسى بن ميمون)
        • Middle Commentary by Averroes – 33,004 words (تلخيص كتاب النفس لأرسطو لابن رشد)
        • Metaphysics by Avicenna – 90,760 words (الإلهيات من الشفاء لابن سينا)
        • Issues in Fiqh by Ibn Hanbal - 72,589 words (مسائل أحمد ين حنبل رواية ابنه عبد الله)
        • Prolegomenon by Ibn Khaldun - 241,374 words (مقدمة لابن خلدون)
      • The 'Hadith Literature' corpus, which has 3,624,346 words, is a subsection of the 'Premodern' corpus and consists of the following corpora:
        • Sunan by Al-Bukhari - 617,933 words (سنن للبخاري)
        • Sunan by Abi Daud - 429,227 words (سنن لأبي داود)
        • Sunan by Al-Tarmthi - 425,807 words (سنن للترمذي)
        • Sunan by Al-Darmi - 210,909 words (سنن للدارمي)
        • Sunan by Ibn Maja - 321,924 words (سنن ابن ماجة)
        • Sunan by Muslim - 746,829 words (صحيح مسلم لمسلم بن الحجاج)
        • Sunan by Al-Nasai - 871,717 words (السنن الكبرى للإمام النسائي)
  • Modern Literature
    • The next section of corpora comes from Modern Literature, primarily from novels. The countries represented with the novels are Egypt, Palestine, Algeria, Saudi Arabia, the Sudan, Syria, and Lebanon. Currently about half of the material in the Modern Literature corpus is from Egypt, about a fourth from Algeria, and lesser amounts from the rest of the countries. These are the novels found in the corpus as of Feb 2012:
      • The 'Modern Literature' corpus has 1,026,171 words and consists of the following texts:
        • ريم بسيوني: رائحة البحر with 38,485 words
        • ريم بسيوني: مدبولي with 40,042 words
        • إبراهيم عبد المجيد: لا أحد ينام في الاسكندرية with 98,580 words
        • علاء الأسواني: عمارة يعقوبيان with 55,191 words
        • علاء الأسواني: شيكاجو with 77,026 words
        • خالد الخميسي: تاكسي with 28,391 words
        • نجيب محفوظ: ميرامار with 35,579 words
        • نجيب محفوظ: الكرنك with 13,742 words
        • نجيب محفوظ: صدى النسيان with 6,276 words
        • نجيب محفوظ: أصداء السيرة الذاتية with 11,542 words
        • نجيب محفوظ: أولاد حارتنا with 104,684 words
        • أحلام مستغانمي: ذاكرة الجسد with 71,065 words
        • أحلام مستغانمي: عابر سرير with 53,925 words
        • أحلام مستغانمي: فوضى الحواس with 57,418 words
        • رجاء عبدالله الصانع: بنات الرياض with 56,173 words
        • الطاهر وطار: الولي الطاهر يعود إلى مقامه الزكي with 19,271 words
        • الطاهر وطار: الولي الطاهر يرفع يديه بالدعاء with 19,957 words
        • التطاهر وطار: الحوات والقصر with 26,756 words
        • الطيب صالح: عرس الزين with 16,283 words
        • إدوار الخرات: ترابها زعفران with 20,475 words
        • لطيفة الزيات: الشيخوخة وقصص أخرى with 19,556 words
        • يحيى حقي: قصص ليحيى حقي with 9,409 words
        • إلياس خوري: مملكة الغرباء with 19,790 words
        • غسان كنفاني: أم سعد with 6,623 words
        • غسان كنفاني: عائد إلى حيفا with 11,326 words
        • غسان كنفاني: مسرحية الباب with 9,263 words
        • نجاة حالو: سر الحياة with 11,211 words
        • سعد الله ونوس: مغامرة رأس المملوك جابر with 17,659 words
        • تميم صائب: لا تفقأ عينيك يا أوديب with 7,607 words
        • غادة السمان: ختم َلذاكرة بالشمع الأحمر with 4,070 words
        • علي سالم: أولادنا في لندن with 16,865 words
        • أمجد ناصر: حيث لا تسقط الأمطار with 41,931 words
  • Egyptian Colloquial
    • There is a small combined corpus with some Egyptian Colloquial data.
      • The 'Egyptian Colloquial' corpus has 164,457 words and consists of the following corpora:
        • The Egyptian play Awaladna fi Landan – 16,865 words
        • Material from the Egypt Chat website – 140,234 words
        • An interview with Hosni Mubarak – 7,358 words
      • It is important to note that these texts have some colloquial and a lot of Fusha and mixed Fusha and colloquial. The literature and the newspapers also contain some colloquial, as that is the nature of Arabic.
  • Nonfiction
    • The last group of corpora is in the Nonfiction category, which includes a large amount of material from an Islamic Discourse web-site along with some literary criticism, other scholarly (and not so scholarly) works, some political speeches, and some official UN and other diplomatic documents. These are the corpora found in this genre:
      • The 'Nonfiction' corpus has 27,945,460 words and contains all of the following texts, which are also divided into smaller subsections:
      • The 'Islamic Discourse' corpus, which has 27,365,915 words, is a subsection of the 'Nonfiction' corpus and contains the following corpus:
        • Material from Sayd.net (an Islamic Discourse web-site) - 27,365,915 words (صيد الفوائد)
      • The 'Other Nonfiction' corpus, which has 579,545 words, is a subsection of the 'Nonfiction' corpus and contains the following texts:
        • Landed from the Sky by Anis Mansour – 40,322 words (الذين هبطوا من السماء لأنيس منصور)
        • The Absent Truth by Farag Fouda – 33,234 words (الحقيقة الغائبة لفرج فودة)
        • We March Forward edited by Dr. Hala Esbanyuli – 11,113 words (نسير إلى الأمام لدكتور هالة اسبانيولي)
        • Myth of Delusion by Muhammad Khalil Al-Hakaymah – 46,959 words (أسطورة الوهم لمحمد خليل الحكايمه)
        • Milestones by Sayyid Qutb – 38,842 words (معالم في الطريق لسيد قطب)
        • Black Book of Capitalism translated by Dr. Anton Hamsy – 145,362 words (الكتاب الأسود للرأسمالية ترجم الدكتور أنطون حمصي)
        • American Time from New York to Kabul by Muhammad Hasanayn Haykal – 73,778 words (الزمن الأمريكي من نيويورك إلى كابول لمحمد حسنين هيكل)
        • Al-Naba’ Magazine – 5,204 words (مجلة النبأ)
        • UN Resolution 1636 – 1,870 words (نص القرار 1636 للأمم المتحدة)
        • Media Arabic in Egypt by Farouq Shousha – 3,990 words (اللغة العربية في الإذاعة والتلفاز والفضائيات في جمهورية مصر العربية للاستاذ فاروق شوشة)
        • Sharm Al-Sheikh Resolution 1999 – 1,852 words (نص مذكرة شرم الشيخ من أيلول 1999)
        • Sana Speech by Hosni Mubarak – 1,110 words (نص خطاب السيد رئيس الجمهورية في قمة تجمع صنعاء اديس ابابا)
        • Dramatic Structure by Sabah Al-Anbari - 23,976 words (البناء الدرامي لصباح الأنباري)
        • Dialogue with an Atheist by Mustafa Mahmoud – 12,068 words (حوار مع صديقي الملحد للدكتور مصطفى محمود)
        • Yusuf Idris Literary Criticism by Dr. Abir Salama – 54,549 words (نداهة الكتابة: نصوص مجهولة فى إبداء يوسف إدريس للدكتور عبير سلامة)
        • In Pre-Islamic Poetry by Taha Hussein - 35,199 words (في الشعر الجاهلي لطه حسين)
        • Political Necessity Volume 5 - 15,178 words (الضرورة السياسية)
        • Place Aesthetics in Contemporary Arab Literary Criticism by Abdullah Abu Heif - 6,090 words (جماليات المكان في النقد الأدبي العربي المعاصر لعبد الله أبو هيف)
        • Letters from Kanafani to Ghada - 18,253 words (رسائل غسان كنفاني إلى غادة السمان)
        • Interview with Muammar Al-Gaddafi - 2,714 words (نص حديث الأخ قائد لاثورة لمعمر القذافي)
        • New Military Service Law - 524 words (نص قانون خدمة العلم الجديد)
        • Interview with Hosni Mubarak – 7,358 words (نص مقابلة الرئيس حسني مبارك مع قناة العربية)
  • All
    • The 'All' corpus consists of all the texts mentioned and has a total of 173,600,000 words.
    • Remember that searching in the larger corpora such as 'All' or 'All Newspapers' will lead to much longer wait times.
      • If you combine a larger corpus with a search that will yield hundreds of thousands of results, the machine may balk and not return anything.
        • You should only use the 'All' corpus to search for rare or unusual forms. Searching for common words can lead to inaccurate results and may cause problems in subsequent searches.
      • In general, the 'All' category is mainly appropriate for researchers looking for broad quantitative or comparative data.
      • The smaller categories are usually going to be more appropriate for pedagogical data, or data for students, since it will come in a quantity that you can understand and deal with.
Performing Searches in arabiCorpus
Analyzing the Results of a Search
  • Basic Information about Results
    • The results come back after waiting for a few seconds.
      • If you search for a single noun or adjective in a single corpus, it should come back in about 10 seconds
      • If you search for multiple nouns, or verbs, or search in the combined corpora, the results will take longer.
    • If you search for a very common word that produces tens of thousands of hits or more, the program will sometimes balk because the database it is using will limit what can be inserted.
      • If the search takes too long and returns no results, the best solution is to log out of the program and log back in. You may have to wait a few minutes before the program returns back to normal.
    • The results of a search are saved temporarily in a database, and you can access them by clicking on the words that appear in the red bar.
  • Summary Page
    • This page comes up first automatically after a search, and you can return to it by clicking on 'summary' in the red bar.
    • This page gives you summary information about your search:
      • The search string you typed in, in both English and Arabic scripts
      • The search strings actually used by the program after its 'figuring'
      • The corpus you searched
      • The time it took the search engine to do the basic search (this will generally be less than the actual time experienced by you, since it doesn't include the time it takes for the server to receive your request or to serve the results back to you)
      • The POS filter you chose
      • The POS filter that was actually used
      • The number of 'hits'
      • What that number translates to in terms of words per 100,000 words of your corpus
    • The latter bit of information is useful for comparative purposes. The corpora are of vastly different sizes, so the actual number per corpus may be misleading, but the number per 100,000 words can be compared.
    • This page also displays various alerts according to the search terms you entered. Usually, these alerts will appear when you have used the 'verb' POS filter and will inform you of any guessing the program has performed and how you might need to change your search.
      • It is important to read these alerts so you will know what the program is searching for and how to change your search if you do not find the results you are looking for.
  • Citations Page
    • Click on 'citations' in the red bar to see the 'hits' with the 10 words before and the 10 words after.
    • By default, these citations are sorted by the word that appears directly before the word you searched for.
      • Note that this sorting is done by the whole word, not by the root, so AlktAb الكتاب, ktAb كتاب, and wAlktAb والكتاب are nowhere near each other.
      • The word directly before is repeated at the left hand side of the page so you can quickly glance through the citations and notice patterns, collocations, etc.
    • If you would rather see the citations sorted by the word directly after the 'hit', click on the red sentence near the top of the page that says 'sort by word after.'
    • The citations are shown 100 at a time.
      • If your search returned more than 100 citations, they will be organized into 'pages' which you can access by clicking on the red numbers at the top.
    • The far right column on the Citations page shows you the subsections of the corpus you searched.
      • If you searched a single corpus, the subsection column will indicate what subsection of the corpus the example comes from.
        • The newspapers have self-labeled subsections, so sometimes you need to figure out what the codes mean, but they are related to the various sections of the newspaper.
      • If you have searched a combined corpus, the subsection column will tell you which specific corpus the example came from.
      • If you want to see the exact reference of your example, click on the red 'subsection' heading and it will change to 'reference.'
        • The references in many of the corpora are basically incomprehensible and not very helpful, but the references to some, like the Quran and the novels, are useful.
          • The references in the Quran are to Sura and verse.
          • The references in the novel are to chapter, section of chapter (divided by stars), and paragraph number.
          • The references in 1001 Nights are to Nights and paragraphs (but in an odd kind of way, which I'll explain if you ask me by e-mail).
    • If you want to see more context than the 10 words before and the 10 words after, click on the number at the far left of the citation.
      • This will bring up a separate window that will display the whole verse from the Quran, the whole paragraph from the novel and 1001 Nights, or the whole article from the newspapers.
      • You may have to use your browser's Arabic script search function to find the place in the larger context where your item appears, since it does not highlight it.
  • Subsections Page
    • Click on 'subsections' in the red bar to see the totals for the various subsections of the corpus you searched.
      • Again, if you searched a single corpus, those subsections will be those defined for that corpus.
      • Some corpora, particularly the non-news ones, have no subsections defined.
      • If you searched a combined corpus, the subsections will be the names of the corpora in that combined corpus.
    • The results in this section are ordered from most frequent to less frequent in terms of number.
    • The frequency per 100,000 for each subsection is also given.
    • This page allows you to compare, say, Premodern Arabic with Modern, or to compare newspapers from various places.
      • Type lA bd لا بد and choose All Newspapers, for example, click on subsections, and you will find that the Ahram uses it much less frequently than the Hayat.
      • Type lAbd لابد, though, and choose All Newspapers and you will find the opposite (the Ahram apparently typically types the phrase without a space, and the Hayat does the opposite).
  • Word Forms Page
    • Click on 'word forms' in the red bar to see the exact forms that your search produced, orderd by frequency.
    • Examining this list can give you important hints about normal usage.
    • Examining this list can also help you see easily what kinds of false hits you are getting so you can work to eliminate them.
    • It is strongly suggested that you examine this list for every search before you take the rest of the results seriously.
      • See if there are any forms you expected to find that aren't there.
      • See if there are any forms you did not expect to find that are there.
      • See if you can figure out what it is about your search expression (and Arabic morphology) that would have created these problems.
    • Click on any word in the word form list and a window will open up with the citations for just that word form.
  • Words Before/After Page
    • Click on 'words before/after' in the red bar to see a list of the common words directly before and directly after the 'hit' ordered by frequency.
    • Examining these lists is a good way to scope out the main usages and collocations of the word you searched for.
    • Examining the lists can also help you identify structures that you had not intended to search for that you want to cut out.
    • If you want to see only the citations with that particular word before or after, click on a word in the list.
      • A separate window will open showing you just those citations.
  • Collocates Page
    • Click on 'collocates' in the red bar to see a list of the most common collocates that appear with your search term.
      • For this program, collocates are defined as the most frequent word forms that appear up to four positions to the left or right of your search term.
      • The program only lists those collocates that appear at least 4 times in the given corpus.
    • You cannot click on the word to see the occurrences in context like with other results pages.
    • Like the Words Before/After Page, this page allows you to see what words occur immediately before or after your search term, but it also allows you to see those words that commonly appear near your search term and not just right before or after.
Tutorial
  • First Search: Look for a single noun
    • Click on the red Instructions above the submit button to get these instructions in a separate window (otherwise they will go away as you follow them).
    • Perform a simple search for a noun.
      • Type mktb into the leftmost box. Do NOT type any vowels (i.e. do NOT type maktab). Do NOT type the definite article (do NOT type Almktb).
      • Choose 'noun' from the POS list.
      • Choose Ahram 1999 from the corpus list.
      • Click on 'Submit'.
      • Wait about 10 seconds.
    • Analyze the results of your search.
      • Examine summary of search that appears, and notice how many examples it found, and how frequent these are in the corpus (per 100,000 words).
      • Click on 'citations' in the dark red bar.
        • Scroll down and look at a few of the citations. Scroll back to the top.
        • Note that there are about 40 pages of results. Click on page 25.
        • Note that each example gives you the word in context with 10 words before and 10 after.
        • Note that the examples are organized by the word that comes before (here fy), and that this word is also listed at the beginning of each line so it can be easily picked out.
      • Click on 'sort by word after' to sort the examples by the word after instead.
        • Click on page 25 again, and see what kind of information you get with this order.
        • Click on one of the numbers at the left, and see a new window open with even more context. Close that window
      • Click on 'subsections' in the dark red bar.
        • Examine the frequencies and relative frequencies of this word in the various sections of the Ahram.
      • Click on 'word forms' in the dark red bar.
        • Notice the different forms in which this word was found: alone, with wa-, bi-, fa-, with the definite article, and with various pronoun endings. Notice which of these forms were more common and which relatively rare.
        • Click on مكتبك in the middle of the second column to see the citations just for that word form in a separate window. Examine briefly and close the second window.
      • Click on 'words before/after' in the dark red bar.
        • Examine the most common words that come before our search word and the most common words that come after. Any surprises? Is this what you would have predicted?
        • Click on التنسيق in the second column to see the 200 examples of this word coming after our search word in a new window. After examining for a few moments, close the new window.
      • Click on 'summary' in the dark red bar to go back to the summary page.
  • Try other single words
    • Type jmyl into the leftmost box, choose the 'adj' POS filter, the Ahram 1999 corpus, and click submit.
      • Go through the various choices (as under First Search) and see the differences. Note particularly under 'word forms' that you get examples with and without the feminine ending, but no examples with pronoun endings or prepositions.
      • Now try running the same word on the same corpus but with the 'noun' POS filter chosen. Note under 'word forms' that you don't get the feminine forms, but you do get forms with prepositions and pronouns.
    • Type tqrybA into the leftmost box, choose the 'adv' POS filter and the Ahram 1999 corpus and run through the various options once you get the results.
    • Type ktAbhm into the leftmost box, choose the 'adv' POS filter and the Ahram 1999 corpus, and look to see what you got under word forms.
      • 'adv' is a good POS filter when you are looking for a very specific form, and don't want to see other possible prefixes and suffixes.
    • Type prb into the leftmost box, choose the 'verb' POS filter, the Ahram 1999 corpus, and run through the various options looking at the results. Note particularly the large number of forms under 'word forms'.
    • Now run prb again, but this time choose the 'string' POS filter.
      • Note under word forms that besides all the verb forms you got a moment ago, you are now getting a bunch of other things that happen to contain the sequence prb, including a bunch of typos.
  • Account for variabilty in the corpus
    • There are several ways to have the corpus look for alternates (either this OR that). The easiest involves square brackets.
    • The square brackets tell the program to search for any one of the characters in the brackets.
    • Type lbnAne into the leftmost box, choose the 'noun' POS filter, and the Ahram 1999 corpus.
      • Notice that you get 20 hits that end with an alif maqsura. If you had searched for lbnAny and wanted to see all examples of use of this word, the program would have missed these 20.
      • In general, if you want ALL examples of forms that end in y, you would need to use [ey] instead.
      • The Ahram itself is a special case because for whatever reason they use the yaa' for the alif maqsuura much of the time. So if you search for mqhe (with 'noun') in the Ahram 1999 corpus you get zero results, but if you search for mqhy you get 158, and examining their citations reveals that they are meant to be mqhe.
      • Therefore, if the Ahram is included in your search, you should use [ye] for words that end in y AND e.
    • Now type msWwl into the leftmost box, choose the 'adj' POS filter and the Newspaper combined corpus. You will have to wait longer for the results because of the larger corpus.
      • If you check out the subsections page, you will see that there are a huge number of hits in the other papers, but very few in the Ahram.
      • Now type msYwl into the leftmost box, choose the 'adj' POS filter, and the Newspaper combined corpus.
      • Notice in the subsections page that almost all the examples are from the Ahram, and very few from the others.
      • If you wanted to catch ALL instances of this word in the Newspaper combined corpus, you would need to type ms[WY]wl.
    • Now type tlfwn into the leftmost box, choose the 'noun' POS filter, and the Ahram 1999 corpus. Look how few the results are.
    • Now type tlyfwn into the leftmost box, choose the 'noun' POS filter, and the Ahram 1999 corpus. Look how many the results are.
      • If you want to find tlfwn and tlyfwn together you could type tly?fwn. The question mark tells the program the yaa' can either be there or not.
    • The program is really just an automoton. It doesn't know what you want. It just looks for exactly what you type in. If you need it to pay attention to spelling variation, then you must tell it to do so.
    • Moral: PAY ATTENTION to spelling variations in Arabic, and account for them in your searches, often by using square brackets or question marks.
  • Search for more than one word at once
    • Type mktb,mkAtb into the leftmost box. Note that there is no space after the comma. Choose the 'noun' POS filter and the Ahram 1999 corpus.
      • Notice on the word forms page that the two words have been combined into a single search. You can string any number together with commas.
    • Now try the following combinations:
      • hjwm,hjmAt,hjwmAt,hjmQ,mhAjmQ (check out the word forms page to get an idea of relative frequency)
      • bTbycQ AlHAl,TbcA,bAlTbc (ditto)
      • hAtf,tlfwn,tlyfwn,mwbAyl,mHmwl,xlwy
  • Filter results 'by hand'
    • Type lbnAn into the leftmost box, choose the 'noun' POS filter, and Ahram 1999 corpus.
      • Examine the word forms. Notice that you have quite a few examples of lbnAny, which the program got by putting the y suffix for 'my' on the end of the noun.
      • However, if you click on this word on the word forms page and examine these citations, you will see that none of them mean 'my Lebanon'. The y here is a nisba adjective suffix.
      • In other words, you have gotten a bunch of things you don't want because of the morphological ambiguity of Arabic.
    • Let's say you want to look through all the citations with lbnAn, and you don't want to have to deal with all those lbnAny's.
      • Type lbnAn -- y$ into the leftmost box, choose the 'noun' POS filter and Ahram 1999 corpus.
      • The two dashes tell the program that whatever you type after them is NOT wanted (filter it out).
      • The dollar sign means that this is at the end of the word.
      • So this means, search for all examples of lbnAn with normal noun suffixes, but if you find one the ends in a y, filter it out.
      • A look at the word forms list will convince you that it has now gotten rid of the unwanted lbnAny's.
    • Try the following searches with and without the dash section and check the word forms page to see what it cuts out:
      • mktb -- ^m (this will cut out all word forms that begin with a miim, i.e. all that don't have some other prefix)
      • mktb -- h (this one will cut out all word forms that have a haa' in them, meaning that mktb with certain pronoun endings will be cut out.
      • mktb -- h|y|n (this one will cut out all word forms with a haa', kaaf, yaa', or nuun, meaning that maktb with all but second person pronoun endings will be filtered out. Note that the vertical bar represents alternation.
      • mktb -- h|y|n|km?A?$ (this one also cuts out the second person forms, but you have to be careful: if you just list kaaf by itself it will cut out everything, since kaaf is in the word itself. So you need the question marks and the dollar sign.
    • A good strategy is to perform plain searches first, and then examine the word form list to see if there are any forms you would like to filter out by hand. If there are, create a 'dash' phrase to do the job.
    • Note for those with a 'regular expression' background: You can consider the list after the dashes to be a regular expression for which the program automatically provides an outside set of grouping parentheses.
  • Search for a phrase
    • Type lA gbAr into the leftmost box, choose the 'adv' POS filter, and Hayat 1997.
      • Check out some of the citations.
    • Try these other phrases:
      • Type (Al)?bnyQ (Al)?tHtyQ into the leftmost box, choose the 'string' POS filter and Hayat 1997.
      • Type pyk bdwn rSyd into the leftmost box, choose the 'string' POS filter and Hayat 1997 (one hit only, check out the citation)
      • Type ttHlb lh Al[LA]fwAh into the leftmost box, choose the 'string' POS filter and All (one hit only, check out the citation)
      • Type \bwzArQ Al\w+ into the leftmost box, choose the 'string' POS filter and ShuruqColumns (to see a list of different ministries)
      • Try AlwzArQ Al\w+ (obviously not as useful)
      • Type mA lbV [LA]n into the leftmost box, choose the 'string' POS filter and ShuruqColumns.
      • Try (mA lbV|lm ylbV) [LA]n with the 'string' POS filter and Hayat 1997 (be prepared to wait)
    • You are basically limited only by your imagination. However, if you want to be sure to get the results you want, remember to account for hamza variations and other spelling variations.
    • Also remember that some writers do not put spaces where others do (some write lAbd and others lA bd). You can capture this with lA ?bd (try it with ShuruqColumns).
  • Compare results of different corpora
    • Type hAtf into the leftmost box, choose the 'noun' POS filter and All Newspapers.
      • Go to subsections and see what papers use it more commonly than others.
    • Now type tly?fwn into the leftmost box, choose the 'noun' POS filter and All Newspapers.
      • Compare the subsection results with the ones you got for hAtif.
    • Type lqd into the leftmost box, choose the 'adv' POS filter and All.
      • Note in subsections that the Quran and the novel have a much higher rate of use than the other corpora, and that the Ahram uses it more frequently that the Hayat.
    • Type shl into the leftmost box, choose the 'adj' POS filter and Hayat 1997.
      • Look in subsections and see which parts of the paper like this adjective more than others.
      • Then try it with Ahram 1999 and see if the same pattern holds. Look at the words before words after page to see if you get any hints as to why this may be so.
    • Try the same thing with hjwm (with 'noun') and see if the subsection patterns make sense to you.
Miscellaneous Info
  • Recommended Browser
    • You should be able to use this program with any of Google Chrome, Firefox, or Safari.
    • Internet Explorer is the only browser that seems to cause some problems, so it is recommended to use one of the others mentioned.
  • Note About Numbers
    • Because of how the numbers were originally entered into many of the texts making up the corpora, some of the numbers appear reversed or backwards after performing a search (i.e. 1999 may actually appear as 9991). We were unable to fix this problem, so you will simply have to accept that the digits of numbers might be reversed in any particular case.
    • Therefore, we do not recommend that you use this program to search for numbers.
  • Warning
    • This tool can easily become annoying if you are searching for a common word in a large corpus (you will have to wait a long time or the program may do nothing at all).
      • Only use the large corpora 'All' and 'All Newspapers' when searching for rare or unusual words.
    • When learning to use the tool, try a small corpus and search for slightly less common words.
      • You may want to switch to Advanced Search mode and select an individual text from the corpus list.
  • Common Errors to Avoid
    • Typing a noun or adjective with the definite article (the tool automatically looks for nouns and adjectives with and without the article; if you type it with the article, it will only find examples with it).
    • Choosing the wrong POS filter (looking for an adjective with 'noun').
    • Choosing 'noun' when looking for a phrase (you should usually choose 'adv' or 'string').
    • Searching for a common word in the 'All' or 'All Newspapers' combined corpora.
  • Responses to Common Questions/Problems
    • My results seem to be showing a mix of different searches. What is wrong?
      • You have probably just performed a search that took a very long time or appeared to time out.
        • If you used the large combined corpora 'All' or 'All Newspapers' to search for a common word, the search may have appeared to fail because it took so long. If you perform another search in this time, both searches will be running simultaneously and your results may show both searches.
        • Avoid using these large corpora when searching for common words. Limit searches in these corpora to rare or unusual forms.
      • The best way to manage this problem is to log out and wait a few minutes until performing another search.
        • It is a good idea to perform a few simple searches with few results first to check when the search results have returned to normal.
      • It is also possible that you were logged in as 'guest' at the same time someone else was using the site as 'guest.' This can cause the same issue.
        • This is easy to solve! Simply register with your e-mail address and this problem will occur far less often. It is free and easy and will be much more beneficial for you.
    • Why do I want to filter my results?
      • If you are trying to find examples of a particular word or construction, and the program finds hundreds of results of something else that happens to match what you are looking for (because of the morphological ambiguity of Arabic), it can be very time consuming and annoying to search through the citations one by one looking for the ones you really want. If you can figure out a way to filter out the 'bad' ones, you can save yourself many hours.
    • You want to take care of filtering yourself, and bypass the POS filters.
      • Choose 'string' and type a regular expression that matches what the POS filters do, or which varies them.
        • \bgbAr\b will find ONLY the word gbAr with no prefixes or suffixes.
        • \b[wf]?gbAr\b allows gbAr, wgbAr and fgbAr and nothing else.
        • \b[wf]?(Al)?gbAr\b allows gbAr, wgbAr and fgbAr with and without the definite article.
        • \b[wf]?(Al)?gbAr(h|hA|k|y)?\b allows all of the above and the singular pronoun endings.
        • \b[ytnLA]ktb[wyA]?[nA]?\b is one way to find present tense verb forms, here without any other prefixes or suffixes.
      • The program can deal with quite complex expressions efficiently, as long as parentheses are matched, so you can get quite a bit of control over what you are searching for if you want it.
    • You want to find all examples of the verb twj, but looking at the results you realize that you have many instances of twjh, which could conceivably mean 'he crowned him' but which 99% of the time is in fact another verb: tawajjaha 'to head for'.
      • Choose 'verb' and search for twj -- h. This will delete all the twjh's. You may miss a couple of 'he crowned him's but it's worth it because you can now see that your results are mainly what you want and expected.
    • You want to find all examples of the verb qAl, including passives.
      • In Advanced Search, searching for qAl,ql,qwl,ql with 'verb4' will give you all the actives. Searching for qyl,ql,qAl,ql will give you all the passives. Searching for q[Ay]l,ql,q[wA]l,ql will give you both all sorted together. Try this on a smaller corpus like 1001 Nights first, since it generates a huge number of hits.
    • You want to find all examples of the verb Lyd ('to support).
      • In Basic Search, the program will often guess incorrectly when you search for this verb, thinking it is supposed to be a hollow verb.
      • In Advanced Search, searching for Lyd using 'verb' will miss all the imperfects and all the perfects where the hamza was not typed on the alif.
      • Searching for [LA]yd will get the hamza/alif variation, but will still miss the imperfects. To get those you need to choose 'verb2' instead and type: [LA]yd,Wyd i.e. the perfect stem and the imperfect stem with a comma in between. This should give you what you are looking for.
    • You want to get comparative statistics on use of the verb forms yumkinu and tumkinu when they have a feminine verbal noun subject, but when you search for tmkn (using 'adv') you get overwhelmed with things that are really tamakkana, tamakkun, tumakkinu and the like.
      • This one is a good example of the inherent ambiguity of much of Arabic morphology, particularly graphemically. To get a statistic you can rely on, particularly if you don't have the stomach to actually read through hundreds of citations, you probably need to do what I call a 'limited search.' Instead of looking for all instances of tmkn, do a search for [yt]mkn \w+(Q|t(h|hA|k|hm|hmA|nA)), which will give you a list of all the words that end in a taa' marbuuta or a taa' followed by a pronoun ending after one of these verbs, and look at the list of words before and after. Look down the list of words after and pick out (say) the top 25 (or 10 or 5) feminine verbal nouns that come directly after ymkn or tmkn, and make sure they 'feel' right (i.e. you are pretty sure that the nouns you are choosing really are likely to function as the subjects of these verbs).
      • Then do a search for those forms only, once with ymkn before, and once with tmkn. You don't have an overall statistic, but you have a 'subset' statistic that you can trust. Sample search strings with 5 verbs I found would be:
        • ymkn (AstfAd|tsmy|syTr|zyAd|mwAjh)(Q|t(h|hA|k|y|hm|hn|km|kn|nA|hmA|kmA))
        • tmkn (AstfAd|tsmy|syTr|zyAd|mwAjh)(Q|t(h|hA|k|y|hm|hn|km|kn|nA|hmA|kmA))
      • Comparing the totals you get on those two searches should give you a good start at figuring out the relative frequency.
Announcements
  • 13 August 2012: A new Nonfiction corpus of material from an Islamic Discourse web-site has been added.
  • 2 July 2012: The newspapers mentioned below have been changed. Please see the Corpus Information for the updated word counts. The next section of AlGhad will appear on the corpus shortly. Let us know if there are any questions.
  • 20 June 2012: The word count for some of the Newspaper corpora will change in the coming weeks. When we collected the texts originally, we were aware that there were some duplicate articles but we retained them because that is how they were archived on the Newspaper web-sites. We recently decided to remove as many of the duplicate articles as possible from these texts. If you have been using word counts for statistical research, your results may be altered slightly after the change. Therefore, we plan to wait to make the change until July 1, 2012. If you urgently need more time to finish research with the existing data or have any questions or concerns, please e-mail D. Parkinson (dil@byu.edu) immediately. Otherwise, we will proceed with the change as planned. The Newspapers that will be affected are Ahram 1999, Thawra, and AlGhad01. Another large section of AlGhad will also be added at that time.
  • 22 February 2012: Over the next few weeks, arabiCorpus will be undergoing various changes. We are hoping to eliminate bugs and make the program more user-friendly. We will allow you to continue using the site as we make these changes, so be aware that things may change as you are performing searches. We apologize for the inconvenience but we are grateful for your patience as we try to make this program better for everyone.
  • 20 February 2012: The list of all the texts found in the various corpora has been updated.
  • 18 February 2012: Several premodern Arabic grammar texts have been added to the Premodern combined corpus and a Grammarians corpus has been created.
  • 16 February 2012: A glitch causing an error in the frequency appearing on the Subsections page after performing a search has been resolved.
  • 14 February 2012: Instructions have been edited and rewritten to reflect the changes that have taken place in the program. (This process is still ongoing).