Principles of the Transliteration (Add-On Version 3+)
Aramaic and Arabic are close relatives of Hebrew, and transliteration into the Hebrew writing system makes many words recognizable to a Hebrew speaker. The transliteration model is designed to make Arabic words easily recognizable to a Hebrew reader; pronunciation follows Literary Arabic.
For Arabic words, the converter reads letters together with their vowel marks and uses the word form and surrounding text where the spelling is ambiguous. It recognizes the definite article with attached conjunctions or prepositions, but also distinguishes a similar beginning that belongs to the word itself. A silent article lam before a sun letter may appear in square brackets, as in الشَّمْس (א[ל]שַّמס). When the article is clear, the converter can restore an unwritten shadda on the sun letter.
In Hebrew, the letters בגדכפת have, or historically had, two variants of pronunciation: plosive and fricative. Dagesh Qal denotes a plosive variant. The lesser-known overline called rafe may denote a fricative one. The converter uses both signs. It also uses geresh to mark a letter whose sound differs from its Hebrew pronunciation when both variants are plosive or both are fricative. Where possible, the converter does not use modern Hebrew geresh-based Arabic transliteration. In modern Israeli Hebrew, only כ and פ still have two pronunciations. In Liturgical Mizrahi Hebrew, all בכפת letters have two pronunciations; ג and ד retain two pronunciations in Yemeni Hebrew. The converter uses Liturgical Mizrahi and Yemeni pronunciations of בגדכפת to reflect specific Arabic sounds in transliterated text.
In Arabic, פ/ف always represents F, ב/ب always represents B, and כ/ك always represents K. However, د ط ح ص ت, the equivalents of דטחצת, each belong to a pair of sounds, like Hebrew בגדכפת. The variant whose pronunciation differs from the basic Hebrew or Aramaic pronunciation—ordinarily a fricative—is marked with an additional dot that mostly acts like rafe. Although ح has the same sound as Mizrahi Hebrew ח, while خ has the same sound as כֿ, the converter maps خ to ח׳. Earlier versions used the phonetically consistent כֿ, but خ is linked etymologically with ח, not כֿ. It is thought that two Proto-Semitic sounds that remain distinct in Arabic merged into one Hebrew sound and letter. Conversely, one Arabic or Proto-Semitic sound represented by ك is thought to correspond to two distinct Hebrew sounds.
With dagesh and without rafe, דכת in transliterated text denote the same sounds as in modern Israeli Hebrew; ט denotes the Arabic or Mizrahi Hebrew [tˤ]. However, צּ denotes the consonant ض DAD, the plosive variation of tsade, which sounds closer to D, while Hebrew צ always denotes an affricative. For example, أرض, the Arabic word for “earth,” is written with DAD and is transliterated by default as אְרץּ, making it easily recognizable. With rafe and without dagesh, the דטת letters denote the following sounds according to the description:
| דֿ/ذ | the “th” sound in the English word “this” |
| טֿ/ظ | the “th” sound in the English word “thus” |
| תֿ/ث | the “th” sound in the English word “thin” |
The letter غ GHAYN denotes an R-like sound, the voiced uvular fricative. Many Ashkenazi Jews pronounce ר this way. It would be phonetically consistent to transliterate غ as רֿ, but the converter maps it to ע׳ because غ is linked etymologically with ע, not ר. In converted Syriac text, softened gamal is rendered as גֿ; for example, ܒܓ݂ܕܕ → בגֿדד. A special dot “רוככא” marks a “soft” letter in the original text.
The letters ر/ר represent the same sound. In Arabic, it is a trill close to the Amharic or Russian R. This sound is very difficult for most Ashkenazi Jews to pronounce. During the Holocaust, local Nazi collaborators in Eastern Europe used the Russian word “KUKURUZA”/“кукуруза” as a shibboleth to identify Jews for extermination.
The letter and diacritic sign ء HAMZA denotes a glottal stop , while alif without HAMZA mostly denotes a long /aː/ vowel. In historical Hebrew, א necessarily represents a glottal stop when either the alef itself or the consonant immediately preceding it bears a shva or a hataf. A word-initial alef is also consonantal.
By default, the Add-On uses hatafs where possible to represent HAMZA together with a following short vowel. Hataf-patah is used for the corresponding short /a/, while hataf-qamatz can be used to represent short /u/ (Arabic damma). For HAMZA below alif with kasra, the default notation is the historically attested hataf-hiriq, written as shva + hiriq: إِ → אְִ. The user may select hataf-segol instead of hataf-hiriq if preferred.
If the corresponding hataf notation is disabled, the Add-On represents HAMZA separately by an alef with shva and places the vowel on a following alef where necessary. Thus, ordinary hiriq is used instead of hataf-hiriq for short /i/, and kubutz is used instead of hataf-qamatz for short /u/.
Hatafs can represent only short vowels. Long vowels following HAMZA are therefore written using the ordinary Hebrew notation for long vowels, regardless of the hataf settings. Long /iː/ is represented by hiriq male, and long /uː/ by shuruk. When a separate indication of the glottal stop is required, an alef with shva precedes the vowel-bearing letter; for example, medial /ʔiː/ can be written אְאִי.
The Treat word-initial א as always consonantal option, enabled by default, follows the historical Hebrew convention that an alef at the beginning of a word itself represents the consonantal glottal stop. The option therefore removes a separate word-initial אְ when its only purpose would be to mark HAMZA before another vowel-bearing alef. For example, אְאָ becomes אָ, אְאִי becomes אִי, and אְאוּ becomes אוּ. If the option is disabled, the explicit initial alef with shva is retained. This may be useful for modern Hebrew readers, since word-initial alef is often not pronounced as an audible glottal stop in Israeli Hebrew.
This option does not disable or simplify the hataf notation itself. When hatafs are enabled, a single word-initial alef carrying a hataf continues to be used normally. Thus, إِ is still represented as אְִ when hataf-hiriq is selected. The option only removes an otherwise redundant additional alef.
For example, the modern Arabic spelling إِسْرَائِيل is transliterated as אְִסרָ[י]אְאִיל. The Qur'anic spelling إِسْرَٰٓءِيلَ is transliterated as אְִסרָאְאִילַ. In both forms, the initial HAMZA with short /i/ is represented by hataf-hiriq, while the medial HAMZA before long /iː/ requires ordinary hiriq male in addition to the explicit representation of the glottal stop.
The Add-On also provides an Ancient Hijazi reading reconstruction mode, based on research into the early Qur'anic consonantal text and its historical pronunciation. In this mode, HAMZA is generally omitted in positions where it is reconstructed as absent from Ancient Hijazi pronunciation, and the corresponding historical vowel sequences or long vowels are reconstructed instead.
ى ʾalif maqṣūrah (ألف مقصورة) denotes a long /aː/ vowel at the end of a word in Literary Arabic. It looks like a dotless ـي. The converter marks the long vowel with kamatz and may show the silent source letter in square brackets, as in عَلَى → עַלָ[י]. Egyptians often omit the dots from final ي, making it look like alif maqsurah. The converter uses vowel marks, recognized word forms, and related forms in the same text to distinguish the two readings. When the evidence identifies an undotted final yeh, it renders Hebrew י; for example, فِى → פִֿי. Recovery is not applied when the reading remains ambiguous.
The shva in converted Arabic text should never be pronounced: like sukun (سكون), it denotes the absence of a vowel. Kamatz denotes long /aː/ in converted Arabic and Eastern Syriac texts, but /o/ in converted Western Syriac.