Skip to content

00: Sinhala orthography: consolidated rule set ​

This file merges the topic studies (01–06) into one deduplicated rule list. Each rule cites the source rules it was built from (e.g. 02:VS-007 = file 02, rule VS-007). Compiled October 2026.

  • HARD = encoding or orthographic invariant. A correct text never violates it.
  • SOFT = spelling tendency. Use it to rank candidate spellings, never as a hard filter.
  • STYLE = two or more spellings are correct; the choice is a convention.
  • Confidence: H high · M medium · L low (same meaning as in the topic files).

Machine-readable companions:


1. Encoding invariants (HARD) ​

IDRuleConfSources
G-EN-01Store in logical order: consonant (or whole cluster) first, then the vowel sign, even for ෙ ේ ෛ ො ෝ ෞ, which display before the consonantH02:VS-002, 03:HC-023
G-EN-02Use precomposed signs: ේ U+0DDA, ො U+0DDC, ෝ U+0DDD, ෞ U+0DDE. Never their decomposed pieces. NFC-normalize text as a safety netH01:INV-003, 02:VS-003
G-EN-03ෛ U+0DDB has no decomposition. ෙ+ෙ is a different, silently-wrong string. Always encode U+0DDBH01:INV-003, 02:VS-004
G-EN-04Order inside ෝ is ෙ ා ්. The sequence ෙ ් ා normalizes to ේ + ා, which is wrongH02:VS-005
G-EN-05Independent vowels are atomic. Never encode අ+ා (ආ), අ+ැ, අ+ෑ, එ+් (ඒ), ඔ+් (ඕ), ඔ+ෟ, උ+ෟ, ඍ+ෘ, ඏ+ෟ, එ+ෙH01:INV-002, 02:VS-006
G-EN-06Consonants are atomic, including the 5 sanyaka letters, ඥ (U+0DA5, never ජ්‍ඤ) and ෆH01:INV-001, 03:HC-051
G-EN-07Hal alone never joins. C ් C always shows a visible hal. A joined form needs ZWJH03:HC-001
G-EN-08ZWJ position chooses the style: C ් ZWJ C = conjunct / yansaya / rakaransaya / repaya; C ZWJ ් C = touching letters (Pali). Unicode and SLS 1134:2011 agree; the 2004 SLS draft differsH03:HC-002/003, 01:INV-005
G-EN-09Named sequences: yansaya 0DCA 200D 0DBA · rakaransaya 0DCA 200D 0DBB · repaya 0DBB 0DCA 200DH01:INV-006, 03:HC-004
G-EN-10No ZWNJ in normal textH03:HC-005
G-EN-11Contextual glyph shapes (hook-u, රු රූ රැ රෑ ළු, the tail lost on ද before u, a second hal shape) are font matters. Encoding stays ordinary consonant + signH01:INV-009/010, 02:VS-011…017
G-EN-12Expect ZWJ-less text in real corpora (ශ්රී, අක්ෂර) and fold it for matching. Never strip ZWJ when producing textH03:HC-050, 03:§10.12, 05:SP-011

2. Vowels and vowel signs ​

IDRuleKindConfSources
G-VS-0117 dependent signs + al-lakuna. Each consonant has 19 forms: hal, inherent a, and 17 signsHARDH01:§2d, 02:VS-001
G-VS-02An independent vowel letter is used word-initially. Inside a word, a vowel after a consonant is always a signHARDH02:VS-029, 05:PH-009
G-VS-03No vowel sign after an independent vowel (ඉී, ඔා). Shapers usually do not warnHARDH02:VS-007
G-VS-04No vowel sign after hal (ක්ා), and at most one vowel sign per consonant (කාා). Only ං/ඃ may follow a signHARDH02:VS-009/010
G-VS-05A sign with no base is invalid in textHARDH02:VS-008
G-VS-06Vowel hiatus is repaired with a glide: ය after i/ii/e/ee/ae/aee, ව after u/uu/o/oo/aa (රැය, තොප්පිය, මාලිගාව). After short a, the vowel is usually deleted instead (පොතේ)SOFTH (glide), L (after a)02:VS-030, 05:SN-010, 05:PH-009
G-VS-07Spoken diphthongs in native and English words are V + යි / V + වු (අයියා, ළමයි, කවුද, ලයිට්)SOFTH02:VS-022, 05:PH-010
G-VS-08ඏ ඐ ෟ (alone) ෳ are not used in modern Sinhala. The letter ඎ is obsolete (its sign ෲ survives in a few words). ෟ survives only inside ෞ/ඖHARD (normal mode)H01:INV-018, 02:VS-021
G-VS-09Vowel length is phonemic, with one long sign per short sign: a→aa, ae→aee, i→ii, u→uu, e→ee, o→oo, ru→ruuHARDH02:VS-032
G-VS-10ර and ළ have irregular glyphs with u/uu (and ර with ae/aee). Never encode a "visual" substituteHARDH02:VS-011…013
G-VS-11ෛ ෞ (and ඍ ඓ ඖ) are Sanskrit-loan only. ෘ ෲ write Sanskrit ṛ, and are also the usual spelling of /ru, ruː/ after a consonant in other words and English loans (G-VS-15). ැ/ෑ (ඇ ඈ) are distinctly Sinhala and almost absent from Sanskrit wordsSOFTH02:VS-019/020/022/034, 05:SN-012
G-VS-12"ru" has three readings: ර+ු, ෘ, and rakaransaya+ු. ගෘහ (house) ≠ ග්‍රහ (planet). ක්‍රු/ක්‍රූ must be encoded with ු/ූ, never ැ/ෑ (ක්‍රෑර is wrong)HARD (encoding), SOFT (choice)H02:VS-012/016/020/034, 03:HC-024, 05:G2P-011
G-VS-13Per-consonant sign validity (41 × 17): see data/validity.json. Aspirates rarely take ැ/ෑ; ළ rarely takes long signs; sanyaka letters rarely take long signsSOFTM02:§8
G-VS-14On ්‍ය clusters, SLS lists a aa u uu e ee o oo. On ්‍ර clusters, a aa ae aee i ii e ee ai o oo au, plus u uu (SLS omits them; they are valid but rare, see G-VS-15). Other combinations are encodableSOFTH01:INV-018, 02:VS-016, 03:HC-020/021
G-VS-15C + r + u/uu has two correct spellings: C + ෘ/ෲ (කෲර, මෘදු, ගෲප්) and rakaransaya + ු/ූ (ක්‍රූර). Both read /Cru(ː)/. ෘ/ෲ is the usual one: it is the most frequent spelling in 155 of 205 words attested in more than one spelling (CC-02). The look-alike rakaransaya + ැ/ෑ is never correct (G-VS-12)STYLEH02:VS-020/034, 03:HC-024

3. Hal and consonant clusters ​

IDRuleKindConfSources
G-HC-01A consonant with no following vowel gets hal. This covers word-final position (පොතක්, මල්), before another consonant, and geminatesHARDH03:HC-010/012/013
G-HC-02Pronunciation never decides hal. Whether inherent a is said as [a] or [ə] is not writtenHARDH03:HC-010/011, 05:G2P-001
G-HC-03Geminates are C ් C with no ZWJ (අම්මා, අක්කා, පත්තරය)HARDH03:HC-042, 05:PH-008
G-HC-04These letters never geminate: ඟ ඬ ඳ ඹ ඦ ඞ ෆ හ ශ (learned ශ්ශ exists: නිශ්ශබ්ද), and ළHARDH03:HC-042, 05:PH-008, 04:NS-075
G-HC-05Native syllables are (C)V(C) with no initial clusters. Initial clusters are tatsama or English (ප්‍ර, ස්ව, ස්ටේෂන්)SOFTH03:HC-040, 05:PH-001
G-HC-06Sanyaka ඟ ඦ ඬ ඳ ඹ never take hal, so never ්‍ය/්‍ර after them. They are never word-initialHARDH02:VS-025, 03:HC-014, 04:NS-031, 05:PH-003
G-HC-07ළ never takes hal (stem-final ළ stays ළ; geminate l is ල්ල)HARDM03:HC-014, 04:NS-058/075
G-HC-08ඞ appears only as ඞ් (before a velar, Pali or learned). It never takes a vowel sign. Modern spelling uses ංHARDH01:INV-018, 02:VS-023, 04:NS-040
G-HC-09ණ never ends a stem with hal (a final consonant surfaces as න්). ණ් is fine inside clusters (ණ්ඩ)SOFTM04:NS-058
G-HC-10Each joining point of a long cluster is independent: ZWJ only where a reduced form is wanted (ස්ත්‍රී, රාෂ්ට්‍ර, ශාස්ත්‍ර)HARDH03:HC-044
G-HC-11Yansaya is mandatory for C + ය: C ් ZWJ ය (වාක්‍ය, විද්‍යාව). වාක්ය without ZWJ is not acceptedHARDH03:HC-020, 05:SP-012
G-HC-12Rakaransaya is mandatory for C + ර: C ් ZWJ ර (ක්‍රම, ශ්‍රී, ප්‍රශ්නය). After ම න ල, ර normally starts a new syllable and takes plain hal (දුම්රිය, හෙන්රි). ම්‍ර (තාම්‍ර) is valid encoding but rareHARDH (rule), M (m/n/l usage)03:HC-021/052
G-HC-13Repaya is optional style: ර් + C or ර ් ZWJ + C, both correct (කර්ම / කර්‍ම)STYLEH01:INV-006, 03:HC-022
G-HC-14No yansaya after ර (SLS 1134). ර ් ZWJ ය means repaya + ය. kārya has 3 accepted spellings: කාර්‍ය්‍ය (traditional), කාර්‍ය, කාර්ය. ර්‍ර is unattestedHARD (ra+yansaya), DESIGN (which kārya)H03:HC-033/034/054, 05:PH-012
G-HC-15Other conjuncts (bandi akuru) are optional and mostly classical. Still seen: ක්‍ෂ, ක්‍ව, න්‍ද, න්‍ධ, න්‍ථ, ත්‍ථ. Not contemporary: ද්‍ධ, ද්‍ව, ට්‍ඨ, ඤ්‍ච. Modern default is plain halSTYLEM03:HC-030, 03:§9b
G-HC-16ක්ෂ and ක්‍ෂ are the same spelling of kṣ. Both are commonSTYLEH03:HC-032, 05:SP-011
G-HC-17Touching letters C ZWJ ් C are Pali/classical only. They need SLS Level 3 fonts, which barely existSTYLEH03:HC-031/062
G-HC-18Two accepted spellings of -tva words. තත්ත්වය / සත්ත්වයා is the grammatical form (Sanskrit tat + tva); the reduced තත්වය / සත්වයා is the modern form and is accepted as valid. Both are frequent. No official ruling existsSTYLEM03:HC-043, 05:SP-008

4. Nasals and ayogavaha signs ​

IDRuleKindConfSources
G-NS-01ං follows a vowel, a consonant (+ sign) or a sanyaka. Never word-initial, never after hal, never takes a sign, always last in its clusterHARDH01:INV-008, 04:NS-001/002, 05:PH-004
G-NS-02ං = [ŋ]. It has replaced ඞ් (and often ඤ්): ලංකාව, මංගල, වංචාවSOFTM04:NS-003/004
G-NS-03Nasal before a consonant (tatsama): k/g group, and y r l v ś ṣ s h → ං; ට/ඩ group → ණ්; ත/ද group → න්; ප/බ group → ම්; c/j group → ං (modern) or ඤ් (Pali). Native words use න් before ස (පන්සල)SOFTH (retroflex, dental), M (others)04:NS-005/043/055/056, 05:SP-009
G-NS-04sam- + vowel → ම + sign (සමාගම). ං never stands before a vowel inside a wordHARDM04:NS-006/010
G-NS-05ඃ is rare, Sanskrit-only, [h], with the same position rules as ං. Most Sanskrit visarga surfaces as ෝ / ර් / ශ් / ස් / ෂ් through sandhi. It is never predictable from sound aloneSOFTH04:NS-020…023, 05:SN-011
G-NS-06ඁ candrabindu is not modern Sinhala. It has no place in modern textHARDH01:INV-011, 04:NS-025
G-NS-07Sanyaka = a short nasal fused to a voiced stop (ඞ+ග, ඤ+ජ, ණ+ඩ, න+ද, ම+බ). It contrasts with nasal + hal + stop: කඳ trunk ≠ කන්ද hill, අඟල ≠ අංගයHARD (contrast)H01:INV-021, 04:NS-030/032, 05:SP-010
G-NS-08Sanyaka vs cluster is lexical. Tatsama words never use sanyaka. Native and tadbhava words often do. English loans use clusters (ලන්ඩන්). Some words accept both (මග/මඟ)SOFTM04:NS-033/034
G-NS-09Nothing nasal precedes a sanyaka (no අංඹ, no න්ඳ)HARDM04:NS-037
G-NS-10Prenasal + consonant in compounds → ං or a hal nasal (ගඟ + වතුර → ගංවතුර)SOFTH05:SN-007, 04:NS-007
G-NS-11ඦ is practically unused. Keep it in the inventory, but treat it as obsoleteSOFTH01:§8, 04:NS-036
G-NS-12ඥ = Sanskrit jñ (ඥානය, ප්‍රඥාව, the suffix -ඥ). gn is not always ඥ (අග්නි, නග්න, ලග්න)SOFTH04:NS-042/046
G-NS-13ny is usually න්‍ය (අන්‍ය, ශූන්‍ය, ධන්‍ය). ඤ is rare (ඤාණ, පඤ්ච, සඤ්ඤා)SOFTH04:NS-041/047
G-NS-14Word-final ං contrasts with final ම් / න් (දං ≠ දන්). Colloquial final ං stands for formal ම්/න් (මං, එහෙනං)SOFTM04:NS-008/009
G-NS-15A spoken final [ŋ] can be written ං, න් or ම් (මිනිසුන් is said minisuŋ). A romanized -ng at word end is ambiguousSOFTM03:HC-012, 05:PH-007

5. Phonotactics and sandhi ​

IDRuleKindConfSources
G-PH-01Word-initial bans: sanyaka letters, ඞ, ං, ඃ, ඎ. ණ only in ණය (plus archaic words). ළ is fine initially (ළමයා)HARD (except ණ)H04:NS-031/057, 05:PH-003…006
G-PH-02Words end in a vowel, a hal consonant or ං. Never in a sanyaka. ළ/ණ end a word only with a vowelHARDH05:PH-007, 04:NS-058
G-PH-03Vowel hiatus inside a word occurs only at compound boundaries and in place names (ගිරිඋල්ල). A romanization needs an explicit separator to express itSOFTH05:PH-009, 02:VS-029
G-PH-04The 10 school sandhi types, plus frozen Sanskrit sandhi (dīrgha, guṇa, vṛddhi, yaṇ, visarga), explain most learned spellings (ඉත්‍යාදි, නිර්මාණ, දුෂ්කර, පෞද්ගලික). They are lexical facts, not productive rulesSOFTM05:SN-001…014

6. Spelling distinctions (sound-alike groups): lexical, not rule-based ​

IDRuleKindConfSources
G-SP-01These groups are pronounced alike: {ක ඛ} {ග ඝ} {ච ඡ} {ජ ඣ} {ට ඨ} {ඩ ඪ} {ත ථ} {ද ධ} {ප ඵ} {බ භ} {න ණ} {ල ළ} {ස ශ ෂ} {ඤ ඥ}. Sound cannot pick the letter; a corpus-ranked lexicon must (one study reached 82% with ranking alone)SOFTH01:INV-023, 05:G2P-010, 05:SP-003…006
G-SP-02ණ: after ර / ෂ / ඍ in nouns, unless a dental, palatal, retroflex, ල, ශ or ස comes between (the ṇatva rule); before ට/ඩ; in honorifics (-ආණ, -අණි); in past/passive forms (-ඉණි, -උණු)SOFTH04:NS-051…060, 05:SP-003
G-SP-03න: in native verbs even after ර (මරන ≠ මරණ); before dentals and ස; after ස/ශ; in geminates; at compound boundariesSOFTH04:NS-056/061…063, 05:SP-003
G-SP-04ළ: in past forms of ර-roots (කර → කළ); the prefix පිළි-; "small / young" words (ළමා); the first l of an l…l noun; before a sanyaka (M); Pali/Sanskrit retroflex reflexes (පොළොව)SOFTH04:NS-071…079, 05:SP-004
G-SP-05ස / ශ / ෂ: native words use ස. ශ with palatals, before ව, in ශ්‍ර, and first of two sibilants. ෂ before retroflexes, in ක්ෂ, after vowels other than a/aa before k/p (ruki), and in -ඉෂ්ඨ. ස before dentalsSOFTH04:NS-080…087, 05:SP-005
G-SP-06Aspirates have no rule; they are learned per word. Some words have accepted variants (කථා/කතා)SOFTH05:SP-006
G-SP-07Vowel length is the commonest error class. The verbal noun ending is long -ීම (කිරීම)SOFTH02:VS-033, 05:SP-007
G-SP-08Pairs that differ in meaning must stay distinct: කණ/කන, වණ/වන, කල/කළ, පල/පළ, මරණ/මරනSOFTH04:§6, 05:SP-003/004

7. Writing Sinhala in Latin script (informal romanization) ​

Measured from the Dakshina corpus of native-speaker romanizations; see 06-romanization.md. Every corpus number in this file is reproduced by tools/corpus_counts.py (reports/corpus-counts.md).

IDFindingConfSources
G-TY-01th = ත and t = ට are near-universal (93–99%)H06:RS-001
G-TY-02ද is written d (99%); ධ is written dhH06:RS-002/003
G-TY-03ී is written ee 38% of the time and ූ is written oo 22%H06:RS-006/007
G-TY-04Vowel length is rarely marked: ා is written a 97% of the timeH06:RS-009
G-TY-05Informal chat drops vowels (nthi, mta); only a lexicon or fuzzy matching recovers themH06:RS-010
G-TY-06ඳ ඹ ඟ are written nd mb ng (ඳ = nd 91%) and ං is written n (91%)H06:RS-012/014
G-TY-07ඇ is written e or a, rarely aeH06:RS-008
G-TY-08Retroflex ණ / ළ are never distinguished in informal writing (always n, l)H06:RS-025
G-TY-09Schwa [ə] is written a (sometimes e): karanawa / keranawa = කරනවා. Ordered schwa rules (98% accurate) predict itH05:G2P-001…009, 03:HC-011
G-TY-10w and v both stand for ව; w is preferred (72%)H06:RS-011

8. Loanword conventions ​

IDRuleConfSources
G-LW-01f → ෆ (older loans ප: කෝපි)H05:LW-001, 01:INV-022
G-LW-02v/w → ව; z → ස; ʒ → ජ; θ → තM05:LW-002/006/008
G-LW-03æ → ඇ or ඈ (monosyllables tend to ඈ: ෆෑන්, බෑග්; polysyllables ඇ: බැංකුව). Both occurM05:LW-003
G-LW-04ɒ → ඔ; əʊ / ɔː → ඕ; English schwa and -er → අ / -අර්H05:LW-004/005
G-LW-05sh → ෂ (popular), sometimes ශ; x → ක්ස් (never ක්ෂ); -ng → ං (ඉංග්‍රීසි)M05:LW-007/009/012
G-LW-06Loans keep a final hal (බස්, ෆෝන්). Literary forms add -ය/-ව (බෝලය, නෝට්ටුව)H05:LW-011

9. A phonetic romanization built on these rules ​

07-phonetic-romanization.md specifies a phonetic romanization (conventions R-01…R-15) that encodes every rule above: it can produce all 791 letter forms the rules allow and, checked exhaustively, never produces a forbidden sequence.


10. Open gaps and conflicts ​

#GapWhere it matters
1NIE textbooks, the final SLS 1134:2004/2011 texts and the Sinhala Lekhana Rīthiya (1989) were not consulted. School-grammar rules rest on agreeing secondary sourcesG-SP-*, G-HC-07, alphabet counts
2The 41 × 17 validity table is a synthesis, not a corpus count. Only the ෘ / ෲ columns have been checked against a corpus (02:VS-035)G-VS-13 / validity.json
3Touching-letter and yansaya-with-repaya encodings changed between SLS draftsG-EN-08, G-HC-14
4ම්‍ර / න්‍ර / ල්‍ර have no attestation in a Sinhala source; plain ම්ර / න්ර is attested (03:HC-052)G-HC-12, R-07
5How ං before ය ර ල ව ශ ස හ is actually pronouncedG-NS-02/03

Recommended next step: count consonant + sign and cluster bigrams in a Sinhala corpus (Wikipedia dump or UCSC 10M). That turns validity.json from a synthesis into measured data, and it gives the disambiguation layer a better lexicon.