Skip to content

07: A grammar-correct phonetic romanization of Sinhala ​

This document specifies a phonetic romanization of Sinhala: a set of Latin-letter sequences and conversion rules that turn romanized text into Sinhala script. It is designed so that every result obeys the orthographic rules in 00-rules.md, while staying close to how Sinhala is informally written in Latin script (06-romanization.md). Compiled October 2026.

The reference implementation is src/sinhala_orthography/romanization.py. The sequence tables are src/sinhala_orthography/data/{consonants,vowels,specials}.json.


⚠️ Things that will surprise you ​

ConventionExample
1Lower-case t is retroflex but lower-case d is dental, matching informal usaget ට, th ත · d ද, D ඩ
2ai and au are written with a glide, not as ෛ / ෞlait → ලයිට්, kauda → කවුද
3ru / ruu straight after a consonant is written with ෘ / ෲ, the usual spelling; rakaransaya + ු / ූ is an optionkruura → කෲර, mrudu → මෘදු, karu → කරු
4Lower-case n before a velar becomes ංlankaava → ලංකාව
5Repaya is plain ර් by defaultkarma → කර්ම, kaarya → කාර්ය
6A rule-forbidden form is never produced; a valid fallback is used insteadkazd → කඳ (no hal on ඳ), Ba → බ (no word-initial ඹ)
7Case is significant: capitals mark retroflex, aspirate, ඇ/ඈ and diphthong lettersN ණ, L ළ, A ඇ, E ෛ

1. Design principles ​

  1. Never produce a forbidden sequence. All HARD rules in 00-rules.md hold for every output. This is checked exhaustively (§5).
  2. Every allowed letter form is producible. All 791 forms that data/validity.json allows can be written.
  3. Follow majority informal usage where it is unambiguous (06-romanization.md): th ත, t ට, d ද, w/v ව, ee for long e.
  4. Mark spelling-only distinctions explicitly. Distinctions speakers do not hear (aspiration, ණ/න, ළ/ල, ශ/ෂ, sanyaka) get an explicit marker: a capital, an h, or a z prefix. Without the marker, the plain letter is written, and a lexicon can recover the standard spelling (§4).
  5. Logical order. Romanization is written in pronunciation order, which is also Unicode storage order, so pre-base signs (ෙ, ෛ, ො…) need no reordering.

2. Sequence tables ​

Generated from the JSON tables by tools/build_spec_tables.py. The Notes column cites the convention (R-) or rule (G-) behind a row. Matching is case-sensitive and longest-first.

Consonants ​

A consonant written without a following vowel takes hal (්): k → ක්, ka → ක.

LetterRomanizationNotes
කk · calias
ඛkh · K · Calias
ගg
ඝgh · G
ඞXG-HC-08: hal only
ඟzgsanyaka
චch
ඡchh
ජj
ඣjh · J
ඤzkR-12
ඥzhR-12
ඦzjR-14; archaic option
ටt
ඨT
ඩDR-01
ඪDhR-01
ණN
ඬzDR-01 sanyaka
තth
ථthh
දd · qR-01; alias
ධdh · dhhR-01; alias
නn
ඳzd · zdh · zqR-01 sanyaka; alias
පp
ඵph · P
බb
භbh
මm
ඹBsanyaka
යy
රr
ලl
වw · v · W · V
ශsh
ෂSh · S
සs
හh
ළL
ෆf

Vowels ​

The independent letter is used at the start of a word; after a consonant the vowel is written as a sign.

VowelIndependentSignRomanizationNotes
aඅ(inherent)a
aaආ◌ාaa
aeඇ◌ැA · aeR-04
aeeඈ◌ෑAa · AA · aeeR-04
iඉ◌ිi
iiඊ◌ීii · IR-02
uඋ◌ුu · U
uuඌ◌ූuu · UU · UuR-02
ruඍ◌ෘRR-06
ruuඎ◌ෲRRR-06 (independent ඎ archaic only)
eඑ◌ෙe
eeඒ◌ේeeR-02
aiඓ◌ෛER-05b
oඔ◌ොo · O
ooඕ◌ෝoo · OO · OoR-02
auඖ◌ෞAu · AUR-05b
iluඏ◌ෟ~lR-14; archaic option
iluuඐ◌ෳ~llR-14; archaic option

Signs and joins ​

OutputRomanizationNotes
ංxG-NS-01: needs a vowel base
ංznG-NS-01
ංMG-NS-01
ඃHG-NS-05
ඁ~nR-14 / G-NS-06; archaic option
touching letters (C ZWJ ් C)+R-14 / G-HC-17: C ZWJ ් C; archaic option

3. Conversion rules ​

#RuleExampleBasis
C-1A consonant with no following vowel takes halpin → පින්G-HC-01
C-2C + y → yansaya (C ් ZWJ ය), except after රvaakya → වාක්‍යG-HC-11, G-HC-14
C-3C + r → rakaransaya (C ් ZWJ ර) for every consonant that takes hal, except ම න ල (plain hal; rakaransaya under classical)kramaya → ක්‍රමය, dumriya → දුම්රියG-HC-12, R-07
C-4ර + C → plain ර් (repaya with ZWJ under repaya_zwj)karma → කර්මG-HC-13, R-08
C-5Geminates and other clusters use plain hal (ZWJ conjuncts under classical)amma → අම්ම, akShara → අක්ෂරG-HC-03, R-10
C-6A vowel after a vowel is joined by a glide: ය after front vowels, ව after back vowels; after a / aa the following vowel decidestoppia → ටොප්පිය, dua → දුව, ai → අයිG-VS-06, R-05
C-7n + k / kh / g / gh after a vowel → ං; word-final ng → ංingriisi → ඉංග්‍රීසි, sing → සිංG-NS-03, R-11
C-8ං needs a vowel base: a consonant before x keeps its inherent a; a vowel after ං turns it into ම + signkx → කංG-NS-01, G-NS-04
C-9Sanyaka and ළ never take hal: without a vowel they keep the inherent akaL → කළG-HC-06, G-HC-07
C-10No sanyaka at the start of a word: the plain stop is writtenzda → දG-PH-01
C-11ඞ only as ඞ් before a consonant; ඞ + vowel → ඟ + sign; ං before a sanyaka is droppedaXa → අඟ, axzda → අඳG-HC-08, G-NS-09
C-12A sequence that cannot be written is left in Latin scriptx (no base) → xG-NS-01
C-13C + r + u / uu → C + ෘ / ෲ where that form is attested (valid, loan or rare in validity.json: මෘ, not ලෘ); otherwise C-3 applies (rakaransaya + ු / ූ, plain hal after ම න ල). Rakaransaya + ු / ූ throughout under rakaransaya_u; R / RR always write ෘ / ෲkruura → කෲර, mrudu → මෘදු, gruup → ගෲප්, dilrukshi → දිල්රුක්ශිG-VS-15, R-06

Options ​

OptionEffectConvention
repaya_zwjර්‍ + C instead of ර් + CR-08
classicalZWJ conjuncts for ක්‍ෂ ක්‍ව ග්‍ධ ට්‍ඨ ත්‍ථ ත්‍ව ද්‍ධ ද්‍ව න්‍ථ න්‍ද න්‍ධ න්‍ව ඤ්‍ච, and rakaransaya after ම න ල (ම්‍ර න්‍ර ල්‍ර)R-10, G-HC-15, R-07
archaicඏ ඐ ෟ ෳ (~l, ~ll), standalone ඎ (RR), ඁ (~n), ඦ (zj), touching letters (+)R-14
rakaransaya_uC + r + u / uu as rakaransaya + ු / ූ (ක්‍රු, ක්‍රූ) instead of C + ෘ / ෲ. The name means "rakaransaya + u": the u is the vowel sign ු. Rakaransaya itself is always written; the option changes only C + ru / ruuR-06
retroflex_dd ඩ · dh ද · D ඪ · Dh ධ · zd ඬ, the older keyboard convention (dhh ධ, zdh ඳ, q ද keep their letters)R-01

4. Disambiguation with a lexicon ​

A romanization fixes one spelling, but about 14 groups of Sinhala letters are pronounced alike (G-SP-01). People rarely mark vowel length or aspiration in Latin script, so the standard spelling often needs a word list. src/sinhala_orthography/lexicon.py:

romanized ──► conversion (§3) ──► one spelling ──► sound key ──► words of a frequency list
                                                       │            with the same key,
                                                       │            most frequent first
                                                       └─ explicit marker present and the
                                                          spelling is a word, or a lone
                                                          vowel letter → it stays first

The sound key erases: aspiration, ණ/න, ළ/ල, ශ/ෂ/ස, ද/ඩ, ඤ/න, ඥ/ග්න, sanyaka vs nasal + stop, ං/න්, vowel length, ඇ/එ, ෛ/යි, ෞ/වු, ෘ/්‍රු and ZWJ.

A lone vowel letter keeps its spelling. When an input with an explicit marker converts to a single independent vowel (A ඇ, Aa ඈ, R ඍ, E ඓ, Au ඖ), that letter stays first although the list has no such word: a letter typed on its own is meant as that letter, and frequency would otherwise turn it into a sound-alike word (ඍ → රු, ඓ → අයි). Unmarked vowels are still matched by sound (e → ඒ).

Options apply to the words too. A word list is written in the usual style (කෲර, කර්ම), so candidates() rewrites each word in the style of the converter options before returning it (restyle()): rakaransaya_u turns C + ෘ/ෲ into rakaransaya + ු/ූ (ක්‍රූර), repaya_zwj joins repaya with ZWJ (කර්‍ම) and classical joins the bandi akuru pairs (අක්‍ෂර). Otherwise the lexicon would undo the options on every word. With no options nothing changes.

Evaluation. 27 everyday words, written as people informally romanize them, with the University of Moratuwa NLPC frequency list (2.1M words, pinned; CC-14 in reports/corpus-counts.md). The list is not included here: tools/corpus_counts.py downloads it, and tests/test_lexicon.py runs the same words against any list.

Correct spelling ranked first
Conversion rules alone9 / 27
Rules + lexicon27 / 27
RomanizedRules aloneRules + lexicon
hondaහොන්දහොඳ
sinhalaසින්හලසිංහල
thiyenawaතියෙනවතියෙනවා
badaබදබඩ
vaidyaවයිද්‍යවෛද්‍ය
krushiකෘශිකෘෂි
lamayaලමයළමයා

17.7% of the list's sound keys are shared by two or more words (e.g. කළ කල කාල කලා; CC-15), so frequency alone cannot always choose. Context from the previous word would help.


5. Verification ​

tools/check_romanization.py (report: reports/romanization-check.md):

CheckResult
Coverage: letter forms allowed by data/validity.json that can be produced791 / 791
Safety: inputs of 1–3 sequences (plus space), all six option sets3.98 million inputs, 0 violations of 14 forbidden patterns
Fixtures: tests/test_romanization.pyall pass

6. Conventions and their rationale ​

IDConventionRationale
R-01d ද · dh ධ · D ඩ · Dh ඪ; sanyaka zd ඳ, zD ඬ; q = ද aliasInformal writing uses d for ද 99% of the time and dh for ධ (G-TY-02). Capitals mark retroflex letters, as with N ණ and L ළ. ද is also 5 times as frequent as ඩ in running text (74% vs 15% of d-letters, CC-12), so the unmarked key goes to the common letter. Older keyboard schemes write d ඩ / dh ද, mirroring t ට / th ත (06:RS-030); retroflex_d offers that convention
R-02ee ඒ · oo ඕ · ii ඊ · uu ඌConsistent doubling for length. Informal ee = ී (G-TY-03) is recovered by the lexicon
R-03nd / mb / ng + vowel are written as clusters; the lexicon chooses sanyaka (ඳ ඹ ඟ) or clusterSanyaka vs cluster is lexical (G-NS-08), and informal writing uses nd/mb/ng for both (G-TY-06)
R-04A / ae ඇ · Aa / AA / aee ඈae matches ISO 15919 7-bit; A keeps a one-letter form
R-05ai / au → glide spelling (අයි / අවු, C + යි / C + වු); ෛ ඓ via E, ෞ ඖ via AuTwo vowels are never written side by side (G-VS-03, G-VS-06). In the NLPC list, words with C + යි outnumber C + ෛ 11:1 and C + වු outnumber C + ෞ 1.7:1 (CC-05). A word that has ෛ or ෞ is nearly always written with it (වෛද්‍ය 46,508 vs වයිද්‍ය 142, CC-06), and the lexicon restores it
R-06C + ru / ruu = C + ෘ / ෲ; R / RR also write ෘ / ෲ; rakaransaya_u gives ක්‍රු / ක්‍රූ instead; standalone R = ඍ; ru after a vowel = ර + ුෘ / ෲ is the usual spelling of /Cru(ː)/ in Sanskrit and English loans alike (G-VS-15): in the NLPC list it is the most frequent spelling in 155 of 205 words written more than one way (කෲර 1,667, ක්‍රූර 9; CC-02, CC-03). ගෘහ and ග්‍රහ stay distinct (G-VS-12)
R-07Rakaransaya after every consonant that takes hal, except ම න ල: there ර starts a new syllable and takes plain hal (දුම්රිය, හෙන්රි); classical writes ම්‍ර (තාම්‍ර)Running text writes ම්ර / න්ර plain (දුම්රිය 30,800 vs දුම්‍රිය 82; 03:HC-052, CC-10); a ZWJ would draw a rakaransaya under ම. ම්‍ර is attested only for Sanskrit tatsama words
R-08Repaya defaults to plain ර් + CBoth forms are standard (G-HC-13); the plain form dominates the NLPC list (2,182,087 vs 26,785 tokens, CC-07)
R-09kārya → කාර්යFollows from R-08; 18,722 occurrences in the NLPC list (CC-08)
R-10ක්ෂ and other bandi akuru default to plain hal; ZWJ conjuncts are optionalConjuncts other than yansaya/rakaransaya are optional and mostly classical (G-HC-15/16)
R-11n + velar → ං automatically; before h s sh y r l v the lexicon chooses ං or න්G-NS-03; native and Sanskrit words differ before ස (පන්සල vs සංසාරය)
R-12ny → න්‍ය; gn → ග්න; ඥ = zh; ඤ = zkන්‍ය and ග්න are the common readings (G-NS-12/13); ඥ and ඤ are recovered by the lexicon
R-13Case-sensitiveCapitals carry the spelling-only distinctions (principle 4)
R-14Archaic letters and touching letters only under the archaic optionNot used in modern Sinhala (G-VS-08, G-HC-17, G-NS-06)
R-15Spelling-only distinctions are resolved by a frequency lexiconG-SP-01: sound cannot decide them

7. Limitations ​

  • Disambiguation quality depends on the word list. The NLPC list is drawn from news text, so formal words outrank colloquial ones (ඔය before ඔයා). Some copies of it have lost their ZWJs; normalize() repairs yansaya and rakaransaya on load. It runs only on a list with no ZWJ at all, because it cannot see word boundaries and also joins බවත් + ය → බවත්‍ය.
  • The per-consonant validity table (data/validity.json) is a synthesis of the sources, not a corpus count (00-rules.md §10).
  • Informal romanization often drops vowels altogether (G-TY-05). Recovering those needs fuzzy matching, which is out of scope here.