Skip to content

06: Romanization of Sinhala ​

This document surveys how Sinhala is written in Latin script. It covers (a) the formal romanization systems (ISO 15919 and its 7-bit ASCII form, the Sri Lanka national system, the UN 1972 table, ALA-LC and KNAB) and (b) informal romanized Sinhala ("Singlish") as people actually write it, measured from the Sinhala portion of the Dakshina corpus and described in the published research on Singlish back-transliteration. It closes with a per-letter comparison, a set of conventions (RS-xxx) and recommendations for a phonetic romanization of Sinhala.

Compiled October 2026.

Legend ​

MarkMeaning
✅Taken from a primary source (the standard, an official report, or a published mapping table)
🧪Measured from a corpus (the Dakshina Sinhala data, §2a)
❓Not confirmed. The source says nothing, or the value is an inference. Don't treat it as a mapping
-No value recorded for this letter in the sources used here

Confidence for each RS convention: H means several primary sources agree, M means one primary source or an inference, L means a guess.

Letter IDs follow the rest of this repository: consonants ka kha ga gha nga(ඞ) nnga(ඟ) ca cha ja jha nya(ඤ) jnya(ඥ) nyja(ඦ) tta ttha dda ddha nna nndda(ඬ) ta tha da dha na nda(ඳ) pa pha ba bha ma mba(ඹ) ya ra la va sha(ශ) ssa(ෂ) sa ha lla(ළ) fa; vowels a aa ae aee i ii u uu ru ruu ilu iluu e ee ai o oo au; signs hal anusvara visarga yansaya rakaransaya repaya.


Summary ​

  • Informal writing uses th = ත and t = ට. ත is written th 93% of the time and ට is t 99%. ISO 15919 does the opposite (t = ත, ṭ = ට) (RS-001).
  • Informal d is ද first. People write d for both ද (99%) and ඩ (100%), and keep dh for ධ (69%), which agrees with ISO (RS-002, RS-003).
  • Informal ee/oo mean ī/ū, not ē/ō. ී is written ee 38% of the time, ේ almost never. ISO's 7-bit form uses ee/oo for ē/ō (RS-006, RS-007).
  • Vowel length is rarely marked. ා is written a 97% of the time, and ී is i 58% (RS-009).
  • The formal systems disagree on the nasals. ISO, the Sri Lanka national system and ALA-LC each romanize ං and ඞ differently (RS-029).

⚠️ Things that will bite you ​

#GotchaRule
1Informal kohomada means කොහොමද: d is ද, not ඩ. Reading d as ඩ gives the wrong wordRS-002
2Informal ee is ඊ/ී (kireemata, pawathee), not ඒ as in ISO 7-bitRS-006
3Informal oo is ඌ/ූ (soodanam), not ඕ as in ISO 7-bitRS-007
4Plain nd, mb, ng are ambiguous: න්ද or ඳ, ම්බ or ඹ, න්ග or ඟ or ංග. Informal writing uses them for the prenasalized letters (ඳ is nd 91% of the time)RS-012
5Informal n for ං (91%) can't be told apart from න්RS-014
6c is ච in ISO, but informal writing uses ch for ච (95%)RS-005
7Chat-style Singlish drops vowels (nthi, mta, kynna), so one spelling can stand for several wordsRS-010
8ඏ ඐ ෟ ෳ have values only in the formal systems; there is no informal data for themRS-018

1. Formal romanizations ✅ ​

1a. ISO 15919 (2001), with its 7-bit ASCII form ​

Source: the ISO 15919 table on Wikipedia, read from the page source. https://en.wikipedia.org/wiki/ISO_15919. The 7-bit ASCII form is given in brackets.

  • Vowels: අ a · ආ ā (aa) · ඇ æ (ae) · ඈ ǣ (aee) · ඉ i · ඊ ī (ii) · උ u · ඌ ū (uu) · ඍ r̥ (,r) · ඎ r̥̄ (,rr) · ඏ l̥ (,l) · ඐ l̥̄ (,ll) · එ e · ඒ ē (ee) · ඓ ai · ඔ o · ඕ ō (oo) · ඖ au
  • Consonants: ක k · ඛ kh · ග g · ඝ gh · ඞ ṅ (;n) · ඟ n̆g (^ng) · ච c · ඡ ch · ජ j · ඣ jh · ඤ ñ (~n) · ඦ n̆j (^nj) · ට ṭ (.t) · ඨ ṭh · ඩ ḍ · ඪ ḍh · ණ ṇ · ඬ n̆ḍ (^n.d) · ත t · ථ th · ද d · ධ dh · න n · ඳ n̆d (^nd) · ප p · ඵ ph · බ b · භ bh · ම m · ඹ m̆b (^mb) · ය y · ර r · ල l · ළ ḷ (.l) · ව v · ශ ś (sh) · ෂ ṣ (.s) · ස s · හ h · ෆ f
  • Signs: ං ṁ (;m) · ඃ ḥ (.h) · ් has no symbol
  • ඥ isn't listed separately. It's normally romanized as the conjunct jñ ❓.

1b. Sri Lanka national system (Survey Department, Cabinet-approved 4 Sep 2018, amended later) ​

Source: UNGEGN Working Group on Romanization Systems report, v5.0, Oct 2021. https://arhiiv.eki.ee/wgrs/rom2_si.htm. The report states there is still no UN-approved system for Sinhala.

  • Based on ISO, but these values differ:
    • ඞ = ṁa and ං = ṅ (ISO has these the other way round)
    • ඓ = ĩ (it was ai in 2018)
    • ඣ = qa (it was jha)
    • ඨ = ṯa (it was ṭha)
    • ෂ = sha, ශ = ś
    • ඥ = gna
    • ඹ = ḅa (it was m̆ba)
    • ඦ = n̆ǰa
  • Other listed values: ඟ n̆ga · ඬ n̆ḍa · ඳ n̆da · ඇ æ · ඈ ǣ · ඏ ḷ · ඐ ḹ · ඍ ṛ · ඎ ṝ
  • Ligatures: ර්‍ is r-, ්‍ර is -r, ්‍ය is -y. The Survey Department's online converter is at https://www.survey.gov.lk/RomanizationConverter/.

1c. UN 1972 "Sharma" table (never approved) ​

Source: interscript map un-sin-Sinh-Latn-1972, citing the 1972 UN conference papers. https://github.com/interscript/maps

  • Short vowels carry a breve: ඇ æ̆ · ඈ æ · එ ĕ · ඒ e · ඔ ŏ · ඕ o
  • ච ch · ඡ chh · ශ sh · ෂ ṣh · ං ṁ · ඞ ṅ

1d. ALA-LC (Library of Congress) Sinhalese ​

Source: interscript maps alalc-sin-Sinh-Latn-1997 and -2011. The LoC PDF itself was not consulted, so this is medium confidence.

AreaALA-LC 1997
Vowelsඇ ă · ඈ â · ඏ ḷ · ඐ ḹ · ෟ ḷ · ෳ ḹ · ē / ō for the long forms
Consonantsච ca · ඡ cha · ශ ś · ෂ ṣ
Anusvaraං ṃ
Prenasalizedඟ ṅga · ඦ ñja · ඬ ṇḍa · ඳ nda · ඹ ṃba

Rules:

  • ALA-LC writes anusvara as the nasal of the following consonant's class: ṅ, ñ, ṇ, n or m.
  • A sanyaka (prenasalized) letter followed by an aspirate is written as unaspirated + aspirate.

The 2011 test strings use ă / â for ඇ / ඈ.

1e. KNAB (Estonian place-name database), 1989 ​

Source: https://arhiiv.eki.ee/knab/lat/kblsi1.pdf

  • Vowels: ඇ è · ඈ ê · එ ĕ · ඒ e
  • Consonants: ච cha · ශ sha · ෂ ṣha
  • ං ṁ (ń)
  • It marks prenasalized letters with a middle dot (·mba).
  • It notes that in local practice è / ê are written e, and v is written w.

2. Informal Singlish: how people actually write Sinhala in Latin script ​

2a. Measured: the Dakshina corpus, Sinhala portion 🧪 ​

Corpus: 10,000 Wikipedia sentences romanized by native speakers (Roark et al., LREC 2020), plus a lexicon of attested variants: the Sinhala files of the official Dakshina v1.0 release, pinned by hash.

Method: tools/corpus_counts.py (with tools/dakshina_counts.py) splits each Sinhala word into letters and signs and aligns it to its romanization by dynamic programming over a broad set of candidate spellings per letter. The candidate probabilities are re-estimated from the alignments for four rounds, so the counts do not depend on the order of a hand-written table. The percentages are shares of the aligned occurrences of each letter (claims CC-17 and CC-18 in reports/corpus-counts.md).

  • "Sentences" = sentence word tokens: 127,072 of 139,154 aligned (tokens with digits or no Sinhala letter left out).
  • "Lexicon" = lexicon variants, weighted by the number of annotators: 88,994 of 93,505 aligned.
LetterSentencesLexicon variants
කk 99%, c 1%k 99%
ඛ / ඝkh 71%, k 29% / gh 70%, g 30%kh 94%, k 6% / gh 93%, g 7%
ඟng 52%, g 48%ng 85%, g 15%
ච / ඡch 95%, c 5% / ch 98%, chh 1%ch 59%, c 41% / ch 99%
ජj 100%j 100%
ඤ / ඥn 70%, gn 30% / gn 98%, kn 1%n 82%, gn 16%, kn 1% / gn 94%, ny 4%, gy 2%
ට / ඨt 99% / t 92%, th 8%t 99% / th 81%, t 19%
ඩ / ඪd 100% / d 73%, dh 27%d 99% / dh 70%, d 30%
ණn 100%n 100%
ඬnd 83%, d 17%nd 91%, d 9%
ත / ථth 93%, t 7% / th 100%th 51%, t 49% / th 96%, t 4%
ද / ධd 99% / dh 69%, d 31%d 97%, dh 3% / dh 94%, d 6%
ඳnd 91%, d 9%nd 82%, d 12%, dh 4%
ඵ / භp 56%, ph 44% / bh 91%, b 9%ph 83%, p 17% / bh 89%, b 11%
ඹmb 88%, b 11%mb 93%, b 7%
වw 72%, v 28%v 83%, w 17%
ශ / ෂsh 85%, s 15% / sh 91%, s 9%s 53%, sh 47% / sh 52%, s 48%
ළl 100%l 100%
ෆf 87%, ph 13%f 90%, ph 10%
inherent aa 99%a 100%
ාa 97%, aa 3%a 95%, aa 5%
ැ / ඇe 64%, a 36%, ae ≈0% / a 60%, e 40%, ae ≈0%a 45%, ae 43%, e 12% / a 53%, ae 39%, e 8%
ෑe 68%, a 29%, aa 3%a 49%, ae 36%, e 10%
ී / ඊi 58%, ee 38%, ii 2% / i 72%, ee 27%i 83%, ee 15% / i 99%
ූ / ඌu 73%, oo 22%, uu 5% / u 94%, uu 6%u 93%, oo 6% / u 100%
ේ / ඒe 100% / e 99%, ee 1%e 99% / e 100%
ෝ / ඕo 98%, oo 2% / o 97%, oo 3%o 100% / o 100%
ෘru 93%, r 7%r 55%, ru 45%
ෛ / ඓai 99%, ei 1% / ai 89%, ei 11%ai 99% / ai 100%
ෞ / ඖau 97%, ou 3% / au 89%, ow 7%, ou 4%au 97%, ou 2%, ow 1% / au 100%
ංn 91%, ng 6%, m 3%n 60%, m 34%, ng 6%

Caveats:

  • These are Wikipedia sentences romanized carefully on request, not chat. Chat drops vowels much more (§2b).
  • In the lexicon, ත splits th/t about 50/50, so some annotators used ISO-like plain t for ත.

2b. Papers on Singlish back-transliteration ​

  1. Athukorala & Sumanathilaka, "Swa Bhasha: Message-Based Singlish to Sinhala Transliteration" (2022 / 2024). https://arxiv.org/abs/2404.13350
    • Writers drop vowels. One word appears as kiyanna, kianna, kynna, kynn and kiynna.
    • The authors handle this with numeric letter codes and fuzzy matching.
    • They note that rule-based converters need vowels to be written, and they compare their system against existing tools.
  2. Perera, Prabhath, Sumanathilaka, Anuradha, "IndoNLP 2025 Shared Task: Romanized Sinhala to Sinhala Reverse Transliteration Using BERT" (2025). https://aclanthology.org/2025.indonlp-1.16.pdf
    • Romanization is "ad-hoc" in their terms, with vowels omitted. For example, තාත්තා is written Thaaththaa, Thaththa, Thattha, Thatta or Tatta.
    • Their system combines a dictionary, rules for out-of-vocabulary words, and BERT for ambiguity.
    • Test set 2 is mostly vowel-less input.
  3. Perera & Sumanathilaka, "Evaluating Transliteration Ambiguity in Adhoc Romanized Sinhala" (RANLP 2025). https://aclanthology.org/2025.ranlp-1.107.pdf
    • 22 romanized words that each map to two Sinhala words.
    • Examples: nthi නීති/නැති · mta මට/මීට · es ඇස/එසේ · eda ඇද/එදා · bala බල/බාල · badu බඩු/බදු · ud උඩ/උදේ · dnna දන්නා/දෙන්නා · oya ඔය/ඔයා.
    • These show three things: length isn't marked, ඇ and එ collapse into one spelling, and d is ambiguous between ඩ and ද.
  4. Sumanathilaka et al., "Swa-bhasha Resource Hub" (arXiv 2507.09245, 2025). https://arxiv.org/abs/2507.09245
    • A survey of 2020–2025 systems and datasets.
    • Singlish isn't standardised and often drops vowels.
    • Swa Bhasha does better than existing tools and rule-based converters on vowel-less input.

2c. Everyday spellings ​

mama, oya/oyaa, kohomada (d for ද), ayubowan, ganna, thiyenawa (th for ත, no length mark), amma/ammaa, thaththa, api, eka, ekka, mokada, honda (nd for ඳ), lassana, sinhala/lanka (n for ං), wenawa/venava, karanna, kiyanna.


3. Per-letter comparison ​

Notes on the columns:

  • Consonants are shown without the inherent vowel (the national system and ALA-LC sources write it, e.g. ṁa, ca).
  • ISO 7-bit: the bracketed values from §1a. Where the ISO value is already plain ASCII, the 7-bit form is the same.
  • Sri Lanka national and ALA-LC: only the values the sources list (§1b, §1d). "-" means the source excerpt used here gives no value, not that the letter is absent.
  • Informal: the top spelling in the Dakshina sentences (§2a) and its share. Where no share is given, the spelling comes from the per-letter summary without a published percentage. n = … marks very small counts.

Consonants ​

IDSinhalaISO 15919ISO 7-bitSri Lanka nationalALA-LCInformal 🧪
kaකkk--k 99%
khaඛkhkh--kh 71% (k 29%)
gaගgg--g
ghaඝghgh--gh 70% (g 30%)
ngaඞṅ;nṁ-- (no data)
nngaඟn̆g^ngn̆gṅgng 52% (g 48%)
caචcc-cch 95%
chaඡchch-chch 98%
jaජjj--j 100%
jhaඣjhjhq-jh (n = 4)
nyaඤñ~n--n 70% (gn 30%)
jnyaඥjñ ❓-gn-gn 98%
nyjaඦn̆j^njn̆ǰñj- (no data)
ttaටṭ.t--t 99%
tthaඨṭh-ṯ-t 92%
ddaඩḍ---d 100%
ddhaඪḍh---d 73% (dh 27%)
nnaණṇ---n 100%
nnddaඬn̆ḍ^n.dn̆ḍṇḍnd 83% (d 17%)
taතtt--th 93%
thaථthth--th 100%
daදdd--d 99%
dhaධdhdh--dh 69% (d 31%)
naනnn--n
ndaඳn̆d^ndn̆dndnd 91%
paපpp--p
phaඵphph--p 56% (ph 44%)
baබbb--b
bhaභbhbh--bh 91%
maමmm--m
mbaඹm̆b^mbḅṃbmb 88% (b 11%)
yaයyy--y
raරrr--r
laලll--l
vaවvv--w 72% (v 28%)
shaශśshśśsh 85% (s 15%)
ssaෂṣ.sshṣsh 91%
saසss--s
haහhh--h
llaළḷ.l--l 100%
faෆff--f 87% (ph 13%)

Vowels (independent / sign) and signs ​

IDSinhalaISO 15919ISO 7-bitSri Lanka nationalALA-LCInformal 🧪
aඅ / (inherent)aa--a 99%
aaආ / ාāaa--a 97%
aeඇ / ැæaeæăැ e 64%; ඇ a 60%
aeeඈ / ෑǣaeeǣâe 68% (a 29%)
iඉ / ිii--i
iiඊ / ීīii--i 58% (ee 38%)
uඋ / ුuu--u
uuඌ / ූūuu--u 73% (oo 22%)
ruඍ / ෘr̥,rṛ-ru 93%
ruuඎ / ෲr̥̄,rrṝ-ru (n = 7)
iluඏ / ෟl̥,lḷḷ- (no data)
iluuඐ / ෳl̥̄,llḹḹ- (no data)
eඑ / ෙee--e
eeඒ / ේēee-ēe 100%
aiඓ / ෛaiaiĩ-ai 99%
oඔ / ොoo--o
ooඕ / ෝōoo-ōo 98%
auඖ / ෞauau--au 97%
hal්(none)(none)--not written (implied at end of word)
anusvaraංṁ;mṅṃ (class nasal)n 91%
visargaඃḥ.h--h (n = 1)
yansaya්‍ය-y-y-y-y
rakaransaya්‍ර-r-r-r-r
repayaර්‍r-r-r--r

4. Conventions (RS-xxx) ​

IDConvention / conflictEvidenceConf.
RS-001Informal writing uses th = ත and t = ට: ත is th 93%, ට is t 99%. ISO 15919 uses plain t = ත and ṭ = ට, and some lexicon annotators follow it (ත splits th/t about 50/50 in the lexicon)§1a, §2aH
RS-002d is ambiguous in informal writing. People write d for both ද (99%) and ඩ (100%). ISO separates them (d = ද, ḍ = ඩ). The RANLP pair badu බඩු/බදු shows the ambiguity in practice§1a, §2a, §2bH
RS-003dh = ධ. ISO uses dh = ධ, and informal writing agrees (69% in sentences, 94% in the lexicon)§1a, §2aH
RS-004ISO marks aspirates by adding h to the base letter (ch = ඡ, th = ථ). Informal writing barely separates aspirates from their plain letters: ථ and ත are both th, ඨ is mostly t, ඪ mostly d§1a, §2aH
RS-005c: ISO and ALA-LC use c = ච. Informal writing uses ch for ච (95%), c only 5%; ක is c 1%§1a, §1d, §2aM
RS-006ee: in ISO 7-bit, ee = ē (ඒ). In informal writing ee is a common spelling of ī (ී = ee 38%), and almost never of ē (ේ = ee ≈0%)§1a, §2aH
RS-007oo: in ISO 7-bit, oo = ō (ඕ). In informal writing oo is a spelling of ū (ූ = oo 22%), rarely of ō (ෝ = oo 2%)§1a, §2aH
RS-008ඇ: ISO is æ, with ae as its 7-bit form; UN 1972 æ̆, ALA-LC ă, KNAB è. Informal writing mostly uses e or a, rarely ae (ැ = ae ≈0% in sentences)§1, §2aH
RS-009Informal writing rarely marks vowel length (ා = a 97%, ී = i 58%). Length has to be recovered from context or a dictionary§2a, §2bH
RS-010Chat-style Singlish drops vowels heavily (nthi, mta, kynna). The published systems cope with dictionaries, fuzzy matching or BERT§2bH
RS-011ව: the formal systems use v. KNAB notes that local practice writes w, and informal writing prefers w in sentences (72%)§1e, §2aH
RS-012Prenasalized letters: ISO and the national system mark them with a breve (n̆d, m̆b, n̆g), KNAB with a middle dot (·mba), ALA-LC writes nda / ṃba / ṅga. Informal writing uses a plain cluster (nd 91%, mb 88%, ng 52%), which is ambiguous with න්ද, ම්බ, න්ග§1, §2aH
RS-014ං: ISO ṁ, national system ṅ, ALA-LC the class nasal, KNAB ṁ (ń). Informal writing uses n (91%), which can't be told apart from න්§1, §2aH
RS-017ෘ is written ru informally (93%, r 7%); ISO uses r̥. A romanization that writes ෘ as ru cannot distinguish කෘ from ක්‍රු (both kru)§1a, §2aH
RS-018ඏ ඐ ෟ ෳ have values only in the formal systems (ISO l̥ / l̥̄, national and ALA-LC ḷ / ḹ). There is no informal data for them§1, §2aH
RS-019Hal is not written. ISO has no symbol for ්, and informal writing leaves a consonant with no following vowel bare (hal implied, especially at the end of a word)§1a, §2aH
RS-020The inherent vowel is written as a in ISO and in informal writing (99% in Dakshina). Chat breaks this (RS-010)§1a, §2aH
RS-024ශ / ෂ: ISO ś / ṣ, national ś / sh, UN 1972 sh / ṣh, KNAB sha / ṣha. Informal writing uses sh for both (85% / 91%)§1, §2aH
RS-025The formal systems mark retroflex letters with an underdot (ṭ ḍ ṇ ḷ). Informal writing does not distinguish them at all for the nasal and lateral (ණ = n 100%, ළ = l 100%)§1a, §2aH
RS-026ai = ඓ and au = ඖ in ISO, and in informal writing (ai 99%, au 97%). The national system now uses ĩ for ඓ§1a, §1b, §2aH
RS-029The formal systems disagree with each other on the nasals. ISO: ං = ṁ, ඞ = ṅ. Sri Lanka 2018+: ං = ṅ, ඞ = ṁ. ALA-LC: ං takes the class nasal§1H
RS-030Keyboard schemes split from informal writing on d. The open Singlish keymap of the Wikimedia input tools (si-singlish, 2012) types d ඩ, dh ද, D ඪ, Dh ධ, the same pattern as t ට / th ත, and older keyboard schemes share it. Informal writing reads d as ද (RS-002), and ද is about 5 times as frequent as ඩ in running text (CC-12). A converter should default to d ද and offer the keyboard convention as an option§2a, SourcesH

5. Where conventions agree: recommendations for a phonetic romanization ​

  1. Use the values that formal and informal practice already share:
    • k g j p b m y r l s h f n as the base consonants
    • kh gh bh as aspirates, and dh = ධ
    • d = ද (ISO and informal agree)
    • a i u e o short vowels and ai = ඓ, au = ඖ
    • the inherent vowel written as a, hal left unwritten
  2. Where informal practice departs from ISO, decide and state it. The clearest cases:
    • th = ත and t = ට (informal) vs t / ṭ (ISO)
    • ch = ච (informal) vs c (ISO)
    • w (informal, KNAB's local practice) vs v (formal) for ව
    • sh for both ශ and ෂ (informal) vs ś / ṣ (formal)
  3. Mark vowel length explicitly, and avoid ee/oo for it. Use aa ii uu (ISO 7-bit) for ā ī ū. ee and oo mean ī/ū to informal writers but ē/ō in ISO 7-bit, so they are ambiguous.
  4. Use ae / aee for ඇ / ඈ, matching the ISO 7-bit form. Informal e/a collapses ඇ with එ and අ.
  5. Mark the prenasalized letters and anusvara explicitly if the romanization must round-trip. Plain nd, mb, ng and n are ambiguous with න්ද, ම්බ, න්ග and න්. Prior art: the ISO breve and KNAB's middle dot.
  6. Pick one convention for ං and ඞ and document it, since ISO and the national system swap ṁ and ṅ.
  7. Cover the gaps informal writing ignores: a distinction between ෘ and ්‍රු, and values for ඏ ඐ ෟ ෳ (ISO and ALA-LC are the only prior art).

6. Open questions ​

  1. How long-vowel and ඇ use differs between real chat text (social media) and the careful Dakshina romanizations. No chat-corpus counts per letter are published ❓.
  2. Is the Dakshina Hugging Face mirror identical to the official release? ❓
  3. The ALA-LC table is from interscript, not the LoC PDF. The 2011 ඇ value (ă vs æ) is not confirmed ❓.

Removed RS conventions ​

Removed: out of scope: RS-013, RS-015, RS-016, RS-021, RS-022, RS-023, RS-027, RS-028, RS-030, RS-031.


Sources ​