Skip to content

Python and JavaScript packages ​

sinhala_orthography holds the letter inventory, the romanization converter and a frequency-based disambiguator. It needs Python 3.9 or later and has no dependencies.

Install ​

Clone the repository and import from src/, or install it:

sh
git clone https://github.com/SrilalS/Sinhala-Phonetic-Orthography
cd Sinhala-Phonetic-Orthography
pip install .

Windows consoles

Set PYTHONIOENCODING=utf-8 before printing Sinhala text.

Converting romanization ​

python
from sinhala_orthography import to_sinhala

to_sinhala("lankaava")                  # 'ලංකාව'
to_sinhala("kruura")                    # 'කෲර'
to_sinhala("kruura", rakaransaya_u=True)  # 'ක්‍රූර'
to_sinhala("lait")                      # 'ලයිට්'
to_sinhala("akShara", classical=True)   # 'අක්‍ෂර'
to_sinhala("karma", repaya_zwj=True)    # 'කර්‍ම'

to_sinhala(source, archaic=False, repaya_zwj=False, classical=False, rakaransaya_u=False) ​

Converts a romanized word or text to Sinhala script. Characters outside the sequence tables (spaces, digits, punctuation) pass through unchanged. Every output obeys the HARD rules of the rule set. This is checked over every input of up to three sequences (see Verification).

OptionEffectConvention
archaicAllow ඏ ඐ ෟ ෳ, standalone ඎ, ඁ, ඦ and touching lettersR-14
repaya_zwjWrite repaya as ර්‍ + C instead of plain ර් + CR-08
classicalZWJ conjuncts for the classical bandi akuru pairs, and rakaransaya after ම න ලR-07, R-10
rakaransaya_uWrite C + r + u/uu as rakaransaya + ු/ූ (ක්‍රූර) instead of the usual C + ෘ/ෲ (කෲර). The name reads "rakaransaya + u": the u is the vowel sign ු (rakaransayaU in JavaScript). Rakaransaya itself is always on; only ru / ruu changeR-06
retroflex_dd types ඩ and dh types ද (D ඪ, Dh ධ), the older keyboard conventionR-01

transliterate() is the same function under its original name. The sequence tables live in src/sinhala_orthography/data/ (consonants.json, vowels.json, specials.json), and the romanization is specified in 07 · Phonetic romanization. Try it in the playground.

Disambiguating with a word list ​

About 14 groups of letters sound alike (G-SP-01), so one romanization can't always pick the standard spelling. Lexicon indexes a word-frequency list by sound key, and candidates() ranks the words that share the input's key.

python
from sinhala_orthography import Lexicon, candidates

lex = Lexicon("word-frequency-list.tsv")   # one "word<TAB>count" per line
candidates(lex, "honda")                   # ['හොඳ', ...]
candidates(lex, "hon", partial=True)       # completions of a word prefix
FunctionReturns
candidates(lex, roman, limit=5, partial=False, **options)Ranked spellings. The converter's own spelling is always included. When the romanization has an explicit marker (a capital, a z prefix, a doubled vowel …) and that spelling is a word or a lone vowel letter (R ඍ, E ඓ), it comes first
sound_key(text)The text folded so sound-alike spellings match: aspirates → plain, ණ → න, ළ → ල, ශ/ෂ → ස, vowel length dropped, sanyaka → nasal + stop. sound_key("හොඳ") is 'හොන්ද'
normalize(word)Restores the mandatory ZWJ in C ් ය and C ් ර (never after ර), for word lists that dropped it
restyle(word, **options)A word in the style of the converter options (ක්‍රූර for කෲර with rakaransaya_u …). candidates() applies it, so the word list never undoes an option
Lexicon(path).count (word → frequency), .exact(key), .prefix(key, limit=200)

The evaluation in 07 used the University of Moratuwa NLPC Word Frequency List for Sinhala, which isn't redistributed here.

Letter inventory ​

python
from sinhala_orthography import VOWELS, CONSONANTS, SIGNS, entries

len(entries())   # 923
entries()[30]    # {'id': 'ka.ru', 'kind': 'syllable', 'text': 'කෘ', 'consonant': 'ka',
                 #  'vowel': 'ru', 'codepoints': ['U+0D9A', 'U+0DD8']}
NameWhat
VOWELS(id, independent letter, dependent sign) for hal, the inherent a and the 17 signs
CONSONANTS(id, letter) for the 41 consonants in Unicode order
SIGNSං ඃ ඁ
entries()Every letter form, as in data/letters.json

JavaScript and TypeScript ​

The same package is available for JavaScript and TypeScript, for Node.js and the browser, with no dependencies. It lives in js/ and reads the same sequence tables. A golden test checks that it gives exactly the Python output for more than 50,000 inputs, and the playground runs on it. It is published on npm as sinhala-orthography:

sh
npm install sinhala-orthography
js
import { toSinhala, Lexicon, candidates } from "sinhala-orthography";

toSinhala("kruura");                          // 'කෲර'
toSinhala("kruura", { rakaransayaU: true });  // 'ක්‍රූර'
const lex = new Lexicon(wordListText);        // "word<TAB>count" lines, or [word, count] pairs
candidates(lex, "honda", { limit: 5 });
PythonJavaScript
to_sinhala(s, repaya_zwj=True)toSinhala(s, { repayaZwj: true }) (also archaic, classical, rakaransayaU, retroflexD)
Lexicon(path)new Lexicon(text) or new Lexicon(pairs): no file access, so it also runs in a browser
candidates(lex, roman, limit=5, partial=False)candidates(lex, roman, { limit, partial })
sound_key, normalize, entries, VOWELS, CONSONANTS, SIGNSsoundKey, normalize, entries, VOWELS, CONSONANTS, SIGNS

The Python package stays the reference: change it first, run python tools/build_js_golden.py, then make npm test in js/ pass.

Tools ​

sh
python tools/build_data.py            # data/letters.json, data/validity.json
python tools/check_romanization.py    # coverage + exhaustive safety (about 20 s)
python tools/build_spec_tables.py     # tables in docs/07
python tools/check_translations.py    # Sinhala translations whose English source changed
python -m unittest discover tests     # set SINHALA_WORD_LIST to also run the lexicon tests