RU-RZ/NT 1.0
Standard for romanized and native-script text styles for russian language.
This document contains many rare Unicode characters, whose name is not mentioned. If you need to know the Unicode name of some character, you can copy the character from this document to some website that identifies Unicode characters.
Main versions of writing styles:
RU-RZ — fluent romanized russian:
 
RU-RZ-P Precise romanized russian. Differentiates the pronunciations and literal spelling preferences a/ȧ, ă/ya/ặ, e/ĕ/ye/ė/ŏ/yo, o/ọ, ŭ/yu, v/ṽ, g/ḡ, d/đ, j/ī, l/ł, t/ŧ and ḫ/h. Abbreviates letter combination ij as í (which is not done in any other of these standards). Foreign non-cyrillic proper nouns are written literally (letter by letter), not based on pronunciation, as in RU-RZ-S. In foreign proper nouns uses letters that are outside of the basic russian alphabet, such as C, H, Q, W, X, Ä, Å, Ö or Ü. (RU-RZ-S imitates the selection of letters that is available in the basic russian alphabet, and refrains from using any letters outside of it.)
 
RU-RZ-S Simple romanized russian. Does not differentiate any of the above-mentioned pronunciations or literal spelling preferences. Differentiating the letters е and ё is supported but not required by this standard. Any ordinary russian text can be automatically converted into RU-RZ-S text (and typically does not differentiate letters е and ё). Foreign non-cyrillic proper nouns can be written either literally (letter by letter) or based on pronunciation.
 
RU-NT — russian with the native script:
 
RU-NT-P Precise russian with cyrillic script. Differentiates all the above-mentioned pronunciations and literal spelling preferences. Does not abbreviate "ij", as RU-RZ-P does. Foreign non-cyrillic proper nouns are written literally (letter by letter), not based on pronunciation, as in RU-NT-S. In foreign proper nouns uses letters that are outside of the basic russian alphabet, such as the equivalents of C, H, Q, W, X, Ä, Å, Ö or Ü. (RU-NT-S imitates the selection of letters that is available in the basic russian alphabet, and refrains from using any letters outside of it.)
 
RU-NT-S Standard russian with cyrillic script. Does not differentiate any of the above-mentioned pronunciations or literal spelling preferences. Differentiating the letters е and ё is supported but not required by this standard. Foreign non-cyrillic proper nouns are written based on pronunciation.
 
Foreign non-cyrillic proper nouns are usually written based on pronunciation in the cyrillic script. The russian style of writing foreign proper nouns is imitated also in such of these romanized standards, whose primary emphasis is replicating the cyrillic text as such in latin script, rather than writing optimally convenient text with the latin script.
Only such latin and cyrillic Unicode characters have been deemed acceptable for these standards, which do not force the text row to be any higher than normal. Sometimes a cyrillic text standard uses a latin Unicode character, even if a similar character were possible to achieve as a combination of a cyrillic letter (visually identical to the latin letter) and a separate diacritical mark. Single Unicode characters are always favoured, because they produce the expected visual look more reliably and precisely in various software.
Sample text, for comparing these text styles:
 
RU-RZ-P Janice [{Dženis}] pọznakomilsă s Yevgeniĕm desặtᶥ let nazad. Ọna vidėla čto u nĕḡo dobroĕ serđċe.
 
RU-RZ-S Dženis {[Janice]} poznakomilsă s Еvgeniem desătᶥ let nazad. Ona videla čto u nego dobroe serdċe.
 
RU-NT-P Йаниџе [{Дженис}] пọзнакомился с Êвгениӗм деся̇ть лет назад. Ọна видėла что у нӗѓо доброӗ серд̄це.
 
RU-NT-S Дженис {[Жанице]} познакомился с Евгением десять лет назад. Она видела что у него доброе сердце.
Logical convertibility between these text standards (based on the text itself only, without using any dictionary data or human help):
 
  RZ-P RZ-S NT-P NT-S
RZ-P   ➔   YES YES YES
RZ-S   ➔   YES
NT-P   ➔ YES YES   YES
NT-S   ➔ YES  
 
A dash — or arrow → is used in this diagram, if the conversion between two text styles is not supported, because the source text does not contain enough information for logically concluding the correct ortography in the target text style (in all possible scenarios). If such an unsupported conversion is requested, the default option is not to convert the source text at all. However, if it is deemed preferable to convert the text into the closest possible text style (most notably, when the script should change from native to romanized, or vice versa), the arrows point towards the text style that is the recommendable substitute for the unsupported conversion. A dash — is used in the diagram, when no other logically supported substitutes are available (in the same script) than the source text style itself.
Foreign proper nouns cannot be automatically converted between a literal format (replicated letter by letter) and a transliteration based on pronunciation. Such a conversion would be reliably possible only if both the literal and the pronounced form of the name are documented in the source text, using some kind of tags or footnotes. This table of logical convertibility between the text styles ignores this aspect of foreign proper nouns, and promises convertibility from a text style to another, if no other logical obstacles exist for the conversion than the writing of foreign proper nouns being based on different principles.
The sample texts afore use such a notation that the form based on pronunciation is given in [{curly brackets inside square brackets}], and the literal form is given in {[square brackets inside curly brackets]}, after the spelling that is chosen for the main text. Thus it would be possible to automatically recognize, which of the two formats is the literal one. These codes and alternative spellings are not intended to be seen by the human reader in the main text.
These text styles use a strict logical correlation between latin letters and cyrillic script letters, so that the text can be converted back and forth between cyrillic script and latin script, and the text should stay exactly similar through all these conversions, without any changes caused by the conversion process back and forth. However, foreign proper nouns are not always fully compatible with this system. In some scenarios it is possible that converting a foreign proper noun from latin script to cyrillic script, and then back into latin script, produces a different spelling in latin script than was the original form of the name. This can happen because these romanized cyrillic text styles are optimized for fluent reading and strict logical compatibility with the cyrillic script, not strict compatibility with the way how other languages use the latin script.
Transliterated RU-RZ-P standard uses alternative esthetic spelling preferences ă/ya, ĕ/ye, ŏ/yo and ŭ/yu, whose purpose is to make the text more stylish and convenient to read. While this feature was originally not intended for text that is written in cyrillic script, it is included in RU-NT-P standard (mostly as diacritical marks over or under vowels), to make this standard fully convertible to RU-RZ-P format without the loss of any information.
Technical reliability of these text styles, as Unicode characters:
To increase the grammatical information content of the text, these text styles use many uncommon diactirical marks and special characters, both in latin and in cyrillic script. This causes a higher risk that some fonts or software will fail to display the text correctly and beautifully. One of the esthetic risks is that some letters in the text are displayed with a different font than the rest of text. This happens if the primary font does not include some rare character. In that case the software will use any other font that contains the character. If none of the available fonts contains the character, then the software probably displays some generic character, such as a square or a question mark, for example.
Below is a guide for transliterating foreign proper nouns from latin script to russian cyrillic script literally, letter by letter — regardless of the language, or how the word is pronounced. The standard RU-NT-P uses the primary variant only, which is not in parentheses. Some ambiguity and taking the pronunciation into consideration is allowed in standard RU-NT-S, whose acceptable alternatives are listed in parentheses, using the most basic russian alphabet only.
 
LAT CYR LAT CYR
A a А а P p П п
B b Б б Q q Қ қ (К к)
C c Џ џ (Ц ц / С с / К к) R r Р р
D d Д д S s С с
E e Е е (Е е / Э э) T t Т т
F f Ф ф U u У у
G g Г г (Г г / Ж ж) V v В в
H h Ӽ ӽ (Х х) W w Ў ў (В в / У у)
I i И и X x Ӿ ӿ (КС Кс кс)
J j Й й (Й й / Ж ж) Y y Ѝ ѝ (Й й / И и / Ы ы / УЕ Уе уе)
K k К к Z z З з (З з / Ц ц)
L l Л л Ä ä Ӓ ӓ (А а / Я я / АЕ Aе ае)
M m М м Å å Å å (О о / АA Aа аа / А а)
N n Н н Ö ö Ӧ ӧ (О о / ОЕ Ое ое)
O o О о Ü ü Ӱ ӱ (Ы ы / У у / УЕ Уе уе)
CYR LAT SPECIAL CASES
А а A a => Ȧ ȧ, when pronounced "i": dvėnadċȧtᶥ.
— In RU-NT-P: Ȧ ȧ => (no conversion: use these latin characters).
Б б B b  
В в V v => Ṽ ṽ, when left unpronounced by most people, for example in letter combination vstv.
— In RU-NT-P: Ṽ ṽ => Ɓ ᴃ (these are latin characters).
Г г G g => Ḡ ḡ, in the ending -ḡo, which is pronounced "-vo": ĕḡo.
— In RU-NT-P: Ḡ ḡ => Ѓ ѓ.
Д д D d => Đ đ, when left unpronounced by most people, for example in letter combinations ndsk, zdn or rdċ.
— In RU-NT-P: Đ đ => Д̣  д̣  – using combined characters: letter Д д and a separate character "combining dot below". => Ԭ ԭ. A secondary character, which is also supported but not preferred by the text conversion algorithm.
Е е E e => Ĕ ĕ, when pronounced "ye": ĕsli, nĕ, načalĕ.
=> ye, when pronounced "ye" immediately after ă, ĕ, ŏ or ŭ: gulăĕm => gulăyem. In proper nouns the preferred transliteration might be "ye" also in other situations (depending on personal esthetic preferences): Ĕlena => Yelena, Fedĕnka => Fedyenka.
NOTE: ĕ and ye are synonyms: they mean the same sound, and behave similarly if text is converted into RU-NT-S.
=> Ė ė, when pronounced "i": smọtrėli, vidėl, iŝėtĕ, mėnă, pọnėdelᶥnik.
— In RU-NT-P:
Ĕ ĕ => Ӗ ӗ (this is a conversion from latin script to cyrillic script, even if the characters look identical).
YE/Ye ye => Ẽ ẽ (these are latin characters).
Ė ė => (no conversion: use these latin characters).
Ё ё Ŏ ŏ => yo, as the first letter of a word, or when immediately after ă, ĕ, ŏ or ŭ: ŏlka => yolka, ĕŏ => ĕyo. In proper nouns the preferred transliteration might be "yo" also in other situations (depending on personal esthetic preferences): Alŏna => Alyona, Artŏm => Artyom.
NOTE: ŏ and yo are synonyms: they mean the same sound, and behave similarly if text is converted into RU-NT-S.
— In RU-NT-P, ONLY in cases of personal preference, when not the first letter of a word, and not immediately after ă, ĕ, ŏ or ŭ: YO/Yo yo => Ê ê (these are latin characters). => Ȅ ȅ (these are latin characters): a secondary character, which is also supported but not preferred by the text conversion algorithm.
NOTE: It is very common in cyrillic texts to misspell ё as е. The standards RU-NT-S and RU-RZ-S accept such ambiguity, but any other of these standards do not accept it.
Ж ж Ž ž  
З з Z z  
И и I i => Í í is an abbreviation for IJ Ij ij / ИЙ Ий ий, which is used in RU-RZ-P only, and only at the end of a word, not in other positions: sinij => siní / Ijsselmeer, Rijkaard.
NOTE: í and ij are synonyms: they mean the same sound and letter combination, and behave similarly if text is converted into RU-NT-S.
Й й J j => Ī ī, when left unpronounced by most people, for example in word pọžaluīsta.
— In RU-NT-P: Ī ī => Ӣ ӣ.
К к K k  
Л л L l => Ł ł, when left unpronounced by most people, for example in letter combination lnċ.
— In RU-NT-P: Ł ł => Ԯ ԯ.
М м M m  
Н н N n  
О о O o => Ọ ọ, when prononuced nearly similarly as "a": ḫọrọšo.
— In RU-NT-P: Ọ ọ => (no conversion: use these latin characters).
П п P p  
Р р R r  
С с S s  
Т т T t => Ŧ ŧ, when left unpronounced by most people, for example in letter combinations ntsk, ntst, stsk or stn, possibly also stl.
— In RU-NT-P: Ŧ ŧ => Ŧ ҭ (no conversion in uppercase).
У у U u  
Ф ф F f  
Х х Ḫ ḫ  
Ц ц Ċ ċ  
Ч ч Č č  
Ш ш Š š  
Щ щ Ŝ ŝ  
Ъ ъ ˺   ·  
Ы ы Ý ý  
Ь ь ᴵ   ᶥ  
Э э Ḙ ḙ  
Ю ю Ŭ ŭ => yu, as the first letter of a word, or immediately after ă, ĕ, ŏ or ŭ: Ŭrí => Yurí, gulăŭŝij => gulăyuŝij. In proper nouns the preferred transliteration might be "yu" also in other situations (depending on personal esthetic preferences): Katŭška => Katyuška.
NOTE: ŭ and yu are synonyms: they mean the same sound, and behave similarly if text is converted into RU-NT-S.
— In RU-NT-P, ONLY in cases of personal preference, when not the first letter of a word, and not immediately after ă, ĕ, ŏ or ŭ: YU/Yu yu => Ю̣ ю̣  – using combined characters: letter Ю ю and a separate character "combining dot below". => Ю̱ ю̱  – a secondary character, "combining macron below", which is also supported but not preferred by the text conversion algorithm.
Я я Ă ă => ya, as the first letter of a word, or immediately after ă, ĕ, ŏ or ŭ: ă => ya, Ăna => Yana, ăblọkọ => yablọkọ, sinăă => sinăya. In proper nouns the preferred transliteration might be "ya" also in other situations (depending on personal esthetic preferences): Ženă => Ženya.
NOTE: ă and ya are synonyms: they mean the same sound, and behave similarly if text is converted into RU-NT-S.
=> Ặ ặ, when pronounced "i": desặtᶥ, verặŝiĕ, svặŝennik, prặmoj.
— In RU-NT-P, ONLY in cases of personal preference, when not the first letter of a word, and not immediately after ă, ĕ, ŏ or ŭ:
YA/Ya ya => Я̣ я̣  – using combined characters: letter Я я + separate character "combining dot below". => Я̱ я̱  – a secondary character, "combining macron below", which is also supported but not preferred by the text conversion algorithm.
— In RU-NT-P:
Ặ ặ => Я̉ я̇ , using combined characters: letter Я я + separate character "combining hook above" (uppercase) or "combining dot above" (lowercase).
RU-RZ/NT 1.0 standard for romanized and native-script text styles for russian language. Ion Mittler, 29 March 2024. Released in the public domain under CC0-1.0 license (Creative Commons 0 version 1.0). http://creativecommons.org/publicdomain/ zero/1.0/