TTS Input Tips
How you format the text in a TTS request affects pronunciation and pacing. If a word comes out wrong or the output audio doesn’t sound natural, try the tips below.
Spell words phonetically
If a word is mispronounced, replace it with a phonetic spelling that guides the model toward the right sound.
Write acronyms in capitals
If an acronym isn’t read correctly, write it in capital letters, separate the letters with periods, or both.
Write out numbers
Write numbers in words rather than numerals. This gives more stable results.
Use punctuation for pauses
Commas, dashes and ellipses change the rhythm and add natural pauses.
Break up long sentences
Long or complex sentences are harder to read clearly. Split them into shorter ones.
Avoid special symbols
Replace special symbols like ( ), @, #, ", $, etc., with words, or rephrase
the sentence without them.
Mark stresses
With the Ukrainian model, stress marks let you control where the stress falls in a word. This is useful for homographs or names where the model might otherwise guess wrong.
A stress mark is placed immediately before the vowel that should be stressed. You can mark stress in any of three equivalent ways:
- Stress token — insert a
<stress>token before the vowel, e.g.йог<stress>урт. - Plus sign — insert a
+before the vowel, inside the word, e.g.йог+урт. - Unicode combining accent — place a combining acute accent (U+0301) after the
vowel, e.g.
йогу́рт.
If a single mark doesn’t take effect, you can repeat it to push the model harder
toward placing the stress, e.g. йог<stress><stress>урт or йог++урт.
Repeated marks only increase how strongly the stress is enforced — they do not add a second or secondary stress. Start with a single mark and only add more if it isn’t honored. More than three repeats is not recommended.