Subjective Evaluation Methods for Thai TTS Models

Evaluating Thai Text-to-Speech: A Human-Centric Approach

Introduction: The Necessity of the Human Touch

As Text-to-Speech (TTS) technology advances, the challenge lies not just in generating speech, but in assessing its quality. While automated tools exist, they often fail to capture the nuances of human judgment—especially in tonal languages like Thai.

To bridge this gap, we have developed a specialized TTS Evaluation Platform. Our focus is on subjective preference evaluation, as human listeners remain the gold standard for assessing systems such as agent assistants and audiobooks. Even if human scoring varies, it provides the most reliable reflection of real-world expectations.

Mean Opinion Score (MOS)

The Mean Opinion Score (MOS) is a common method in speech synthesis evaluation. This process relies on human evaluators to provide subjective qualitative judgments. It utilizes a 5-point scale (1: Poor to 5: Excellent) to rate specific dimensions of audio. Multi-dimensional aspects are included because a single score cannot capture the specific strengths and weaknesses of a speech model; for instance, a voice may have perfect pronunciation yet sound robotic due to poor rhythm.

To ensure consistency among evaluators, our platform provides standardized definitions and scoring criteria for each of the four key dimensions:

A comprehensive table detailing the standardized scoring criteria for evaluating Thai Text-to-Speech models across four dimensions: Naturalness, Pronunciation, Tones, and Rhythm.
Table 1: Standardized definitions and scoring criteria for each of the four key dimensions

By breaking evaluation into these attributes, we can identify exactly where a model struggles. For instance, a model might achieve a 5 in pronunciation but only a 2 in rhythm, signaling a need for better spacing between words. Ultimately, the scores from every evaluator are averaged across each specific dimension to calculate the final Mean Opinion Score (MOS) for the model. This statistical average provides a balanced representation of the model's overall quality.

The Elo Rating System

While MOS tells us how "good" a model is, it can be difficult to choose between two high-performing models. For this, we use Comparative MOS (CMOS) powered by the Elo Rating System. The Elo Rating System was originally designed for ranking chess players. It is a mathematical method for calculating the relative skill levels of competitors.

In our implementation, two models are paired against each other using the same text. A human evaluator chooses which model performed better in a specific dimension. If a lower-rated model "defeats" a higher-rated one, it gains a significant number of points, while the leader loses an equal amount. This creates a self-correcting leaderboard where the most "human-like" models naturally rise to the top through direct competition.

Our Evaluation Platform: Streamlined and Unbiased

To facilitate this research, we built a web-based application designed for ease of use and scientific integrity. Key features include:

  • Dual Evaluation Modes: Users can perform traditional MOS scoring or head-to-head Elo comparisons.
  • Bias Reduction: Models are anonymized (e.g., "Audio 1" vs. "Audio 2") to ensure evaluators judge the sound, not the brand.
  • Scalability: When a new model is developed, it can be integrated into the Elo system immediately without needing to re-evaluate every previous model from scratch.
  • Data Export: Users can download raw scores for in-depth statistical analysis.
The Elo evaluation interface demonstrates the platform's bias-reduction strategy by anonymizing model names as 'Audio 1' and 'Audio 2' during the assessment.
Figure 1: A Snippet of the Evaluation Platform using Elo System
The overview of text-to-speech evaluation platform consists of the system loads text and audio files, user selection of text and model to evaluate, functionality of elo and mos evaluation method, and saved final result.
Figure 2: Overview Flowchart of Evaluation Platform

Comparative Performance Results

The evaluation was performed by two in-team evaluators across 6 models and 20 text samples, resulting in a total of 120 assessed audio clips. The following bar chart presents the Mean Opinion Scores (MOS) across four critical dimensions. Scores range from 1.0 (Poor) to 5.0 (Excellent):

A bar chart presenting the detailed results of the subjective human evaluation for six TTS models across four specific dimensions: Naturalness, Pronunciation, Tone, and Rhythm. The table provides precise numerical scores for each service, highlighting ElevenLabs as the highest performer in all categories.
Figure 3: Comprehensive MOS Results

The next following bar chart presents the Average Mean Opinion Score (MOS). This chart ranks the models from highest to lowest overall performance, making it easy to see the clear lead ElevenLabs and Cartesia have over the standard cloud services:

A horizontal bar chart displaying the overall ranking of six Text-to-Speech models based on their average Mean Opinion Score (MOS). The chart ranks models from highest to lowest performance, with ElevenLabs and Cartesia leading the group, followed by Google Chirp 3 HD, Botnoi 3.0, Microsoft, and Google Standard.
Figure 4: Overall ranking of TTS model

The following table presents a side-by-side analysis of Time to First Byte (TTFB) latency, and Estimated Cost per million characters, providing the context necessary to select an optimal model. As highlighted, Google Chirp 3 HD emerges as the strongest candidate for low-latency requirements at a competitive price point.

A comparison table of 6 TTS models highlighting technical metrics. Google Chirp 3 HD is highlighted as the top performer with the lowest latency of 428ms at a cost of ~$30/1M characters.
Table 2: Latency and Cost Analysis

Key Insights from the Evaluation

The Industry Leader: ElevenLabs emerged as the top performer in every category. It achieved near-perfect scores in Pronunciation (4.90) and Tone (4.95), suggesting it is exceptionally capable of navigating the complex pitch requirements of Thai speech.

The Challenge of Rhythm: Across all models, Rhythm consistently received the lowest scores. Even the highest-rated models struggled to maintain a natural flow, often pausing in unnatural places or maintaining an inconsistent speed.

Tonal Accuracy vs. Naturalness: Interestingly, several models (such as Microsoft and Google Standard) scored quite well in Tone and Pronunciation but significantly lower in Naturalness. This indicates that while they can speak individual words correctly, they still sound noticeably "robotic" to native listeners.

The Latency Champion: Google Chirp 3 HD emerged as the fastest for real-time applications, boasting the lowest TTFB at 428ms. This makes it the most viable option for low-latency, interactive AI voice agents.

The High-Value Middle Ground: Google Chirp 3 HD presents the best balance of performance and cost. At $30/1M chars, it provides high-quality speech and superior latency at a fraction of the cost of premium tiers like ElevenLabs or Cartesia.

Conclusion: Making Evaluation Accessible

Our platform aims to transform TTS assessment from a technical hurdle into a streamlined, understandable process. By combining the granular detail of MOS with the competitive ranking of Elo, we provide a comprehensive toolkit for developing Thai speech synthesis that sounds less like a machine and more like a neighbor.

Collaborate and partner with our AI Lab at Amity Solutions here

Collaborate and partner with our AI Research & Application Center

Partner with Amity’s AI Lab to co-develop AI solutions through collaboration opportunities.

Collaborate with Us