SpeechEvalPro

1
5 0 Reviews 1 Saved
Introduction: SpeechEvalPro provides pronunciation assessment and scoring API solutions. Using a proprietary educational voice AI model, it integrates speech recognition and voice evaluation technologies to deliver high-quality, multi-dimensional pronunciation scoring for Chinese and English. The platform enables developers to build intelligent learning products that support human-computer interaction.

SpeechEvalPro Product Information

What is SpeechEvalPro?

SpeechEvalPro is a platform offering pronunciation assessment and scoring API solutions. It utilizes an independently researched and developed educational voice AI model, integrating voice evaluation, speech recognition, and other core technologies to provide high-quality, multi-dimensional Chinese and English pronunciation evaluation APIs. It helps customers create intelligent learning products for human-computer interaction.

How to use SpeechEvalPro?

Access the SpeechEvalPro API using HTTP or WebSocket protocols. Submit audio files in the recommended formats (16-bit sample size, 16K sample rate, 1 channel; supported formats include opus_raw, pcm, wav, and mp3). Consult the API documentation for specific details regarding question types, time limits, and text length restrictions.

SpeechEvalPro's Core Features

  • Pronunciation assessment and scoring API
  • Voice evaluation
  • Speech recognition
  • Multi-dimensional Chinese and English pronunciation evaluation
  • Support HTTP and WebSocket protocols

SpeechEvalPro Use Cases

#1 Intelligent learning products for human-computer interaction
#2 Homework assessment
#3 Examination assessment

FAQ from SpeechEvalPro

Is an SDK available? +

An SDK is not currently available. You can access the service directly via WebAPI, which offers lightweight, cross-platform, and streaming capabilities.

What audio formats are supported for pronunciation evaluation? +

We recommend using a 16-bit sample size, 16K sample rate, and 1 channel. Supported formats include opus_raw, pcm, wav, and mp3. Using other audio formats may negatively impact scoring accuracy.

What question types are supported, and what are the time and text length restrictions? +

Phoneme and word modes support up to 20 seconds. Sentence mode supports up to 40 seconds and less than 300 characters. Chapter (paragraph) mode supports up to 300 seconds and less than 10,000 characters. Please refer to the documentation for specific differences between Chinese and English language restrictions.

SpeechEvalPro Pricing

Free

$0

Free plan available.

Related Model Comparison Pages

Use these comparison pages to understand the trade-offs between the models most relevant to SpeechEvalPro.

Compare Gemini 1.0 Pro Deprecated and Gemini 2.0 Flash across pricing, context window, capabilities, benchmarks, and API access to choose the better fit for long-context workloads versus long-context workloads.

Compare Gemini 2.0 Flash Lite and Gemini 2.0 Flash across pricing, context window, capabilities, benchmarks, and API access to choose the better fit for long-context workloads versus long-context workloads.

Compare Gemini 1.0 Pro Deprecated and Gemini 1.5 Flash Deprecated across pricing, context window, capabilities, benchmarks, and API access to choose the better fit for long-context workloads versus general-purpose AI workloads.