speech-to-text
Speech-to-text is technology that converts spoken audio into written text, also known as automatic speech recognition.
Speech-to-text is technology that analyzes recorded or live audio and transcribes it into text. It is also called automatic speech recognition (ASR) and works in the opposite direction from text-to-speech, which turns text into audio.
Speech-to-text systems mainly rely on neural network models trained on large amounts of audio paired with transcripts. The technology is used for meeting notes, captions, voice assistants and voice commands, and some models also return word-level timestamps.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
…de acts as the project manager, and it hands each job to a specialist: a speech-to-text engine finds the clean takes, a rendering engine builds the motion graph…
Dictation Fixer, as an honorable mention: cleans up speech-to-text errors before the text reaches Claude
When the model detects speech, or the wearer activates Live Rewind, an encrypted 15-second snippet goes to the paired iPhone for speech-to-text processing.
For speech-to-text, Whisper ran through MLX, a framework used here to run models on Apple Silicon.