speech-to-text

Speech-to-text is technology that converts spoken audio into written text, also known as automatic speech recognition.

5 articles
Last mentioned

Speech-to-text is technology that analyzes recorded or live audio and transcribes it into text. It is also called automatic speech recognition (ASR) and works in the opposite direction from text-to-speech, which turns text into audio.

Speech-to-text systems mainly rely on neural network models trained on large amounts of audio paired with transcripts. The technology is used for meeting notes, captions, voice assistants and voice commands, and some models also return word-level timestamps.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

Hold a push-to-talk key, speak, and speech-to-text turns the command into text.

…de acts as the project manager, and it hands each job to a specialist: a speech-to-text engine finds the clean takes, a rendering engine builds the motion graph…

Dictation Fixer, as an honorable mention: cleans up speech-to-text errors before the text reaches Claude

When the model detects speech, or the wearer activates Live Rewind, an encrypted 15-second snippet goes to the paired iPhone for speech-to-text processing.

For speech-to-text, Whisper ran through MLX, a framework used here to run models on Apple Silicon.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.