# Audio to Text Transcriber (AI)

Transcribe speech from audio (wav/mp3/m4a/flac/ogg/webm/aac) to text, SRT, VTT or JSON using the grok-stt AI model. Up to 10 minutes.

> Canonical page: https://elysiatools.com/en/tools/audio-to-text-transcriber

- **Category:** AI Tools

- **Keywords:** audio, transcribe, speech to text, stt, subtitle, srt, vtt, caption, grok, whisper, transcription

## Overview

Upload an audio or video file containing speech and the tool transcribes it with x-ai/grok-stt-1.0 through OpenRouter. Long inputs are split near silence into bounded FLAC chunks. Subtitle formats use native word or segment timestamps when the provider returns them; otherwise they use clearly labelled silence-aligned approximate timing. Supports WAV, MP3, M4A, FLAC, OGG, WebM, and AAC. Audio duration follows your plan limit (free tier: 10 min, enforced by the platform). Choose automatic detection or a language hint for better accuracy.

## Inputs

- **Audio File** (file): Upload an audio/video file with speech
- **Output Format** (select)
- **Language Hint** (select)
- **Maximum chunk duration (s)** (number)

## When to use

- Convert interviews, meetings, lectures, or voice recordings into searchable text.
- Create SRT or VTT subtitles from spoken audio for video content.
- Generate a structured JSON cue list when you need timestamped transcription data.

## How it works

- Upload a supported audio or video file containing speech, such as WAV, MP3, M4A, FLAC, OGG, WebM, or AAC.
- Choose Plain text, SRT subtitles, VTT subtitles, or JSON cue list as the output format.
- Select automatic language detection or provide a language hint for English, Chinese, Spanish, Russian, French, German, or Portuguese.
- For longer inputs, the tool splits audio near silence into bounded FLAC chunks before transcription.

## Use cases

- Transcribe interviews, calls, lectures, and recorded notes into plain text.
- Add SRT or VTT captions to short videos and spoken-content recordings.
- Create timestamped JSON cue data for reviewing or organizing spoken audio.

## Frequently asked questions

### What file formats are supported?

Supported audio formats include WAV, MP3, M4A, FLAC, OGG, WebM, and AAC. The tool also accepts supported WebM and MP4 video uploads containing speech.

### What output formats are available?

You can export plain text, SRT subtitles, VTT subtitles, or a JSON cue list.

### How long can the uploaded file be?

The free tier supports up to 10 minutes, with the applicable duration limit enforced by your plan.

### Can the tool detect the spoken language?

Yes. Choose Auto detect or provide a language hint to help guide transcription.

### Are subtitle timestamps exact?

SRT and VTT outputs use provider word or segment timestamps when available. Otherwise, they use clearly labelled silence-aligned approximate timing.

## Related tools

- [Audio Melody Contour Extractor](https://elysiatools.com/en/tools/audio-melody-contour-extractor): Extract a dominant melody path from audio and export MIDI, note events, a pitch contour, SVG, and JSON in one ZIP.
- [Audio Chord Progression Detector](https://elysiatools.com/en/tools/audio-chord-progression-detector): Estimate an audio file's chord sequence locally and download CSV, JSON, SVG, and method notes in one ZIP.
- [Batch Audio Converter](https://elysiatools.com/en/tools/audio-batch-converter): Convert multiple audio files between different formats like MP3, WAV, FLAC, AAC, OGG, Opus with quality settings and batch processing
- [Audio Codec Swap](https://elysiatools.com/en/tools/audio-codec-swap): Change the codec of an audio file within a chosen output format
- [Audio Format Converter](https://elysiatools.com/en/tools/audio-converter): Convert audio files between different formats like MP3, WAV, FLAC, AAC, OGG with quality settings and volume adjustment
- [Audio to Multitrack MIDI (Draft)](https://elysiatools.com/en/tools/audio-to-multitrack-midi): Split a full mix into stems (drums/bass/other/vocals) and transcribe each to MIDI — a multi-track starting point for transcription
- [Data URI Generator](https://elysiatools.com/en/tools/data-uri-generator): Convert files into Data URIs (Base64 or percent-encoded) for inlining images, fonts, and assets directly into HTML, CSS, or Markdown
- [Audio Declipper](https://elysiatools.com/en/tools/audio-declip): Repair clipped audio by reconstructing peaks that exceeded the maximum amplitude. Restore distorted audio from overdriven recordings or excessive gain

## Samples

- [Copyright-Free MP3 Audio Samples](https://elysiatools.com/en/samples/mp3-samples): Collection of royalty-free audio samples for testing and development purposes including nature sounds, meditation music, and ambient audio
- [Copyright-Free FLAC Audio Samples](https://elysiatools.com/en/samples/flac-samples): Lossless FLAC audio samples for testing and development, mirrored from MP3 set with nature sounds and meditation music
- [Copyright-Free WAV Audio Samples](https://elysiatools.com/en/samples/wav-samples): Uncompressed PCM WAV audio samples for testing and development, mirrored from MP3 set with nature sounds and meditation music
- [Video Samples](https://elysiatools.com/en/samples/video-samples): Sample video files in different formats and orientations
