How Transcription Works in 12scribe | 12scribe

How Transcription Works in 12scribe

From Speech to Searchable Text

Audio transcription in 12scribe converts spoken words into text that you can read, search, and navigate. The process uses AI-powered speech recognition to analyze your recording and produce a timestamped transcript. Once transcribed, every word in your recording becomes searchable — type a phrase into the search bar and jump to the exact second it was spoken.

Transcription is the foundation for many of 12scribe's most powerful features. AI Smart Notes analyze the transcript to generate structured markers at key moments. Global search indexes the transcript text across all your projects, letting you find any word in any recording instantly. Even casual review is improved — you can read the text instead of re-listening to audio, which is significantly faster.

The transcription process is fully automated. You tap one button (or enable auto-transcription), and the AI handles everything — splitting the audio, recognizing speech, aligning timestamps, and delivering the finished text back to your device.

Process Flow

Here is what happens step by step when you tap Transcribe or when auto-transcription triggers after a recording session:

  1. Upload — the audio file is sent from your device to the 12scribe server over your internet connection. Upload speed depends on file size and your connection quality. A typical one-hour lecture in .m4a format uploads in under a minute on a decent Wi-Fi connection.
  2. Chunking — the server splits your audio into smaller segments for parallel processing. This is the key to fast transcription — instead of processing one long file sequentially, multiple chunks are transcribed simultaneously by separate AI workers.
  3. Recognition — each chunk is processed by an AI speech-to-text model that converts audio waveforms into text. The model handles various accents, speaking speeds, technical vocabulary, and even multiple speakers within a single recording.
  4. Assembly — the transcribed chunks are merged back into a single continuous transcript. Timestamps from each chunk are aligned to your original recording timeline, ensuring that every word maps to the correct second in the audio.
  5. Delivery — the completed transcript is sent back to your device and stored locally inside the project. You can read it immediately, search through it, and use it as the basis for AI Smart Notes generation.
Tip: Transcription quality depends heavily on audio clarity. A quiet room with the speaker positioned near the microphone produces much better results than a noisy lecture hall with your phone buried in a bag. Even small improvements in recording conditions create noticeably better transcripts.

Language and Accuracy

12scribe supports multiple languages for transcription. You set the transcription language per project in the project settings screen. Choosing the correct language is important — the AI model optimizes its speech recognition for the selected language's phonetics, grammar patterns, and vocabulary.

Accuracy is typically 90-98% for clear audio in a quiet environment with a single speaker. Several factors affect how accurate the transcription will be:

  • Audio quality — clearer audio produces better results. An external microphone helps significantly in noisy environments like lecture halls or conference rooms.
  • Speaker clarity — normal conversational pace works best. Very fast speech, heavy accents, mumbling, or speaking away from the microphone all reduce accuracy.
  • Technical terms — highly specialized or domain-specific jargon may not be recognized perfectly on the first occurrence. Common terms in established fields (medicine, law, engineering) are generally well-handled.
  • Multiple speakers — the model handles alternating speakers well. Overlapping speech (two people talking simultaneously) reduces accuracy for the overlapping portion.
  • Background noise — steady background noise (air conditioning, traffic) is filtered reasonably well. Sudden loud noises (doors slamming, coughing) may cause brief gaps in recognition.

Processing Time

Transcription speed depends on recording length, audio complexity, and current server load. Here are typical processing times you can expect:

  • 15-minute recording — approximately 1 minute of processing
  • 1-hour recording — approximately 2-5 minutes of processing
  • 3-hour recording — approximately 8-12 minutes of processing

Pro users get priority processing, which means faster turnaround during peak usage hours when many users are submitting transcriptions simultaneously. During off-peak hours, both Free and Pro users experience similar speeds.

You can monitor transcription progress in real time from the project view, where a progress indicator shows the percentage complete. See Transcription Status: Tracking Progress for details on the status indicators and what to do if something goes wrong.

You do not need to keep the app open while transcription processes. The work happens entirely on the server, and you receive a push notification when the transcript is ready. If you are offline when you request transcription, the request is queued automatically and sent when you reconnect. Learn more in Offline Recording and Queued Transcription.

Tip: Enable auto-transcription in your project settings to skip the manual step entirely. Every recording session gets transcribed automatically the moment you tap Stop — by the time you pack up after a lecture, your transcript may already be ready to read.

Download for iOS

Start right now. 150 minutes free.

Try for free