How transcripts are made
Sutam uses automatic speech recognition to generate transcripts from the Dharma Seed archive. Every transcript is generated with Whisper's large-v3-turbo model and a glossary of Pali and Sanskrit terms.
Automatic transcripts are a starting point, not the final word. We mark every transcript as automated so you know to listen for context and accuracy. The audio is always the authoritative version.
You can suggest fixes using the "suggest fix" link on each talk page. We review suggestions and update transcripts based on corrections that match the audio more closely.
For technical details on the transcription pipeline, see the project repository.