Turn-level speaker labels
Neutral markers separate one participant’s words from the next participant’s words. They help readers follow a question-and-answer exchange without claiming to know a person’s name, role, or identity from the voice alone.
Live transcript
Transcript ready to review
Speaker A
Thanks for joining. Let us start with the interview questions.
Speaker B
当然,我已经准备好了。The first topic is the launch timeline.
Speaker A
Great. I will keep the transcript open while we talk.
Speaker B
Perfect. We can review the final text before sharing it.
This process answers “who spoke when” by dividing audio into speaker turns. The prepared transcript above lets you inspect a concrete two-speaker layout with timestamps and export controls. It is an interface demonstration: live microphone mode on this page labels speech as “You” and does not automatically identify multiple participants.
Speaker diarization groups segments by distinct voices and assigns neutral labels such as Speaker A and Speaker B. It is different from speech-to-text, which determines what was said, and from speaker identification, which attaches a real identity to a known voice. Use diarization labels as a review aid, then verify turn boundaries and names against the original audio.
Muse Voice Transcribe is an independent browser tool and is not affiliated with Meta. The two-speaker sample is prepared product content, not the result of a claimed live diarization model on this page.
Last updated:
The workbench opens with a fixed 21-second, four-turn transcript. Speaker A starts the interview, Speaker B replies with a mixed-language sentence, and both speakers take one additional turn. Each segment has a neutral label and an elapsed timestamp. Copy or download the TXT output to see how those labels survive outside the interface.
Speaker A [00:04]: Thanks for joining. Let us start with the interview questions. Speaker B [00:09]: 当然,我已经准备好了。The first topic is the launch timeline. Speaker A [00:15]: Great. I will keep the transcript open while we talk. Speaker B [00:21]: Perfect. We can review the final text before sharing it.
This example proves the display and export format only. It does not prove automatic speaker count estimation, voice matching, or a diarization accuracy score. If you reset the sample and start the browser microphone, the live result uses the single label “You.” That boundary prevents a prepared transcript from being mistaken for a live model result.
| Fact | Value |
|---|---|
| Prepared transcript | 4 turns / 21 seconds |
| Prepared speaker labels | 2: Speaker A and Speaker B |
| Live speaker label | You |
| Timestamp granularity | 1 timestamp per turn |
| Export | TXT only |
| Diarization accuracy | Not measured |
| Speaker identification | Not supported |
Source: the visible workbench and prepared transcript on this page. Verified September 3, 2026.
These terms describe different jobs. Use the table to keep a transcript format from being mistaken for a speaker identity feature.
| Concept | Question answered | Current page status |
|---|---|---|
| Speech-to-text | What was said? | Live microphone draft |
| Speaker diarization | Who spoke when? | Prepared Speaker A/B example |
| Speaker identification | Which known person spoke? | Not supported |
Good diarization makes a conversation easier to navigate, but each label still needs context and review. These are the practical benefits illustrated by the sample.
Neutral markers separate one participant’s words from the next participant’s words. They help readers follow a question-and-answer exchange without claiming to know a person’s name, role, or identity from the voice alone.
Elapsed times create reference points for returning to a decision, quote, or disputed boundary in the source audio. The sample uses one timestamp per turn; it does not provide word-level timing or a synchronized media player.
Copy and TXT download preserve each visible label before the line of dialogue. That makes the draft easy to move into research notes, an editor, or a meeting document while keeping speaker turns distinct.
The example exposes every label and segment directly. Reviewers can merge a split turn, divide a combined turn, or replace Speaker A with a verified name in another editor instead of relying on hidden speaker identity assumptions.
A speaker diarization result is best treated as structured draft data. A careful review focuses on boundaries first and identities second.
Listen to the source and note how many people actually speak. Background television, audience reactions, and the same person changing distance or tone can be mistaken for additional voices. Do not equate the highest label count with the participant count without checking.
Compare the transcript near interruptions, short acknowledgements, and overlapping speech. Those are common places for a phrase to be attached to the wrong participant or for one turn to be split unnecessarily. Timestamps make the suspicious moments easier to revisit.
Only replace Speaker A or Speaker B when the audio or meeting context confirms the identity. Voice clustering is not proof of a legal identity. Keep neutral labels when the speaker cannot be established confidently or when anonymity is part of the workflow.
Fix names, numbers, quotations, and technical terms as well as speaker assignments. Then copy or download the transcript. Preserve the original audio separately when an audit trail matters, because the sample workbench exports text but does not package audio evidence.
Diarization is most valuable when several voices contribute meaningfully and readers need to attribute statements without replaying the whole recording.
Separate interviewer questions from participant answers so researchers can scan themes and extract candidate quotes. Consent, accurate identity mapping, and quote verification remain separate responsibilities; neutral labels alone do not establish who a person is.
Follow decisions, objections, and handoffs across several participants. Short overlaps and similar voices can confuse a model, so review action owners and commitments against the recording before placing them in minutes or a customer record.
Use turn labels as a first pass for editing, show notes, or quote selection. Music, inserted clips, remote audio quality, and guests speaking at once can create extra clusters or incorrect switches that require an editor’s correction.
A clean-looking label can still be wrong. Evaluate the source, the processing path, and the intended use before trusting an attribution.
Speaker A means that segments were grouped as one voice; it does not reveal a name. Attaching an identity requires verified context or an explicitly enrolled speaker-recognition system. Avoid presenting a neutral cluster as proof that a particular person made a statement.
Interruptions, laughter, “yes” or “right,” crosstalk, and two people speaking together can produce missed or misplaced turns. Similar voices and one person moving around a room can also cause labels to merge or split unexpectedly.
The prepared transcript displays two speakers, but the live microphone mode uses only “You.” This page does not claim automatic live diarization, a maximum supported speaker count, file upload, or benchmark accuracy. Evaluate those capabilities only when a real processing implementation exposes them.
Multi-person audio may contain confidential statements and biometric voice information. Obtain consent, control access, and understand where processing occurs. This browser page does not claim universal on-device handling or provide a storage-retention control for uploaded recordings.
The live workflow depends on browser APIs. These public specifications explain the permission and recognition interfaces referenced on this page.
Speaker diarization is the process of dividing audio into segments according to who spoke when. Its output normally uses neutral clusters such as Speaker A and Speaker B. Speech recognition can then place words inside those segments, producing a transcript that is easier to follow than one uninterrupted block.
No. Diarization groups similar voice segments within a recording. Speaker identification or recognition attempts to match a voice to a known person, usually using enrollment data or verified context. A diarization label should not be treated as a person’s real identity.
No. The workbench initially displays a prepared two-speaker sample so you can inspect the format. After you reset it and start the browser microphone, live speech is labeled “You.” The page does not claim that the browser-only recognition path separates multiple speakers automatically.
Accuracy depends on microphone quality, room acoustics, speaker similarity, overlap, segment length, and the diarization system. This page does not publish an accuracy percentage for its prepared sample. For consequential use, compare every important attribution with the original audio and correct it manually.
Yes. The prepared example can be copied or downloaded as TXT with Speaker A and Speaker B labels. The export is plain text, not a subtitle file, editable timeline, audio package, or proof of identity. Review it before publishing or assigning statements to named people.
Use the live transcription page to test supported microphone recognition, read the meeting transcription page for an end-to-end review workflow, or open the prepared Muse Voice Transcribe demo for the original product example.
Review all four turns, copy the transcript, and check how speaker diarization labels appear in plain text. Reset the workspace only when you are ready to switch to single-speaker browser recognition.
Return to the speaker example